WO2026014266A1 - 情報処理装置、情報処理方法、情報処理システム及びプログラム - Google Patents
情報処理装置、情報処理方法、情報処理システム及びプログラムInfo
- Publication number
- WO2026014266A1 WO2026014266A1 PCT/JP2025/023179 JP2025023179W WO2026014266A1 WO 2026014266 A1 WO2026014266 A1 WO 2026014266A1 JP 2025023179 W JP2025023179 W JP 2025023179W WO 2026014266 A1 WO2026014266 A1 WO 2026014266A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- frame
- motion vector
- amount
- change
- information processing
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/124—Quantisation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/136—Incoming video signal characteristics or properties
- H04N19/14—Coding unit complexity, e.g. amount of activity or edge presence estimation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/503—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal prediction
- H04N19/51—Motion estimation or motion compensation
- H04N19/513—Processing of motion vectors
- H04N19/517—Processing of motion vectors by encoding
Definitions
- This disclosure relates to an information processing device, an information processing method, and an information processing system.
- Patent Document 1 discloses a technique for improving the accuracy of candidate predictive vectors for blocks adjacent in the temporal direction.
- this technique for a first coordinate in a block to be processed, multiple blocks are determined in adjacent pictures in the temporal direction, including the block closest to the first coordinate, and at least one motion vector is selected from the motion vectors held by each of the multiple determined blocks.
- Patent Document 1 does not address issues that arise when the data size of motion vectors increases, for example, in order to improve the accuracy of motion vectors.
- the purpose of this disclosure is to provide technology that can appropriately compress motion vectors used in video compression.
- a first aspect of the present disclosure provides an information processing device having: a calculation unit that calculates the amount of change in each pixel of a second frame based on a frame that is earlier than a first frame in a moving image relative to surrounding pixels; and a compression unit that compresses the motion vector with a code amount that corresponds to the amount of change at a position indicated by a motion vector that corresponds to one or more pixels of the second frame, estimated based on the first frame and the second frame.
- a second aspect of the present disclosure provides an information processing method that calculates the amount of change relative to surrounding pixels for each pixel of a second frame based on a frame that is earlier than a first frame in a moving image, and compresses the motion vector with a code amount that corresponds to the amount of change at the position indicated by the motion vector, which is estimated based on the first frame and the second frame and corresponds to one or more pixels of the second frame.
- a third aspect of the present disclosure provides an information processing device having: an acquisition unit that acquires information indicating the amount of change relative to surrounding pixels of each pixel of a second frame based on a frame that is earlier than a first frame in a moving image; and a motion vector that is estimated based on the first frame and the second frame and compressed with a code amount corresponding to the amount of change at a position indicated by a motion vector corresponding to one or more pixels of the second frame; and a restoration unit that restores the compressed motion vector acquired by the acquisition unit based on the amount of change.
- a fourth aspect of the present disclosure provides an information processing method that acquires information indicating the amount of change relative to surrounding pixels of each pixel of a second frame based on a frame that is earlier than a first frame in a moving image, and a motion vector estimated based on the first frame and the second frame and compressed with a code amount corresponding to the amount of change at a position indicated by the motion vector corresponding to one or more pixels of the second frame, and restores the compressed motion vector based on the amount of change.
- a fifth aspect of the present disclosure provides an information processing system having a first information processing device and a second information processing device, wherein the first information processing device has a calculation unit that calculates an amount of change for each pixel of a second frame based on a frame that is earlier than the first frame in a moving image relative to surrounding pixels, and a compression unit that compresses the motion vector with a code amount corresponding to the amount of change at a position indicated by a motion vector corresponding to one or more pixels of the second frame, estimated based on the first frame and the second frame, and the second information processing device has an acquisition unit that acquires the motion vector compressed with a code amount corresponding to the amount of change at a position indicated by the motion vector corresponding to one or more pixels of the second frame, and a restoration unit that restores the compressed motion vector acquired by the acquisition unit based on the amount of change.
- motion vectors used in video compression can be appropriately compressed.
- FIG. 1 is a diagram illustrating an example of a configuration of an information processing device that compresses a moving image according to an embodiment.
- FIG. 1 is a diagram illustrating an example of a configuration of an information processing device that restores a moving image according to an embodiment.
- FIG. 1 is a diagram illustrating an example of a configuration of an information processing system according to an embodiment.
- FIG. 1 is a diagram illustrating an example of a hardware configuration of an information processing apparatus according to an embodiment.
- FIG. 2 is a diagram illustrating an example of a video compression process of the information processing device according to the embodiment.
- 10A and 10B are diagrams illustrating an example of a compression process of a motion vector performed by the information processing device according to the embodiment.
- 10A and 10B are diagrams illustrating an example of a compression process of a motion vector performed by the information processing device according to the embodiment.
- FIG. 10 is a diagram illustrating an example of a video restoration process of the information processing device according to the embodiment.
- FIG. 1 is a diagram showing an example of the configuration of the information processing device 10 that compresses moving images according to an embodiment.
- the information processing device 10 includes a calculation unit 11 and a compression unit 12. Each of the calculation unit 11 and the compression unit 12 may be realized by cooperation between one or more programs installed in the information processing device 10 and hardware such as a processor and memory of the information processing device 10.
- the calculation unit 11 calculates the amount of change in each pixel of the second frame relative to surrounding pixels based on a frame that occurs earlier than the first frame in the video.
- the compression unit 12 compresses the motion vector with a code amount that corresponds to the amount of change at the position indicated by the motion vector, which corresponds to one or more pixels in the second frame and is estimated based on the first and second frames.
- FIG. 2 is a diagram showing an example of the configuration of the information processing device 20 that restores a moving image according to the embodiment.
- the information processing device 20 has an acquisition unit 21 and a restoration unit 22.
- Each of the acquisition unit 21 and the restoration unit 22 may be realized by cooperation between one or more programs installed in the information processing device 20 and hardware such as a processor and memory of the information processing device 20.
- the acquisition unit 21 acquires information indicating the amount of change relative to surrounding pixels for each pixel of the second frame based on a frame that occurs earlier than the first frame in the video, and motion vectors estimated based on the first and second frames and compressed with a code amount corresponding to the amount of change at the position indicated by the motion vector corresponding to one or more pixels in the second frame.
- the restoration unit 22 restores the compressed motion vectors acquired by the acquisition unit 21 based on the amount of change.
- FIG. 3 is a diagram illustrating an example of the configuration of the information processing system 1 according to an embodiment.
- the information processing system 1 includes an information processing device 10 and an information processing device 20.
- the number of information processing devices 10 and the information processing device 20 is not limited to the example of FIG. 3 .
- the technology of the present disclosure can be used in various services that encode and transmit video, such as, for example, distribution of stored video, compression of video for recording (saving) on a recording medium, and distribution of video in real time.
- the technology of the present disclosure can also be used in remote medical treatment, for example, in which video captured by a camera (imaging device) such as an endoscope is distributed to a remote specialist and instructions are received from the specialist.
- a camera imaging device
- the video of the present disclosure is not limited to video captured by an imaging device, but may also be, for example, a video of distance images captured by a stereo camera or LiDAR (Light Detection and Ranging).
- the video of the present disclosure may also be, for example, a video of thermal images captured by an infrared camera.
- information processing device 10 and information processing device 20 are connected so as to be able to communicate via network N.
- network N include, for example, the Internet, mobile information processing systems, wireless LANs (Local Area Networks), short-range wireless communications such as BLE (Bluetooth Low Energy, registered trademark), LANs, and buses.
- mobile information processing systems include, for example, fifth-generation mobile information processing systems (5G, 5th Generation), fourth-generation mobile information processing systems (4G), and third-generation mobile information processing systems (3G).
- Each of the information processing devices 10 and 20 may be, for example, a personal computer, a smartphone, a server, a cloud server, or other device.
- the information processing device 10 compresses video.
- the information processing device 10 may also transmit the compressed video to the information processing device 20.
- the information processing device 20 expands (restores) the video compressed by the information processing device 10.
- Fig. 4 is a diagram showing an example of the hardware configuration of the information processing device 10 and the information processing device 20 according to the embodiment.
- each of the information processing device 10 and the information processing device 20 includes a processor 101, a memory 102, and a communication interface 103. These components may be connected via a bus or the like.
- the memory 102 stores at least a part of a program 104.
- the communication interface 103 includes an interface required for communication with other network elements.
- the memory 102 may be of any type. As a non-limiting example, the memory 102 may be a non-transitory computer-readable storage medium. The memory 102 may also be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. While only one memory 102 is shown in the computer 100, several physically distinct memory modules may be present in the computer 100.
- the processor 101 may be of any type.
- the processor 101 may include one or more of a general-purpose computer, a special-purpose computer, a microprocessor, a digital signal processor (DSP), and, as a non-limiting example, a processor based on a multi-core processor architecture.
- the computer 100 may have multiple processors, such as application-specific integrated circuit chips that are time-slaved to a clock that synchronizes the main processor.
- Embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device.
- the present disclosure also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium.
- the computer program product includes computer-executable instructions, such as instructions included in program modules, that execute on a target real or virtual processor device to perform the processes or methods of the present disclosure.
- Program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform particular tasks or implement particular abstract data types.
- the functionality of the program modules may be combined or divided among program modules as desired in various embodiments.
- the machine-executable instructions of the program modules may be executed in local or distributed devices. In a distributed device, the program modules may be located in both local and remote storage media.
- Program code for carrying out the methods of the present disclosure may be written in any combination of one or more programming languages.
- the program code is provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus.
- the program code is executed by the processor or controller, the functions/acts in the flowcharts and/or implementing block diagrams are performed.
- the program code may be executed entirely on the machine, partly on the machine, as a standalone software package, partly on the machine and partly on a remote machine, or entirely on a remote machine or server.
- the program includes instructions (or software code) that, when loaded into a computer, cause the computer to perform one or more functions described in the embodiments.
- the program may be stored on a non-transitory computer-readable medium or a tangible storage medium.
- computer-readable medium or tangible storage medium includes random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technology, CD-ROM, digital versatile disc (DVD), Blu-ray (registered trademark) disc or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device.
- the program may also be transmitted on a transitory computer-readable medium or communication medium.
- transitory computer-readable medium or communication medium includes electrical, optical, acoustic, or other forms of propagated signals.
- FIG. 5 is a diagram showing an example of video compression processing by the information processing device 10 according to the embodiment. Note that, hereinafter, the same components are denoted by the same reference numerals to avoid redundant description.
- step S101 the compression unit 12 estimates a motion vector based on the frame to be compressed this time (first frame) 501 and the second frame 502, lossy compresses the estimated motion vector, and decompresses the compressed motion vector. This generates a restored (compressed and then decompressed) motion vector 503.
- the second frame 502 is, for example, a frame from an earlier point in time than the first frame 501, and may be a compressed and then decompressed frame. Details of the processing of step S101 in this case will be described later using Figure 6.
- the second frame 502 may be, for example, a frame that is earlier than the first frame 501, and may be a frame based on a compressed and restored frame. Details of the processing of step S101 in this case will be described later using Figure 7.
- the compression unit 12 performs motion compensation based on the restored (expanded after compression) motion vector 503 and the second frame 502, and generates a motion-compensated frame 504 (step S102).
- the compression unit 12 generates a residual image 505 by calculating the difference (residual) for each pixel between the first frame 501 and the motion-compensated frame 504 (step S103).
- the compression unit 12 may determine the code rate of the motion vector and the code rate of the residual using, for example, rate-distortion optimization (RDO).
- RDO rate-distortion optimization
- the accuracy of the motion vector restored using the technology disclosed herein is improved, and the code rate of the residual required to achieve a specific image quality for the restored frame can also be reduced. Therefore, the total value of the code rate of the motion vector and the code rate of the residual can be further reduced, thereby improving the video compression efficiency.
- the compression unit 12 then performs lossy compression on the residual image 505 and decompresses the compressed residual image (step S104).
- the compression unit 12 then adds, for each pixel, the residual image restored by the processing of step S104 and the motion-compensated frame 504 to generate the currently restored frame 506 (step S105). Note that the restored frame 506 is used as the previously restored frame 502 when compressing the next frame to be compressed.
- Fig. 6 is a diagram showing an example of the motion vector compression process of the information processing device 10 according to the embodiment.
- the compression unit 12 estimates a motion vector corresponding to one or more pixels of the second frame 502 based on the first frame 501 and the second frame 502.
- the compression unit 12 may, for example, use the learning results of deep learning to estimate (infer) a motion vector for each pixel or for each block containing multiple pixels based on the first frame 501 and the second frame 502. This can, for example, improve the accuracy of motion vector estimation.
- the calculation unit 11 calculates the amount of change for each of one or more pixels in the second frame 502 relative to its surrounding pixels (step S202). This generates a change amount MAP 602 indicating the amount of change for each of one or more pixels in the second frame 502 relative to its surrounding pixels.
- the calculation unit 11 may use, for example, a convolutional neural network (CNN) to calculate the amount of change for each pixel or for each block containing multiple pixels.
- CNN convolutional neural network
- the convolutional neural network may determine (estimate) a larger value for the amount of change according to at least one of the following conditions: the greater the change in color of the subject relative to the surrounding pixels, the smaller the subject appears, the greater the unevenness of the subject, the more in-focus the subject is, and the less blur there is in the subject.
- a relatively large value is determined for the amount of change in an area in which a person or the like that is relatively far away from the camera is captured.
- this convolutional neural network may be generated by supervised learning, for example, using an image as input (explanatory variable) and a change map as correct answer data, in which the change value is set to a larger value the higher the degree of 1 or higher as described above.
- the calculation unit 11 may calculate the amount of change for each pixel using, for example, a Laplacian filter, which is a filter that detects image edges using second-order derivatives. This makes it possible to set a larger value for the amount of change in areas with sharp edges, for example.
- a Laplacian filter which is a filter that detects image edges using second-order derivatives. This makes it possible to set a larger value for the amount of change in areas with sharp edges, for example.
- the compression unit 12 lossy compresses the motion vector with a code amount according to the amount of change at the position indicated by the estimated motion vector (step S203).
- the compression unit 12 may, for example, determine a larger code amount the greater the amount of change at the position indicated by the estimated motion vector. This reduces the degree of compression of the motion vector at locations with greater amount of change, thereby reducing the deterioration in accuracy of the motion vector due to compression of the motion vector.
- the compression unit 12 may, for example, determine a smaller QP (Quantization Parameter) value or quantization width to be used for motion vector compression the greater the amount of change at the position indicated by the estimated motion vector.
- the position indicated by the motion vector may, for example, be the position of the data to be referenced in the second frame 502.
- the direction of the motion vector may be calculated as the direction from a certain pixel in the first frame 501 toward the original position of that pixel in the second frame 502.
- the motion vector has an end point at the position of the certain pixel in the second frame 502 and a start point at the position to which that pixel has moved in the first frame 501.
- the position indicated by the motion vector is the end point of the motion vector.
- the direction of movement of a certain pixel in the second frame 502 and the direction toward the original position of a certain pixel in the first frame 501 are opposite directions.
- the direction of the motion vector may be defined or calculated using either method.
- the compression unit 12 may, for example, compress a motion vector with a code amount corresponding to the amount of change for the pixel at the position indicated by the motion vector.
- the compression unit 12 may also, for example, compress a motion vector with a code amount corresponding to the amount of change for each of multiple pixels surrounding the position indicated by the motion vector.
- the compression unit 12 may, for example, determine the code amount by taking a weighted average of the amount of change for the pixel at the position indicated by the motion vector and the amount of change for each of a specific number of pixels surrounding the position indicated by the motion vector, with weighting values that increase the closer the pixel is to the position indicated by the motion vector.
- the compression unit 12 expands the compressed motion vector 603 (step S204). This generates the restored motion vector 503.
- Part 2>>> 6 has described an example in which a frame that is an earlier frame than the first frame 501 and that has been compressed and restored is used as the second frame 502.
- FIG. 7 has described an example in which a frame that is an earlier frame than the first frame 501 and that is based on a frame that has been compressed and restored is used as the second frame 502. In the example of FIG. 7 , for example, the processing load increases compared to the example of FIG. 6 , but compression efficiency can be further improved.
- Figure 7 is a diagram showing an example of the motion vector compression process of the information processing device 10 according to the embodiment.
- the compression unit 12 In step S301, the compression unit 12 generates a predicted motion vector 702 by predicting a motion vector for a previously restored frame and the first frame 501, which is the frame to be compressed this time.
- the compression unit 12 may generate the predicted motion vector 702 based on data 701, such as a plurality of previously restored frames and a motion vector based on the plurality of frames and previously restored motion vectors.
- the compression unit 12 performs motion compensation based on the previously restored frame 703 and the predicted motion vector 702 to generate the second frame 502, which is a predicted frame for the first frame 501 (step S302).
- the compression unit 12 estimates a motion vector corresponding to one or more pixels of the second frame based on the first frame 501 and the second frame 502 (step S303).
- the calculation unit 11 calculates the amount of change for each of one or more pixels of the second frame 502 relative to surrounding pixels (step S304).
- the compression unit 12 lossy compresses the motion vector with a code amount corresponding to the amount of change at the position indicated by the estimated motion vector (step S305).
- the compression unit 12 expands the compressed motion vector 705 (step S306).
- the processes of steps S303 to S306 may be the same as the processes of steps S201 to S204 in FIG. 6.
- the process of step S306 generates a motion vector 706 based on the first frame 501, which is the frame to be compressed this time, and the second frame 502, which is the predicted frame.
- the compression unit 12 then performs motion compensation processing based on the predicted motion vector 702 and the motion vector 706 (step S307).
- the compression unit 12 then adds the motion vector 706 to the compensation processing result (step S308). This generates the restored motion vector 503.
- FIG. 8 is a diagram showing an example of video restoration processing by the information processing device 20 according to the embodiment.
- step S401 the acquisition unit 21 acquires a motion vector compressed with a code amount corresponding to the amount of change at the position indicated by the motion vector corresponding to one or more pixels in the second frame 502, and a compressed residual image 505A.
- the motion vector compressed with a code amount corresponding to the amount of change at the position indicated by the motion vector corresponding to one or more pixels in the second frame 502 is motion vector 603 in FIG. 6 or motion vector 705 in FIG. 7.
- the second frame 502 may be, for example, a frame that is earlier than the first frame 501 described in FIG. 6 and that has been compressed and restored.
- the second frame 502 may be, for example, a frame that is earlier than the first frame 501 described in FIG. 7 and that is based on a frame that has been compressed and restored.
- the restoration unit 22 expands the compressed motion vector acquired by the acquisition unit 21 based on the amount of change in each pixel of the previously restored second frame 502 relative to its surrounding pixels (step S402).
- the processing of step S402 may be the same as the processing of step S204 in FIG. 6 or the processing of step S306 in FIG. 7.
- This generates the restored motion vector 503.
- the information indicating the amount of change in each pixel of the previously restored second frame 502 relative to its surrounding pixels is, for example, the change amount MAP 602 in FIG. 6 or FIG. 7.
- the restoration unit 22 performs motion compensation based on the restored motion vector 503 and the second frame 502, and generates a motion-compensated frame 504 (step S403).
- the restoration unit 22 then decompresses the compressed residual image 505A (step S404).
- the restoration unit 22 then adds the restored residual image 505 and the motion-compensated frame 504 for each pixel to generate the currently restored frame 506 (step S405). Note that the restored frame 506 is used as the previously restored frame 502 when restoring the next frame to be restored.
- Compressed video data includes compressed data for motion vectors for each frame. Therefore, if the data size of the motion vectors increases, the data size of the compressed video will also increase if the compression rate of the motion vectors remains the same as before. Furthermore, if the compression rate of the motion vectors is increased, the accuracy of the restored motion vectors will decrease.
- the technology disclosed herein allows for appropriate compression of motion vectors. As a result, for example, it is possible to improve the video compression rate while maintaining video quality. It is also possible to improve video quality while maintaining the same video compression rate as before.
- the information processing device 10 and the information processing device 20 may each be a device included in a single housing, but the information processing device 10 and the information processing device 20 of the present disclosure are not limited to this.
- Each unit of the information processing device 10 and the information processing device 20 may be realized, for example, by cloud computing configured with one or more computers.
- at least a portion of the processing of the information processing device 10 may be executed, for example, by the information processing device 20.
- at least a portion of the processing of the information processing device 20 may be executed, for example, by the information processing device 10.
- the information processing device 10 and the information processing device 20 may be an integrated device.
- Such information processing devices 10 and information processing devices 20 are also included in examples of the "information processing device" of the present disclosure.
- (Appendix 1) a calculation unit that calculates a change amount of each pixel of a second frame based on a frame that is earlier than the first frame in the moving image, relative to surrounding pixels; a compression unit that compresses the motion vector with a code amount corresponding to the amount of change at a position indicated by the motion vector corresponding to one or more pixels of the second frame, the motion vector being estimated based on the first frame and the second frame; An information processing device having the above.
- the compression unit compresses the motion vector with a code amount corresponding to at least one of the amount of change with respect to a pixel at a position indicated by the motion vector and the amount of change with respect to each of a plurality of pixels around the position indicated by the motion vector.
- the information processing device calculates the amount of change for each pixel using a convolutional neural network that determines the value of the amount of change to be larger according to at least one of the conditions that the smaller the subject is photographed, the more in focus the subject is, and the less blur of the subject is present; 3.
- the information processing device according to claim 1 or 2.
- the calculation unit calculates the amount of change for each pixel using a convolutional neural network that determines the amount of change to be a larger value in accordance with at least one of a larger change in color of the subject relative to surrounding pixels and a larger unevenness of the subject; 3.
- the information processing device according to claim 1 or 2.
- an acquisition unit that acquires a motion vector that is estimated based on a first frame of a moving image and a second frame that is based on a frame that is earlier than the first frame, and that is compressed with a code amount that corresponds to an amount of change in each pixel of the second frame relative to surrounding pixels at a position indicated by the motion vector corresponding to one or more pixels of the second frame; a restoration unit that restores the compressed motion vector acquired by the acquisition unit based on the amount of change;
- a first information processing device and a second information processing device are included,
- the first information processing device a calculation unit that calculates a change amount of each pixel of a second frame based on a frame that is earlier than the first frame in the moving image, relative to surrounding pixels;
- a compression unit that compresses the motion vector with a code amount corresponding to the amount of change at a position indicated by the motion vector corresponding to one or more pixels of the second frame, the motion vector being estimated based on the first frame and the second frame;
- the second information processing device an acquisition unit that acquires the motion vectors corresponding to one or more pixels of the second frame, the motion vectors being compressed with a code amount corresponding to the amount of change at the positions indicated by the motion vectors; a restoration unit that restores the compressed motion vector acquired by the acquisition unit based on the amount of change;
- Appendix 12 a motion vector estimated based on a first frame of a moving image and a second frame based on a frame earlier than the first frame, the motion vector being compressed with a code amount corresponding to an amount of change in each pixel of the second frame relative to surrounding pixels at a position indicated by the motion vector corresponding to one or more pixels of the second frame; restoring the compressed motion vector based on the amount of change;
- a program that causes a computer to perform a process.
- Information processing system 10 Information processing device 11 Calculation unit 12 Compression unit 20 Information processing device 21 Acquisition unit 22 Restoration unit
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出する算出部と、前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する圧縮部と、を有する情報処理装置が提供される。これにより、動画圧縮に用いられる動きベクトルを適切に圧縮することができる。
Description
本開示は、情報処理装置、情報処理方法、及び情報処理システムに関する。
特許文献1には、時間方向に隣接するブロックの予測ベクトル候補の精度を上げる技術が開示されている。特許文献1では、処理対象ブロック内の第1座標に対し、時間方向に隣接するピクチャ内で第1座標に最も近いブロックを含む複数のブロックを決定し、決定した複数のブロックがそれぞれ有する動きベクトルの中から、少なくとも1つの動きベクトルを選択する。
しかしながら、特許文献1に記載の技術では、例えば、動きベクトルの精度を向上させる等のために動きベクトルのデータサイズが増大する場合の課題については検討されていない。
本開示の目的は、上述した課題を鑑み、動画圧縮に用いられる動きベクトルを適切に圧縮できる技術を提供することにある。
本開示に係る第1の態様では、動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出する算出部と、前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する圧縮部と、を有する情報処理装置が提供される。
また、本開示に係る第2の態様では、動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出し、前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する、情報処理方法が提供される。
また、本開示に係る第3の態様では、動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を示す情報と、前記第1フレームと前記第2フレームとに基づいて推定され、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で圧縮された前記動きベクトルと、を取得する取得部と、前記取得部により取得された、圧縮された前記動きベクトルを、前記変化量に基づいて復元する復元部と、を有する情報処理装置が提供される。
また、本開示に係る第4の態様では、動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を示す情報と、前記第1フレームと前記第2フレームとに基づいて推定され、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で圧縮された前記動きベクトルと、を取得し、圧縮された前記動きベクトルを、前記変化量に基づいて復元する、情報処理方法が提供される。
また、本開示に係る第5の態様では、第1情報処理装置と第2情報処理装置とを有し、前記第1情報処理装置は、動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出する算出部と、前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する圧縮部と、を有し、前記第2情報処理装置は、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で圧縮された前記動きベクトルと、を取得する取得部と、前記取得部により取得された、圧縮された前記動きベクトルを、前記変化量に基づいて復元する復元部と、を有する情報処理システムが提供される。
一側面によれば、動画圧縮に用いられる動きベクトルを適切に圧縮できる。
本開示の原理は、いくつかの例示的な実施形態を参照して説明される。これらの実施形態は、例示のみを目的として記載されており、本開示の範囲に関する制限を示唆することなく、当業者が本開示を理解および実施するのを助けることを理解されたい。本明細書で説明される開示は、以下で説明されるもの以外の様々な方法で実装される。
以下の説明および特許請求の範囲において、他に定義されない限り、本明細書で使用されるすべての技術用語および科学用語は、本開示が属する技術分野の当業者によって一般に理解されるのと同じ意味を有する。
以下、図面を参照して、本開示の実施形態を説明する。なお、各図面は、1又はそれ以上の実施形態を説明するための単なる例示である。各図面は、1つの特定の実施形態のみに関連付けられるのではなく、1又はそれ以上の他の実施形態に関連付けられてもよい。当業者であれば理解できるように、いずれか1つの図面を参照して説明される様々な特徴又はステップは、例えば明示的に図示または説明されていない実施形態を作り出すために、1又はそれ以上の他の図に示された特徴又はステップと組み合わせることができる。例示的な実施形態を説明するためにいずれか1つの図に示された特徴またはステップのすべてが必ずしも必須ではなく、一部の特徴またはステップが省略されてもよい。いずれかの図に記載されたステップの順序は、適宜変更されてもよい。
(実施の形態1)
<構成>
<<情報処理装置10の構成>>
図1を参照し、実施形態に係る動画を圧縮する情報処理装置10(エンコーダ)の構成について説明する。図1は、実施形態に係る動画を圧縮する情報処理装置10の構成の一例を示す図である。情報処理装置10は、算出部11、及び圧縮部12を有する。算出部11及び圧縮部12のそれぞれは、情報処理装置10にインストールされた1以上のプログラムと、情報処理装置10のプロセッサ、及びメモリ等のハードウェアとの協働により実現されてもよい。
<構成>
<<情報処理装置10の構成>>
図1を参照し、実施形態に係る動画を圧縮する情報処理装置10(エンコーダ)の構成について説明する。図1は、実施形態に係る動画を圧縮する情報処理装置10の構成の一例を示す図である。情報処理装置10は、算出部11、及び圧縮部12を有する。算出部11及び圧縮部12のそれぞれは、情報処理装置10にインストールされた1以上のプログラムと、情報処理装置10のプロセッサ、及びメモリ等のハードウェアとの協働により実現されてもよい。
算出部11は、動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出する。圧縮部12は、第1フレームと第2フレームとに基づいて推定された、第2フレームの1以上の各画素に応じた動きベクトルが指す位置での変化量に応じた符号量で、動きベクトルを圧縮する。
<<情報処理装置20の構成>>
次に、図2を参照し、実施形態に係る動画を復元する情報処理装置20(デコーダ)の構成について説明する。図2は、実施形態に係る動画を復元する情報処理装置20の構成の一例を示す図である。情報処理装置20は、取得部21、及び復元部22を有する。取得部21及び復元部22のそれぞれは、情報処理装置20にインストールされた1以上のプログラムと、情報処理装置20のプロセッサ、及びメモリ等のハードウェアとの協働により実現されてもよい。
次に、図2を参照し、実施形態に係る動画を復元する情報処理装置20(デコーダ)の構成について説明する。図2は、実施形態に係る動画を復元する情報処理装置20の構成の一例を示す図である。情報処理装置20は、取得部21、及び復元部22を有する。取得部21及び復元部22のそれぞれは、情報処理装置20にインストールされた1以上のプログラムと、情報処理装置20のプロセッサ、及びメモリ等のハードウェアとの協働により実現されてもよい。
取得部21は、動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を示す情報と、第1フレームと第2フレームとに基づいて推定され、第2フレームの1以上の各画素に応じた動きベクトルが指す位置での変化量に応じた符号量で圧縮された動きベクトルと、を取得する。復元部22は、取得部21により取得された、圧縮された動きベクトルを、上述した変化量に基づいて復元する。
<システム構成>
次に、図3を参照し、実施形態に係る情報処理システム1の構成について説明する。図3は、実施形態に係る情報処理システム1の構成の一例を示す図である。図3の例では、情報処理システム1は、情報処理装置10、及び情報処理装置20を有する。なお、情報処理装置10、及び情報処理装置20の数は図3の例に限定されない。なお、本開示の技術は、例えば、蓄積している動画の配信、記録媒体へ記録(保存)するための動画の圧縮、及びリアルタイムでの動画の配信等の、動画を符号化して伝送する各種のサービスで用いることができる。この場合、本開示の技術は、例えば、内視鏡等のカメラ(撮影装置)で撮影された動画を、遠隔地の専門家へ配信し、当該専門家からの指示を受ける遠隔診療等に用いることもできる。また、本開示の動画は、撮影装置で撮影された動画に限らず、例えば、ステレオカメラやLiDAR(Light Detection And Ranging)で撮影された距離画像による動画でもよい。また、本開示の動画は、例えば、赤外線カメラで撮影された熱画像による動画でもよい。
次に、図3を参照し、実施形態に係る情報処理システム1の構成について説明する。図3は、実施形態に係る情報処理システム1の構成の一例を示す図である。図3の例では、情報処理システム1は、情報処理装置10、及び情報処理装置20を有する。なお、情報処理装置10、及び情報処理装置20の数は図3の例に限定されない。なお、本開示の技術は、例えば、蓄積している動画の配信、記録媒体へ記録(保存)するための動画の圧縮、及びリアルタイムでの動画の配信等の、動画を符号化して伝送する各種のサービスで用いることができる。この場合、本開示の技術は、例えば、内視鏡等のカメラ(撮影装置)で撮影された動画を、遠隔地の専門家へ配信し、当該専門家からの指示を受ける遠隔診療等に用いることもできる。また、本開示の動画は、撮影装置で撮影された動画に限らず、例えば、ステレオカメラやLiDAR(Light Detection And Ranging)で撮影された距離画像による動画でもよい。また、本開示の動画は、例えば、赤外線カメラで撮影された熱画像による動画でもよい。
図3の例では、情報処理装置10、及び情報処理装置20は、ネットワークNにより通信できるように接続されている。ネットワークNの例には、例えば、インターネット、移動情報処理システム、無線LAN(Local Area Network)、BLE(Bluetooth Low Energy、登録商標)等の近距離無線通信、LAN、及びバス等が含まれる。移動情報処理システムの例には、例えば、第5世代移動情報処理システム(5G、5th Generation)、第4世代移動情報処理システム(4G)、第3世代移動情報処理システム(3G)等が含まれる。
情報処理装置10及び情報処理装置20のそれぞれは、例えば、パーソナルコンピュータ、スマートフォン、サーバ、クラウドサーバ等の装置でもよい。情報処理装置10は、動画を圧縮する。また、情報処理装置10は、圧縮した動画を情報処理装置20へ送信してもよい。情報処理装置20は、情報処理装置10により圧縮された動画を伸長(復元)する。
<ハードウェア構成>
図4は、実施形態に係る情報処理装置10及び情報処理装置20のハードウェア構成例を示す図である。図4の例では、情報処理装置10及び情報処理装置20のそれぞれ(コンピュータ100)は、プロセッサ101、メモリ102、通信インターフェイス103を含む。これら各部は、バス等により接続されてもよい。メモリ102は、プログラム104の少なくとも一部を格納する。通信インターフェイス103は、他のネットワーク要素との通信に必要なインターフェイスを含む。
図4は、実施形態に係る情報処理装置10及び情報処理装置20のハードウェア構成例を示す図である。図4の例では、情報処理装置10及び情報処理装置20のそれぞれ(コンピュータ100)は、プロセッサ101、メモリ102、通信インターフェイス103を含む。これら各部は、バス等により接続されてもよい。メモリ102は、プログラム104の少なくとも一部を格納する。通信インターフェイス103は、他のネットワーク要素との通信に必要なインターフェイスを含む。
プログラム104が、プロセッサ101及びメモリ102等の協働により実行されると、コンピュータ100により本開示の実施形態の少なくとも一部の処理が行われる。メモリ102は、任意のタイプのものであってもよい。メモリ102は、非限定的な例として、非一時的なコンピュータ可読記憶媒体でもよい。また、メモリ102は、半導体ベースのメモリデバイス、磁気メモリデバイスおよびシステム、光学メモリデバイスおよびシステム、固定メモリおよびリムーバブルメモリなどの任意の適切なデータストレージ技術を使用して実装されてもよい。コンピュータ100には1つのメモリ102のみが示されているが、コンピュータ100にはいくつかの物理的に異なるメモリモジュールが存在してもよい。プロセッサ101は、任意のタイプのものであってよい。プロセッサ101は、汎用コンピュータ、専用コンピュータ、マイクロプロセッサ、デジタル信号プロセッサ(DSP:Digital Signal Processor)、および非限定的な例としてマルチコアプロセッサアーキテクチャに基づくプロセッサの1つ以上を含んでよい。コンピュータ100は、メインプロセッサを同期させるクロックに時間的に従属する特定用途向け集積回路チップなどの複数のプロセッサを有してもよい。
本開示の実施形態は、ハードウェアまたは専用回路、ソフトウェア、ロジックまたはそれらの任意の組み合わせで実装され得る。いくつかの態様はハードウェアで実装されてもよく、一方、他の態様はコントローラ、マイクロプロセッサまたは他のコンピューティングデバイスによって実行され得るファームウェアまたはソフトウェアで実装されてもよい。
本開示はまた、非一時的なコンピュータ可読記憶媒体に有形に記憶された少なくとも1つのコンピュータプログラム製品を提供する。コンピュータプログラム製品は、プログラムモジュールに含まれる命令などのコンピュータ実行可能命令を含み、対象の実プロセッサまたは仮想プロセッサ上のデバイスで実行され、本開示のプロセスまたは方法を実行する。プログラムモジュールには、特定のタスクを実行したり、特定の抽象データ型を実装したりするルーチン、プログラム、ライブラリ、オブジェクト、クラス、コンポーネント、データ構造などが含まれる。プログラムモジュールの機能は、様々な実施形態で望まれるようにプログラムモジュール間で結合または分割されてもよい。プログラムモジュールのマシン実行可能命令は、ローカルまたは分散デバイス内で実行できる。分散デバイスでは、プログラムモジュールはローカルとリモートの両方のストレージメディアに配置できる。
本開示の方法を実行するためのプログラムコードは、1つ以上のプログラミング言語の任意の組み合わせで書かれてもよい。これらのプログラムコードは、汎用コンピュータ、専用コンピュータ、またはその他のプログラム可能なデータ処理装置のプロセッサまたはコントローラに提供される。プログラムコードがプロセッサまたはコントローラによって実行されると、フローチャートおよび/または実装するブロック図内の機能/動作が実行される。プログラムコードは、完全にマシン上で実行され、一部はマシン上で、スタンドアロンソフトウェアパッケージとして、一部はマシン上で、一部はリモートマシン上で、または完全にリモートマシンまたはサーバ上で実行される。
プログラムは、コンピュータに読み込まれた場合に、実施形態で説明された1又はそれ以上の機能をコンピュータに行わせるための命令群(又はソフトウェアコード)を含む。プログラムは、非一時的なコンピュータ可読媒体又は実体のある記憶媒体に格納されてもよい。限定ではなく例として、コンピュータ可読媒体又は実体のある記憶媒体は、random-access memory(RAM)、read-only memory(ROM)、フラッシュメモリ、solid-state drive(SSD)又はその他のメモリ技術、CD-ROM、digital versatile disc(DVD)、Blu-ray(登録商標)ディスク又はその他の光ディスクストレージ、磁気カセット、磁気テープ、磁気ディスクストレージ又はその他の磁気ストレージデバイスを含む。プログラムは、一時的なコンピュータ可読媒体又は通信媒体上で送信されてもよい。限定ではなく例として、一時的なコンピュータ可読媒体又は通信媒体は、電気的、光学的、音響的、またはその他の形式の伝搬信号を含む。
<処理>
<<動画圧縮(エンコード)処理>>
次に、図5を参照し、実施形態に係る情報処理装置10の動画圧縮処理の一例について説明する。図5は、実施形態に係る情報処理装置10の動画圧縮処理の一例を示す図である。なお、以下では、同一のものには同一の符号を付すことにより、説明の重複を省略する。
<<動画圧縮(エンコード)処理>>
次に、図5を参照し、実施形態に係る情報処理装置10の動画圧縮処理の一例について説明する。図5は、実施形態に係る情報処理装置10の動画圧縮処理の一例を示す図である。なお、以下では、同一のものには同一の符号を付すことにより、説明の重複を省略する。
ステップS101において、圧縮部12は、今回の圧縮対象のフレーム(第1フレーム)501と、第2フレーム502とに基づいて動きベクトルを推定し、推定した動きベクトルを非可逆圧縮し、圧縮した動きベクトルを伸長する。これにより、復元(圧縮後に伸長)された動きベクトル503が生成される。第2フレーム502は、例えば、第1フレーム501よりも以前の時点のフレームであり、圧縮して復元したフレームでもよい。なお、この場合のステップS101の処理の詳細について、図6を用いて後述する。
また、第2フレーム502は、例えば、第1フレーム501よりも以前の時点のフレームであり、圧縮して復元したフレームに基づくフレームでもよい。なお、この場合のステップS101の処理の詳細について、図7を用いて後述する。
続いて、圧縮部12は、復元(圧縮後に伸長)された動きベクトル503と、第2フレーム502とに基づいて、動き補償を行い、動き補償後のフレーム504を生成する(ステップS102)。
続いて、圧縮部12は、第1フレーム501と、動き補償後のフレーム504との各画素毎の差(残差)を算出することにより残差画像505を生成する(ステップS103)。なお、圧縮部12は、例えば、レート歪み最適化(RDO、Rate-distortion optimization)を用いて、動きベクトルの符号量及び残差の符号量を決定してもよい。この場合、本開示の技術により復元された動きベクトルの精度が向上するため、復元されたフレームの画質を特定の画質とするために必要な残差の符号量も低減できる。そのため、動きベクトルの符号量と残差の符号量との合計値をより低減できるため動画の圧縮効率を向上できる。
続いて、圧縮部12は、残差画像505を非可逆圧縮し、圧縮した残差画像を伸長する(ステップS104)。続いて、圧縮部12は、ステップS104の処理により復元した残差画像と、動き補償後のフレーム504とを各画素毎に加算することにより、今回復元したフレーム506を生成する(ステップS105)。なお、復元したフレーム506は、次回の圧縮対象のフレームを圧縮する際に、以前に復元したフレーム502として用いられる。
<<<動きベクトルの圧縮処理(その1)>>>
次に、図6を参照し、図5のステップS101における、実施形態に係る情報処理装置10の動きベクトルの圧縮処理の一例について説明する。図6は、実施形態に係る情報処理装置10の動きベクトルの圧縮処理の一例を示す図である。
次に、図6を参照し、図5のステップS101における、実施形態に係る情報処理装置10の動きベクトルの圧縮処理の一例について説明する。図6は、実施形態に係る情報処理装置10の動きベクトルの圧縮処理の一例を示す図である。
ステップS201において、圧縮部12は、第1フレーム501と、第2フレーム502とに基づいて、第2フレーム502の1以上の各画素に応じた動きベクトルを推定する。ここで、圧縮部12は、例えば、ディープラーニング(深層学習)の学習結果を用いて、第1フレーム501と第2フレーム502とに基づいて、各画素毎、または複数の画素を含むブロック毎の動きベクトルを推定(推論)してもよい。これにより、例えば、動きベクトルの推定精度を向上できる。
続いて、算出部11は、第2フレーム502の1以上の各画素の周辺画素に対する変化量を算出する(ステップS202)。これにより、第2フレーム502の1以上の各画素の周辺画素に対する変化量を示す変化量MAP602が生成される。ここで、算出部11は、例えば、畳み込みニューラルネットワーク(CNN、Convolutional Neural Network)を用いて、各画素毎、または複数の画素を含むブロック毎の変化量を算出してもよい。当該畳み込みニューラルネットワークは、周辺の画素に対して被写体の色の変化が大きいほど、被写体が小さく写されているほど、被写体の凹凸が大きいほど、被写体にピントが合っているほど、及び被写体のブレが少ないほどの少なくとも一つの条件に従って、変化量の値をより大きな値に決定(推定)してもよい。この場合、例えば、カメラから見て比較的遠くに存在する人物等が写されている箇所は、変化量の値が比較的大きな値に決定される。
これにより、例えば、周辺画素と比較して比較的わずかに位置が異なるだけで残差が比較的大きくなる箇所について、変化量の値をより大きな値に決定できる。なお、当該畳み込みニューラルネットワークは、例えば、画像を入力(説明変数)とし、上述した1以上の度合いが高いほど変化量の値をより大きな値に設定された変化量マップを正解データとした教師あり学習により生成されてもよい。
算出部11は、例えば、2次微分で画像のエッジを検出するフィルタであるラプラシアンフィルタを用いて、各画素に対する変化量を算出してもよい。これにより、例えば、エッジが急激な箇所に対して変化量の値をより大きな値に決定できる。
続いて、圧縮部12は、推定した動きベクトルが指す位置での変化量に応じた符号量で、動きベクトルを非可逆圧縮する(ステップS203)。ここで、圧縮部12は、例えば、推定した動きベクトルが指す位置での変化量が大きいほど、符号量を大きく決定してもよい。これにより、例えば、変化量が大きい箇所ほど動きベクトルの圧縮度合いが低減されることにより、動きベクトルの圧縮による動きベクトルの精度の劣化を低減できる。この場合、圧縮部12は、例えば、推定した動きベクトルが指す位置での変化量が大きいほど、動きベクトル圧縮に使用するQP(Quantization Parameter、量子化パラメーター)値または量子化幅を小さく決定してもよい。動きベクトルが指す位置は、例えば、第2フレーム502における引用するデータの箇所の位置でもよい。
また、動きベクトルの向きは、第1フレーム501のあるピクセルから見た第2フレーム502における当該ピクセルの移動元の位置への方向として算出されてもよい。この場合、動きベクトルは、第2フレーム502のあるピクセルの位置が終点とされ、第1フレーム501における当該ピクセルが移動した先の位置が始点とされる。この場合、動きベクトルが指す位置は、動きベクトルの終点の位置となる。なお、第2フレーム502のあるピクセルの動きの方向と、第1フレーム501のあるピクセルの移動元の位置への方向とは、真逆の向きとなる。本実施形態では、動きベクトルの向きは、どちらの方法で定義または算出されてもよい。
圧縮部12は、例えば、動きベクトルが指す位置の画素に対する変化量に応じた符号量で動きベクトルを圧縮してもよい。また、圧縮部12は、例えば、動きベクトルが指す位置の周辺の複数の画素のそれぞれに対する変化量に応じた符号量で動きベクトルを圧縮してもよい。この場合、圧縮部12は、例えば、動きベクトルが指す位置の画素に対する変化量と、動きベクトルが指す位置の周辺の特定数の各画素に対する変化量とを、動きベクトルが指す位置からの距離が近いほど値が大きい重みを付けた加重平均により、当該符号量を決定してもよい。
続いて、圧縮部12は、圧縮した動きベクトル603を伸長する(ステップS204)。これにより、復元された動きベクトル503が生成される。
<<<動きベクトルの圧縮処理(その2)>>>
図6では、第1フレーム501よりも以前の時点のフレームであり、圧縮して復元したフレームを第2フレーム502として用いる例について説明した。図7では、第1フレーム501よりも以前の時点のフレームであり、圧縮して復元したフレームに基づくフレームを第2フレーム502として用いる例について説明する。図7の例では、図6の例と比較して、例えば、処理負荷は増大するものの、圧縮効率をより向上できる。
図6では、第1フレーム501よりも以前の時点のフレームであり、圧縮して復元したフレームを第2フレーム502として用いる例について説明した。図7では、第1フレーム501よりも以前の時点のフレームであり、圧縮して復元したフレームに基づくフレームを第2フレーム502として用いる例について説明する。図7の例では、図6の例と比較して、例えば、処理負荷は増大するものの、圧縮効率をより向上できる。
図7を参照し、図5のステップS101における、実施形態に係る情報処理装置10の動きベクトルの圧縮処理の他の一例について説明する。図7は、実施形態に係る情報処理装置10の動きベクトルの圧縮処理の一例を示す図である。
ステップS301において、圧縮部12は、以前に復元したフレームと今回の圧縮対象のフレームである第1フレーム501とに対する動きベクトルを予測することにより、予測動きベクトル702を生成する。ここで、圧縮部12は、例えば、以前に復元した複数のフレーム、及び当該複数のフレームに基づく動きベクトルであって以前に復元した動きベクトル等のデータ701に基づいて、予測動きベクトル702を生成してもよい。
続いて、圧縮部12は、以前に復元したフレーム703と予測動きベクトル702とに基づいて、動き補償を行うことにより、第1フレーム501に対する予測フレームである第2フレーム502を生成する(ステップS302)。
続いて、圧縮部12は、第1フレーム501と、第2フレーム502とに基づいて、第2フレームの1以上の各画素に応じた動きベクトルを推定する(ステップS303)。続いて、算出部11は、第2フレーム502の1以上の各画素の周辺画素に対する変化量を算出する(ステップS304)。続いて、圧縮部12は、推定した動きベクトルが指す位置での変化量に応じた符号量で、動きベクトルを非可逆圧縮する(ステップS305)。続いて、圧縮部12は、圧縮した動きベクトル705を伸長する(ステップS306)。なお、ステップS303からステップS306の各処理は、図6のステップS201からステップS204の各処理と同様でもよい。なお、ステップS306の処理により、今回の圧縮対象のフレームである第1フレーム501と予測フレームである第2フレーム502とに基づく動きベクトル706が生成される。
続いて、圧縮部12は、予測動きベクトル702と動きベクトル706とに基づいて動き補償処理を行う(ステップS307)。続いて、圧縮部12は、補償処理結果に動きベクトル706を加算する(ステップS308)。これにより、復元された動きベクトル503が生成される。
<<動画復元(デコード)処理>>
次に、図8を参照し、実施形態に係る情報処理装置20の動画復元処理の一例について説明する。図8は、実施形態に係る情報処理装置20の動画復元処理の一例を示す図である。
次に、図8を参照し、実施形態に係る情報処理装置20の動画復元処理の一例について説明する。図8は、実施形態に係る情報処理装置20の動画復元処理の一例を示す図である。
ステップS401において、取得部21は、第2フレーム502の1以上の各画素に応じた動きベクトルが指す位置での変化量に応じた符号量で圧縮された動きベクトルと、圧縮された残差画像505Aとを取得する。なお、第2フレーム502の1以上の各画素に応じた動きベクトルが指す位置での変化量に応じた符号量で圧縮された動きベクトルは、図6の動きベクトル603または図7の動きベクトル705である。なお、第2フレーム502は、例えば、図6で説明した、第1フレーム501よりも以前の時点のフレームであり、圧縮して復元したフレームでもよい。また、第2フレーム502は、例えば、図7で説明した、第1フレーム501よりも以前の時点のフレームであり、圧縮して復元したフレームに基づくフレームでもよい。
続いて、復元部22は、取得部21により取得された、圧縮された動きベクトルを、以前に復元した第2フレーム502の各画素の周辺画素に対する変化量に基づいて伸長する(ステップS402)。なお、ステップS402の処理は、図6のステップS204の処理、または図7のステップS306の処理と同様でもよい。これにより、復元された動きベクトル503が生成される。なお、以前に復元した第2フレーム502の各画素の周辺画素に対する変化量を示す情報は、例えば、図6または図7の変化量MAP602である。
続いて、復元部22は、復元された動きベクトル503と、第2フレーム502とに基づいて、動き補償を行い、動き補償後のフレーム504を生成する(ステップS403)。
続いて、復元部22は、圧縮された残差画像505Aを伸長する(ステップS404)。続いて、復元部22は、復元した残差画像505と、動き補償後のフレーム504とを各画素毎に加算することにより、今回復元したフレーム506を生成する(ステップS405)。なお、復元したフレーム506は、次回の復元対象のフレームを復元する際に、以前に復元したフレーム502として用いられる。
<その他>
近年研究が進んでいるディープラーニング(深層学習)を用いて、今回の圧縮対象のフレーム501と、以前に復元したフレーム502とに基づいて、各画素毎、または複数の画素を含むブロック毎の動きベクトルを推定(推論)することが考えられる。この場合、従来手法と比較して、推定精度は向上するものの、画素毎または比較的小さいブロック毎の動きベクトルが推定されるため、動きベクトルが高精細化されるためデータサイズが増大化する。
近年研究が進んでいるディープラーニング(深層学習)を用いて、今回の圧縮対象のフレーム501と、以前に復元したフレーム502とに基づいて、各画素毎、または複数の画素を含むブロック毎の動きベクトルを推定(推論)することが考えられる。この場合、従来手法と比較して、推定精度は向上するものの、画素毎または比較的小さいブロック毎の動きベクトルが推定されるため、動きベクトルが高精細化されるためデータサイズが増大化する。
圧縮された動画のデータには、フレーム毎の動きベクトルの圧縮データが含まれる。そのため、動きベクトルのデータサイズが増大化した場合、動きベクトルの圧縮率が従来と同程度であれば圧縮された動画のデータサイズが増大化する。また、動きベクトルの圧縮率を高くすると、復元された動きベクトルの精度が劣化する。
一方、本開示の技術によれば、動きベクトルを適切に圧縮できる。そのため、例えば、動画の品質を保ちながら動画の圧縮率を向上させることができる。また、従来と同等の動画圧縮率としながら、動画の品質を向上させることもできる。
<変形例>
情報処理装置10、及び情報処理装置20のそれぞれは、一つの筐体に含まれる装置でもよいが、本開示の情報処理装置10、及び情報処理装置20のそれぞれはこれに限定されない。情報処理装置10、及び情報処理装置20のそれぞれの各部は、例えば1以上のコンピュータにより構成されるクラウドコンピューティングにより実現されていてもよい。また、情報処理装置10の少なくとも一部の処理は、例えば、情報処理装置20により実行されてもよい。また、情報処理装置20の少なくとも一部の処理は、例えば、情報処理装置10により実行されてもよい。また、情報処理装置10、及び情報処理装置20は、一体の装置でもよい。これらのような情報処理装置10、及び情報処理装置20についても、本開示の「情報処理装置」の一例に含まれる。
情報処理装置10、及び情報処理装置20のそれぞれは、一つの筐体に含まれる装置でもよいが、本開示の情報処理装置10、及び情報処理装置20のそれぞれはこれに限定されない。情報処理装置10、及び情報処理装置20のそれぞれの各部は、例えば1以上のコンピュータにより構成されるクラウドコンピューティングにより実現されていてもよい。また、情報処理装置10の少なくとも一部の処理は、例えば、情報処理装置20により実行されてもよい。また、情報処理装置20の少なくとも一部の処理は、例えば、情報処理装置10により実行されてもよい。また、情報処理装置10、及び情報処理装置20は、一体の装置でもよい。これらのような情報処理装置10、及び情報処理装置20についても、本開示の「情報処理装置」の一例に含まれる。
以上、実施の形態を参照して本開示を説明したが、本開示は上述の実施の形態に限定されるものではない。本開示の構成や詳細には、本開示のスコープ内で当業者が理解し得る様々な変更をすることができる。そして、各実施の形態は、適宜他の実施の形態と組み合わせることができる。
上記の実施形態の一部又は全部は、以下の付記のようにも記載されうるが、以下には限られない。なお、付記1に従属する各付記に記載した要素(例えば構成及び機能)の一部または全ては、他のカテゴリーの独立する付記に対しても同様の従属関係により従属し得る。任意の付記に記載された要素の一部または全ては、様々なハードウェア、ソフトウェア、ソフトウェアを記録するための記録手段、システム、及び方法に適用され得る。
(付記1)
動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出する算出部と、
前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する圧縮部と、
を有する情報処理装置。
(付記2)
前記圧縮部は、前記動きベクトルが指す位置の画素に対する前記変化量、及び前記動きベクトルが指す位置の周辺の複数の画素のそれぞれに対する前記変化量の少なくとも一方に応じた符号量で前記動きベクトルを圧縮する、
付記1に記載の情報処理装置。
(付記3)
前記算出部は、被写体が小さく写されているほど、被写体にピントが合っているほど、及び被写体のブレが少ないほどの少なくとも一つの条件に従って、前記変化量の値をより大きな値に決定する畳み込みニューラルネットワークを用いて、各画素に対する前記変化量を算出する、
付記1または2に記載の情報処理装置。
(付記4)
前記算出部は、周辺の画素に対して被写体の色の変化が大きいほど、及び被写体の凹凸が大きいほどの少なくとも一方に従って、前記変化量の値をより大きな値に決定する畳み込みニューラルネットワークを用いて、各画素に対する前記変化量を算出する、
付記1または2に記載の情報処理装置。
(付記5)
前記算出部は、ラプラシアンフィルタを用いて、各画素に対する前記変化量を算出する、
付記1または2に記載の情報処理装置。
(付記6)
前記圧縮部は、レート歪み最適化(RDO、Rate-distortion optimization)を用いて、動きベクトルの符号量及び残差の符号量を決定する、
付記1または2に記載の情報処理装置。
(付記7)
動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出し、
前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する、
情報処理方法。
(付記8)
動画における第1フレームと、前記第1フレームよりも以前の時点のフレームに基づく第2フレームとに基づいて推定され、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記第2フレームの各画素の周辺画素に対する変化量に応じた符号量で圧縮された前記動きベクトルと、を取得する取得部と、
前記取得部により取得された、圧縮された前記動きベクトルを、前記変化量に基づいて復元する復元部と、
を有する情報処理装置。
(付記9)
動画における第1フレームと、前記第1フレームよりも以前の時点のフレームに基づく第2フレームとに基づいて推定され、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記第2フレームの各画素の周辺画素に対する変化量に応じた符号量で圧縮された前記動きベクトルと、を取得し、
圧縮された前記動きベクトルを、前記変化量に基づいて復元する、
情報処理方法。
(付記10)
第1情報処理装置と第2情報処理装置とを有し、
前記第1情報処理装置は、
動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出する算出部と、
前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する圧縮部と、
を有し、
前記第2情報処理装置は、
前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で圧縮された前記動きベクトルと、を取得する取得部と、
前記取得部により取得された、圧縮された前記動きベクトルを、前記変化量に基づいて復元する復元部と、
を有する情報処理システム。
(付記11)
動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出し、
前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する、
処理をコンピュータに実行させるプログラム。
(付記12)
動画における第1フレームと、前記第1フレームよりも以前の時点のフレームに基づく第2フレームとに基づいて推定され、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記第2フレームの各画素の周辺画素に対する変化量に応じた符号量で圧縮された前記動きベクトルと、を取得し、
圧縮された前記動きベクトルを、前記変化量に基づいて復元する、
処理をコンピュータに実行させるプログラム。
(付記1)
動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出する算出部と、
前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する圧縮部と、
を有する情報処理装置。
(付記2)
前記圧縮部は、前記動きベクトルが指す位置の画素に対する前記変化量、及び前記動きベクトルが指す位置の周辺の複数の画素のそれぞれに対する前記変化量の少なくとも一方に応じた符号量で前記動きベクトルを圧縮する、
付記1に記載の情報処理装置。
(付記3)
前記算出部は、被写体が小さく写されているほど、被写体にピントが合っているほど、及び被写体のブレが少ないほどの少なくとも一つの条件に従って、前記変化量の値をより大きな値に決定する畳み込みニューラルネットワークを用いて、各画素に対する前記変化量を算出する、
付記1または2に記載の情報処理装置。
(付記4)
前記算出部は、周辺の画素に対して被写体の色の変化が大きいほど、及び被写体の凹凸が大きいほどの少なくとも一方に従って、前記変化量の値をより大きな値に決定する畳み込みニューラルネットワークを用いて、各画素に対する前記変化量を算出する、
付記1または2に記載の情報処理装置。
(付記5)
前記算出部は、ラプラシアンフィルタを用いて、各画素に対する前記変化量を算出する、
付記1または2に記載の情報処理装置。
(付記6)
前記圧縮部は、レート歪み最適化(RDO、Rate-distortion optimization)を用いて、動きベクトルの符号量及び残差の符号量を決定する、
付記1または2に記載の情報処理装置。
(付記7)
動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出し、
前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する、
情報処理方法。
(付記8)
動画における第1フレームと、前記第1フレームよりも以前の時点のフレームに基づく第2フレームとに基づいて推定され、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記第2フレームの各画素の周辺画素に対する変化量に応じた符号量で圧縮された前記動きベクトルと、を取得する取得部と、
前記取得部により取得された、圧縮された前記動きベクトルを、前記変化量に基づいて復元する復元部と、
を有する情報処理装置。
(付記9)
動画における第1フレームと、前記第1フレームよりも以前の時点のフレームに基づく第2フレームとに基づいて推定され、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記第2フレームの各画素の周辺画素に対する変化量に応じた符号量で圧縮された前記動きベクトルと、を取得し、
圧縮された前記動きベクトルを、前記変化量に基づいて復元する、
情報処理方法。
(付記10)
第1情報処理装置と第2情報処理装置とを有し、
前記第1情報処理装置は、
動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出する算出部と、
前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する圧縮部と、
を有し、
前記第2情報処理装置は、
前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で圧縮された前記動きベクトルと、を取得する取得部と、
前記取得部により取得された、圧縮された前記動きベクトルを、前記変化量に基づいて復元する復元部と、
を有する情報処理システム。
(付記11)
動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出し、
前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する、
処理をコンピュータに実行させるプログラム。
(付記12)
動画における第1フレームと、前記第1フレームよりも以前の時点のフレームに基づく第2フレームとに基づいて推定され、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記第2フレームの各画素の周辺画素に対する変化量に応じた符号量で圧縮された前記動きベクトルと、を取得し、
圧縮された前記動きベクトルを、前記変化量に基づいて復元する、
処理をコンピュータに実行させるプログラム。
この出願は、2024年7月10日に出願された日本出願特願2024-110850を基礎とする優先権を主張し、その開示の全てをここに取り込む。
1 情報処理システム
10 情報処理装置
11 算出部
12 圧縮部
20 情報処理装置
21 取得部
22 復元部
10 情報処理装置
11 算出部
12 圧縮部
20 情報処理装置
21 取得部
22 復元部
Claims (20)
- 動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出する算出部と、
前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する圧縮部と、
を有する情報処理装置。 - 前記圧縮部は、前記動きベクトルが指す位置の画素に対する前記変化量、及び前記動きベクトルが指す位置の周辺の複数の画素のそれぞれに対する前記変化量の少なくとも一方に応じた符号量で前記動きベクトルを圧縮する、
請求項1に記載の情報処理装置。 - 前記算出部は、被写体が小さく写されているほど、被写体にピントが合っているほど、及び被写体のブレが少ないほどの少なくとも一つの条件に従って、前記変化量の値をより大きな値に決定する畳み込みニューラルネットワークを用いて、各画素に対する前記変化量を算出する、
請求項1または2に記載の情報処理装置。 - 前記算出部は、周辺の画素に対して被写体の色の変化が大きいほど、及び被写体の凹凸が大きいほどの少なくとも一方に従って、前記変化量の値をより大きな値に決定する畳み込みニューラルネットワークを用いて、各画素に対する前記変化量を算出する、
請求項1または2に記載の情報処理装置。 - 前記算出部は、ラプラシアンフィルタを用いて、各画素に対する前記変化量を算出する、
請求項1または2に記載の情報処理装置。 - 前記圧縮部は、レート歪み最適化(RDO、Rate-distortion optimization)を用いて、動きベクトルの符号量及び残差の符号量を決定する、
請求項1または2に記載の情報処理装置。 - 動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出し、
前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する、
情報処理方法。 - 前記動きベクトルが指す位置の画素に対する前記変化量、及び前記動きベクトルが指す位置の周辺の複数の画素のそれぞれに対する前記変化量の少なくとも一方に応じた符号量で前記動きベクトルを圧縮する、
請求項7に記載の情報処理方法。 - 被写体が小さく写されているほど、被写体にピントが合っているほど、及び被写体のブレが少ないほどの少なくとも一つの条件に従って、前記変化量の値をより大きな値に決定する畳み込みニューラルネットワークを用いて、各画素に対する前記変化量を算出する、
請求項7または8に記載の情報処理方法。 - 周辺の画素に対して被写体の色の変化が大きいほど、及び被写体の凹凸が大きいほどの少なくとも一方に従って、前記変化量の値をより大きな値に決定する畳み込みニューラルネットワークを用いて、各画素に対する前記変化量を算出する、
請求項7または8に記載の情報処理方法。 - ラプラシアンフィルタを用いて、各画素に対する前記変化量を算出する、
請求項7または8に記載の情報処理方法。 - レート歪み最適化(RDO、Rate-distortion optimization)を用いて、動きベクトルの符号量及び残差の符号量を決定する、
請求項7または8に記載の情報処理方法。 - 動画における第1フレームと、前記第1フレームよりも以前の時点のフレームに基づく第2フレームとに基づいて推定され、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記第2フレームの各画素の周辺画素に対する変化量に応じた符号量で圧縮された前記動きベクトルと、を取得する取得部と、
前記取得部により取得された、圧縮された前記動きベクトルを、前記変化量に基づいて復元する復元部と、
を有する情報処理装置。 - 動画における第1フレームと、前記第1フレームよりも以前の時点のフレームに基づく第2フレームとに基づいて推定され、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記第2フレームの各画素の周辺画素に対する変化量に応じた符号量で圧縮された前記動きベクトルと、を取得し、
圧縮された前記動きベクトルを、前記変化量に基づいて復元する、
情報処理方法。 - 第1情報処理装置と第2情報処理装置とを有し、
前記第1情報処理装置は、
動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出する算出部と、
前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する圧縮部と、
を有し、
前記第2情報処理装置は、
前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で圧縮された前記動きベクトルと、を取得する取得部と、
前記取得部により取得された、圧縮された前記動きベクトルを、前記変化量に基づいて復元する復元部と、
を有する情報処理システム。 - 前記圧縮部は、前記動きベクトルが指す位置の画素に対する前記変化量、及び前記動きベクトルが指す位置の周辺の複数の画素のそれぞれに対する前記変化量の少なくとも一方に応じた符号量で前記動きベクトルを圧縮する、
請求項15に記載の情報処理システム。 - 前記算出部は、被写体が小さく写されているほど、被写体にピントが合っているほど、及び被写体のブレが少ないほどの少なくとも一つの条件に従って、前記変化量の値をより大きな値に決定する畳み込みニューラルネットワークを用いて、各画素に対する前記変化量を算出する、
請求項15または16に記載の情報処理システム。 - 前記算出部は、周辺の画素に対して被写体の色の変化が大きいほど、及び被写体の凹凸が大きいほどの少なくとも一方に従って、前記変化量の値をより大きな値に決定する畳み込みニューラルネットワークを用いて、各画素に対する前記変化量を算出する、
請求項15または16に記載の情報処理システム。 - 動画において第1フレームよりも以前の時点のフレームに基づく第2フレームの各画素の周辺画素に対する変化量を算出し、
前記第1フレームと前記第2フレームとに基づいて推定された、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記変化量に応じた符号量で、前記動きベクトルを圧縮する、
処理をコンピュータに実行させるプログラム。 - 動画における第1フレームと、前記第1フレームよりも以前の時点のフレームに基づく第2フレームとに基づいて推定され、前記第2フレームの1以上の各画素に応じた動きベクトルが指す位置での前記第2フレームの各画素の周辺画素に対する変化量に応じた符号量で圧縮された前記動きベクトルと、を取得し、
圧縮された前記動きベクトルを、前記変化量に基づいて復元する、
処理をコンピュータに実行させるプログラム。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2024110850 | 2024-07-10 | ||
| JP2024-110850 | 2024-07-10 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2026014266A1 true WO2026014266A1 (ja) | 2026-01-15 |
Family
ID=98386694
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2025/023179 Pending WO2026014266A1 (ja) | 2024-07-10 | 2025-06-27 | 情報処理装置、情報処理方法、情報処理システム及びプログラム |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2026014266A1 (ja) |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2007096540A (ja) * | 2005-09-27 | 2007-04-12 | Sanyo Electric Co Ltd | 符号化方法 |
| JP2013098745A (ja) * | 2011-10-31 | 2013-05-20 | Fujitsu Ltd | 動画像復号装置、動画像符号化装置、動画像復号方法、動画像符号化方法、動画像復号プログラム及び動画像符号化プログラム |
-
2025
- 2025-06-27 WO PCT/JP2025/023179 patent/WO2026014266A1/ja active Pending
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2007096540A (ja) * | 2005-09-27 | 2007-04-12 | Sanyo Electric Co Ltd | 符号化方法 |
| JP2013098745A (ja) * | 2011-10-31 | 2013-05-20 | Fujitsu Ltd | 動画像復号装置、動画像符号化装置、動画像復号方法、動画像符号化方法、動画像復号プログラム及び動画像符号化プログラム |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12184893B2 (en) | Learned B-frame coding using P-frame coding system | |
| JP2023517846A (ja) | 再帰ベースの機械学習システムを使用したビデオ圧縮 | |
| JP2019198092A (ja) | 画像予測方法および関連装置 | |
| US12217388B2 (en) | Image processing method and device, electronic equipment, and storage medium | |
| JP5522174B2 (ja) | 動画像符号化装置 | |
| CN109688413B (zh) | 以支持辅助帧的视频编码格式编码视频流的方法和编码器 | |
| US20190373284A1 (en) | Coding resolution control method and terminal | |
| US7956898B2 (en) | Digital image stabilization method | |
| US20200374564A1 (en) | Method and apparatus for noise reduction in video systems | |
| CN113658073A (zh) | 图像去噪处理方法、装置、存储介质与电子设备 | |
| CN115299048B (zh) | 图像编码、解码方法及装置、编解码器 | |
| CN112203086B (zh) | 图像处理方法、装置、终端和存储介质 | |
| JPWO2012160626A1 (ja) | 画像圧縮装置、画像復元装置、及びプログラム | |
| CN115955565A (zh) | 处理方法、处理设备及存储介质 | |
| US9635385B2 (en) | Methods and systems for estimating motion in multimedia pictures | |
| US10051270B2 (en) | Video encoding method using at least two encoding methods, device and computer program | |
| CN113643209B (zh) | 图像降噪处理方法、装置、存储介质与电子设备 | |
| JP2014003587A (ja) | 画像符号化装置及びその方法 | |
| CN112672150A (zh) | 基于视频预测的视频编码方法 | |
| JP2013157950A (ja) | 符号化方法、復号方法、符号化装置、復号装置、符号化プログラム及び復号プログラム | |
| US20250234012A1 (en) | Image encoding apparatus, image encoding method and non-transitory computer-readable storage medium, image decoding apparatus, image decoding method and non-transitory computer-readable storage medium | |
| CN111970517B (zh) | 基于双向光流的帧间预测方法、编码方法及相关装置 | |
| CN119854504A (zh) | 视频帧编码方法、装置 | |
| CN117061753A (zh) | 预测帧间编码的运动矢量的方法和装置 | |
| CN117119188A (zh) | 图像编码方法、装置、介质和计算设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25836827 Country of ref document: EP Kind code of ref document: A1 |