WO2020003936A1 - フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラム - Google Patents
フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラム Download PDFInfo
- Publication number
- WO2020003936A1 WO2020003936A1 PCT/JP2019/022307 JP2019022307W WO2020003936A1 WO 2020003936 A1 WO2020003936 A1 WO 2020003936A1 JP 2019022307 W JP2019022307 W JP 2019022307W WO 2020003936 A1 WO2020003936 A1 WO 2020003936A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- divergence
- filter
- frame
- moving image
- degree
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/117—Filters, e.g. for pre-processing or post-processing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/136—Incoming video signal characteristics or properties
- H04N19/137—Motion inside a coding unit, e.g. average field, frame or block difference
- H04N19/139—Analysis of motion vectors, e.g. their magnitude, direction, variance or reliability
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/146—Data rate or code amount at the encoder output
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/172—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a picture, frame or field
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/80—Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/85—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using pre-processing or post-processing specially adapted for video compression
Definitions
- the present invention relates to a filter selection method, a filter selection device, and a filter selection program.
- High image quality during image reproduction is intended to express the smooth movement of the subject by approaching the upper limit of the frame rate that can be detected by the visual system (displayable on the display). Therefore, high image quality during image reproduction is based on the premise that the display device reproduces moving images at a constant speed.
- the purpose of improving the accuracy of image analysis is to improve the accuracy of image analysis by using a high frame rate image that exceeds the detection limit of vision.
- Image analysis by slow reproduction of a high-speed moving object such as an athlete, FA / inspection, or car is a typical application example.
- the upper limit of the frame rate of the moving image input system and the upper limit of the frame rate of the moving image output system are asymmetric. That is, the upper limit of the frame rate of the high-speed camera as the moving image input system exceeds 10,000 fps.
- the upper limit of the frame rate of the display device, which is a moving image output system is from 120 fps to 240 fps. For this reason, a moving image captured by a high-speed camera is used for slow reproduction (see Patent Document 1).
- the high frame rate image includes a group of frames sampled at high density in the time direction. If the image generation apparatus generates an image for constant speed reproduction at 30 Hz or the like by using a frame group sampled at a high density such as 1000 Hz, the generation of the image for constant speed reproduction can be controlled with high time resolution. It is possible.
- the image generation device samples a frame at a reproduction frame rate. For this reason, the conventional image generation device does not sample a frame at a time resolution higher than the reproduction frame rate.
- the degree of freedom in filter design is expanded.
- a moving image signal of 62.5 fps is generated from a moving image signal of 1000 fps by filtering a moving image signal of 1000 fps
- the frame to be filtered does not overlap even if the frames to be filtered do not overlap.
- Can be 16 ( 1000 / 62.5) frames, which is more than 2 frames.
- the encoder may be able to improve the coding efficiency.
- the conventional filter selection device reduces the generated code amount of the low frame rate image generated from the high frame rate image by the temporal filter in the pre-processing of the moving image In some cases, it is not possible to select a coefficient of a temporal filter for improving the encoding efficiency of a low frame rate image so that the encoder keeps the objective image quality of the image at or above a certain level.
- the present invention provides an encoder that reduces the amount of generated code of a low frame rate image generated from a high frame rate image by a temporal filter, and further reduces the objective image quality of the low frame rate image. It is an object of the present invention to provide a filter selection method, a filter selection device, and a filter selection program capable of selecting a coefficient of a temporal filter for improving coding efficiency of a low frame rate image so as to keep the coding efficiency at or above a certain value.
- One aspect of the present invention is a filter selection method executed by a filter selection device, wherein an encoding target frame of a low temporal resolution moving image is generated from a high temporal resolution moving image frame according to a filter coefficient. Step, the degree of divergence between the frame of the moving image of the high time resolution and the encoding target frame of the moving image of the low time resolution is obtained for each encoding target frame, and the generated code amount of the encoding target frame Selecting the filter coefficient for minimizing the weighted sum with the divergence from a set of candidates for the filter coefficient such that the sum of the divergence falls within a predetermined range. It is.
- One aspect of the present invention is the above-described filter selection method, further comprising the step of updating the divergence degree, wherein the selecting step repeats the sum of the divergence degrees based on the updated divergence degree.
- the updating step the calculation result of the component representing the degree of divergence up to the previous time is recorded in a storage unit, and the calculation result up to the previous time is obtained from the storage unit, and the obtained calculation result is obtained up to the previous time.
- the divergence is updated based on the calculation result.
- One embodiment of the present invention is a filter that generates a coding target frame of a low temporal resolution moving image from a frame of a high temporal resolution moving image according to a filter coefficient, and a frame of the high temporal resolution moving image.
- the filter that obtains the degree of divergence from the encoding target frame of the low temporal resolution moving image for each encoding target frame, and minimizes the weighted sum of the generated code amount of the encoding target frame and the degree of divergence.
- a selection unit that selects a coefficient from a set of candidates for the filter coefficient such that the sum of the degrees of divergence falls within a predetermined range.
- One aspect of the present invention is the above-described filter selection device, further comprising an update unit for updating the divergence degree, wherein the selection unit repeats the sum of the divergence degrees based on the updated divergence degree.
- the update unit records the calculation result of the component representing the degree of divergence up to the previous time in a storage unit, acquires the calculation result up to the previous time from the storage unit, and obtains the obtained calculation result up to the previous time.
- the divergence is updated based on the calculation result.
- One embodiment of the present invention is a filter selection program for causing a computer to function as the above-described filter selection device.
- the encoder reduces the amount of generated code of a low frame rate image generated from a high frame rate image by a temporal filter, and furthermore, the encoder keeps the objective image quality of the low frame rate image at a certain level or more.
- FIG. 6 is a flowchart illustrating an operation example of the filter selection device.
- FIG. 1 is a diagram illustrating a configuration example of the filter selection device 1.
- the filter selection device 1 is an information processing device that selects a coefficient of a filter.
- the filter selection device 1 includes a storage unit 10, an acquisition unit 11, a filter 12, an encoder 13, an initialization unit 14, an update unit 15, and a selection unit 16.
- Part or all of the functional units are realized by, for example, a processor such as a CPU (Central Processing Unit) executing a program stored in a storage unit.
- a processor such as a CPU (Central Processing Unit) executing a program stored in a storage unit.
- Some or all of the functional units may be realized using hardware such as an LSI (Large Scale Integration) or an ASIC (Application Specific Integrated Circuit).
- the storage unit 10 is a non-volatile recording medium (non-temporary recording medium) such as a flash memory and an HDD (Hard Disk Drive).
- the storage unit 10 may include a volatile recording medium such as a RAM (Random Access Memory) and a register.
- the storage unit 10 stores, for example, a frame group of a moving image with a high time resolution before filtering processing, a data table, and a program.
- the acquisition unit 11 acquires a high-resolution video frame (high frame rate image) as an original image of the encoding target frame.
- the acquisition unit 11 records a frame group of a moving image having a high time resolution in the storage unit 10.
- Filter 12 is a temporal filter. The filter 12 acquires a frame of a moving image with a high time resolution from the acquiring unit 11.
- the filter 12 performs a filtering process (downsampling) on a frame of a moving image having a high temporal resolution based on the coefficient vector and the shift amount selected by the selecting unit 16.
- the filter 12 outputs the encoding target frame (low frame rate image) of the low temporal resolution moving image generated by the filtering process to the encoder 13 and the selection unit 16.
- the encoder 13 performs lossless encoding using motion compensation prediction (Motion Compensation Prediction) on the encoding target frame of the moving image with low temporal resolution.
- the encoder 13 outputs to the selection unit 16 information indicating the amount of generated codes of the encoding target frame of the moving image with low temporal resolution.
- the initialization unit 14 initializes a shift amount of a time position (downsampling time by filtering) at which a filter coefficient is applied by the filtering process with respect to the filtering process of each frame (stage).
- the initialization unit 14 initializes an index k representing the number of repetitions to 0.
- the initialization unit 14 determines a value of a coefficient ⁇ to be multiplied by a predetermined value, which is multiplied by a degree of divergence between a frame of a moving image having a high temporal resolution before the filtering process and an encoding target frame of a moving image having a low temporal resolution after the filtering process. Initialize to a value.
- the update unit 15 acquires the threshold Dth .
- the updating unit 15 acquires the sum Dk of the degrees of deviation from the selecting unit 16.
- the updating unit 15 acquires, from the initializing unit 14, an index k representing the number of repetitions and a value of a coefficient ⁇ to be multiplied by the degree of deviation ⁇ . Updating unit 15 minimizes the amount of code generated for the encoding target frame, as objective image quality of the encoding target frame is kept above a certain, based on a result of comparison between the sum of the discrepancy D k and the threshold D th Then, the coefficient ⁇ that is multiplied by the degree of deviation ⁇ is updated.
- the updating unit 15 outputs the value of the coefficient ⁇ multiplied by the degree of deviation ⁇ to the selecting unit 16.
- the selection unit 16 acquires, from the storage unit 10, a frame group of a moving image having a high time resolution before the filtering process.
- the selecting unit 16 acquires, from the filter 12, a coding target frame of a moving image with a low temporal resolution after the filtering process.
- the selecting unit 16 acquires, from the encoder 13, information indicating the generated code amount of the encoding target frame of the moving image with low time resolution.
- the selection unit 16 calculates, for each encoding target frame, the degree of divergence between the high-resolution video frame before the filtering process and the encoding target frame of the low-resolution moving image after the filtering process.
- the selecting unit 16 calculates a weighted sum of the generated code amount of the encoding target frame of the moving image with low temporal resolution and the degree of deviation.
- the selected coefficient vector candidate is referred to as a “coefficient candidate vector”.
- the selection unit 16 selects a coefficient vector (filter coefficient) that minimizes the weighted sum from a set of coefficient candidate vectors based on the weighted sum of the generated code amount and the degree of deviation. That is, based on the weighted sum of the generated code amount and the divergence, the selection unit 16 selects an optimal coefficient vector from the set of coefficient candidate vectors to generate an encoding target frame with a low time resolution. .
- the selecting unit 16 may select the shift amount of the time position (down sampling time by filtering) at which the filter coefficient is applied, based on the weighted sum of the generated code amount and the divergence.
- the encoding target frame of the low temporal resolution moving image suitable for encoding is obtained by using the temporal filter based on the coefficient vector and the shift amount selected by the selecting unit 16 to obtain the high temporal resolution moving image frame. Generated from a group.
- the selection unit 16 calculates the sum Dk of the divergence degrees for a plurality of frames. When the degree of deviation is updated, the selecting unit 16 updates the sum Dk of the degrees of deviation. The selection unit 16 outputs the divergence sum Dk to the update unit 15.
- each frame of the moving image is represented as a one-dimensional signal.
- the i-th frame generated based on the (2 ⁇ + 1) -tap temporal filter is represented by Expression (1).
- i an index designating a frame after downsampling.
- the index takes a non-negative integer value.
- [delta] t represents the distance between the frame of a moving image to be input to the filter 12.
- w i [j] represents a filter coefficient for a reference frame used for motion compensation prediction.
- w i [j] satisfies the relationship of equation (2).
- the coefficient vector the formula in the text are capitalized as W i.
- p i represents a parameter for correcting the time position at which the filter coefficient is applied.
- M represents a parameter (downsampling ratio) that determines the frame rate of the encoding target frame of the moving image output from the filter 12.
- the frame rate of the encoding target frame of the moving image output from the filter 12 is 1 / (M ⁇ t ).
- the selecting unit 16 selects a coefficient vector optimal for generating a frame to be coded for a moving image with a low temporal resolution from among (N ⁇ P) N types of coefficient candidate vectors, Select for each frame.
- a set of coefficient candidate vectors is referred to as a “dictionary”.
- the selecting unit 16 uses the generated code amount in the frame output from the filter 12 (the frame generated by the time filter) as a criterion for optimization in designing the time filter.
- Information indicating the generated code amount is obtained from the encoder 13 that executes lossless encoding using motion compensation prediction.
- R h represents a code amount of the header information coder 13 is produced.
- R d (d i [0] , ..., d i [K-1]) is the estimated displacement amount d i [0], ..., representing the code amount related to d i [K-1].
- R e (e i (0, W i, W i-1, p i, p i-1), ..., e i (X-1, W i, W i-1, p i, p i-1) ) Indicates the code amount related to the motion compensation inter-frame prediction error.
- the generated code amount ⁇ ⁇ ⁇ ⁇ in Expression (5) is a displacement amount, header information, a coefficient vector W i for the i-th frame, and a correction parameter of a time position at which a filter coefficient is applied to the i-th frame.
- the shift amount p i , the coefficient vector W i-1 for the (i-1) th frame, and the shift amount p i-1 which is a correction parameter of a time position at which a filter coefficient is applied to the (i-1) th frame Is determined by
- the selecting unit 16 determines the degree of divergence between the sampled high-resolution video frame (original image) and the encoding target frame of the video output from the filter 12 (the video frame generated by filtering).
- the value of Expression (6) is calculated as ⁇ .
- the selection unit 16 calculates the value of Expression (7) as a weighted sum of the generated code amount shown in Expression (5) and the degree of divergence shown in Expression (6).
- the value of Expression (7) represents an evaluation scale of the filter design of the filter 12.
- ⁇ is a coefficient by which the degree of deviation ⁇ [W i , p i ] is multiplied as a coefficient of the weighted sum.
- the value of ⁇ is initialized by the initialization unit 14.
- the value of ⁇ is updated by the updating unit 15.
- the selection unit 16 acquires the initialized or updated value of ⁇ .
- the selection unit 16 designs the filter 12 such that the evaluation scale shown in Expression (7) is minimized.
- the filter 12 based on the generated code amount ⁇ [W i , W i ⁇ 1 , p i , p i ⁇ 1 ]
- an effect of reducing the generated code amount of a moving image generated by filtering is expected. it can.
- the filter 12 based on the degree of divergence ⁇ [W i , p i ] an effect of suppressing the divergence between the moving image generated by the filtering and the sampled moving image can be expected.
- the selecting unit 16 solves a minimization problem in which the generated code amount in Expression (5) is used as a cost function so that the filter 12 generates a frame that minimizes the generated code amount.
- the selector 16 calculates (J / M) sets of coefficient vectors and the shift amount that satisfy the expression (8). J represents the number of frames of a moving image with a high temporal resolution.
- equation (8) gives: It can be formulated as an optimization problem in a simple Markov process.
- an optimal solution can be calculated by a dynamic programming method with an operation amount of a polynomial order.
- a solution method using the dynamic programming method will be described.
- S i-1 (W i-1 , p i-1 ) has already been calculated using the same recurrence formula.
- S i ⁇ 1 (W i ⁇ 1 , p i ⁇ 1 ) is stored in a storage unit such as a register as a referenceable value when S i (W i , p i ) is calculated.
- the selection unit 16 determines that [[W i , W i ⁇ 1 , p i , p i ⁇ 1 ] + S i ⁇ 1 (W i ⁇ 1 , p i ⁇ 1 ) S i (W i , p i ) is calculated by selecting from the dictionary ⁇ ⁇ N a coefficient candidate vector and a shift amount p i that minimize S i .
- Selection unit 16 records the index of the coefficient candidate vectors formula (10) becomes a minimum value ⁇ n i-1 (n i , p i) of the storage unit 10 for each n i.
- the selection unit 16 records the shift amount ⁇ p i-1 (n i , p i ) at which the equation (10) becomes the minimum value in the storage unit 10 for each n i . Accordingly, the selection unit 16 determines the index ⁇ n i-1 (n i , p i ) of the coefficient candidate vector that minimizes the expression (10) and the shift amount ⁇ p i-1 (n i , p i ). Can be obtained from the storage unit 10 in the subsequent processing.
- the selection unit 16 can calculate the equation (11) with the operation amount of the polynomial order.
- the selecting unit 16 selects the optimal solution (W * 0 ,..., W * J / M-1 , p * 0 ,..., P * J / M-1 ) are calculated by the backtrack process shown below.
- the index of the coefficient candidate vector representing W * J / M-1 is represented as nJ / M-1 .
- the index of the coefficient candidate vector for the (J / M-1) th frame is represented as nJ / M-1 .
- the shift amount for the (J / M-1) th frame is expressed as pJ / M-1 .
- the index of the optimal coefficient candidate vector for the (J / M-2) th frame is stored in the storage unit 10 as ⁇ n J / M-2 (n J / M ⁇ 1 , p J / M ⁇ 1 ). I have.
- the shift amount for the (J / M-2) th frame is stored in the storage unit 10 as ⁇ p J / M-2 (n J / M ⁇ 1 , p J / M ⁇ 1 ).
- the selection unit 16 selects a combination of a coefficient vector and a shift amount.
- the update unit 15 records the calculation result in the storage unit 10 so as to refer to the calculation result as needed.
- ⁇ M-1 k 0 ⁇ x f (x, (iM + p i + j) ⁇ t) and a value of 2
- the updating unit 15 stores the calculation result of the component representing the divergence up to the previous time as needed. Recorded in the unit 10 in advance. The updating unit 15 acquires the calculation results of such components up to the previous time from the storage unit 10 as necessary. The updating unit 15 updates the degree of divergence based on the obtained calculation result up to the previous time.
- FIG. 2 is a flowchart illustrating an operation example of the filter selection device 1.
- the initialization unit 14 initializes an index k representing the number of repetitions to 0 (Step S101).
- the initialization unit 14 initializes parameters. For example, the initialization unit 14 initializes the value of ⁇ k to the value of ⁇ 0 which is a predetermined initial value (Step S102).
- the selection unit 16 performs the selection processing described in the above description of the “optimization of filter coefficients”.
- the selecting unit 16 selects a coefficient vector and a shift amount in each stage using the coefficient ⁇ k (Step S103).
- the update unit 15 determines whether or not the sum Dk of the divergence degrees is larger than the threshold value Dth (Step S105). When the sum D k of the divergence degrees is larger than the threshold value D th (step S105: YES), the updating unit 15 determines whether ⁇ k ⁇ 1 is smaller than ⁇ k (step S106). When ⁇ k ⁇ 1 is smaller than ⁇ k (step S106: YES), the updating unit 15 sets the value of ⁇ k + 1 to 2 ⁇ k (step S107). When ⁇ k ⁇ 1 is equal to or larger than ⁇ k (step S106: NO), the updating unit 15 sets the value of ⁇ k + 1 to ⁇ k / 2 (step S108).
- step S105 determines that ⁇ k ⁇ 1 Is greater than or equal to ⁇ k (step S109).
- ⁇ k ⁇ 1 is larger than ⁇ k (step S109: YES)
- the updating unit 15 sets the value of ⁇ k + 1 to 2 ⁇ k .
- step S109: NO the updating unit 15 sets the value of ⁇ k + 1 to ⁇ k / 2 (step S111).
- the updating unit 15 increments the value of the index k by one (Step S112).
- the selecting unit 16 determines whether the convergence condition is satisfied (Step S113).
- the convergence condition is, for example, a condition that D th ⁇ ⁇ D k ⁇ D th is satisfied.
- the convergence condition may be, for example, a condition that the number of iterations reaches a specified threshold value.
- the selecting unit 16 returns the process to step S103.
- the selection unit 16 ends the processing illustrated in FIG. In this way, the selection unit 16 selects a combination of the coefficient vector and the shift amount recorded in the dictionary as parameters of the time filter.
- the filter selection device 1 of the embodiment includes the filter 12 and the selection unit 16.
- the filter 12 generates an encoding target frame of a low temporal resolution moving image from a high temporal resolution moving image frame according to a filter coefficient.
- the selecting unit 16 obtains the degree of divergence between the frame of the moving image with the high temporal resolution and the encoding target frame of the moving image with the low temporal resolution for each encoding target frame.
- the selecting unit 16 sets the filter coefficient for minimizing the weighted sum of the generated code amount and the divergence of the encoding target frame within a predetermined constraint range (D th - ⁇ ⁇ D k ⁇ D th ). Are selected from a set (dictionary) of filter coefficient candidates so that the sum D k of the filter coefficients can be accommodated.
- the filter selecting apparatus 1 of the embodiment can reduce the amount of generated code of the low frame rate image generated from the high frame rate image by the temporal filter, and then encode the objective image quality of the low frame rate image. It is possible to select the coefficients of a temporal filter that improves the coding efficiency of low frame rate images so that the transformer keeps it above a certain level.
- the filter selection device 1 according to the embodiment can maintain the subjective image quality of the moving image after the filtering.
- the filter selection device 1 according to the embodiment can reduce the generated code amount of a moving image after filtering.
- the filter selection device 1 of the embodiment is capable of selecting a coefficient vector that minimizes the cumulative value of the weighted sum of frames of a moving image.
- the filter selecting device 1 of the embodiment may further include the updating unit 15.
- the updating unit 15 may acquire the predicted distribution of the generated code amount corresponding to another filter coefficient based on the generated code amount corresponding to the filter coefficient.
- the updating unit 15 may acquire a predicted value of the generated code amount according to another filter coefficient based on the predicted distribution.
- the updating unit 15 may generate a set of filter coefficient candidates based on the generated code amount according to another filter coefficient whose predicted value of the generated code amount is minimized.
- the filter selection device 1 of the embodiment can optimize the dictionary design processing and the selection processing in the temporal filtering involving the dynamic update of the dictionary of the coefficient vector.
- the filter selection device in the above-described embodiment may be realized by a computer.
- a program for realizing this function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read and executed by a computer system.
- the “computer system” includes an OS and hardware such as peripheral devices.
- the “computer-readable recording medium” refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, and a CD-ROM, and a storage device such as a hard disk built in a computer system.
- a “computer-readable recording medium” refers to a communication line for transmitting a program via a network such as the Internet or a communication line such as a telephone line, which dynamically holds the program for a short time.
- a program may include a program that holds a program for a certain period of time, such as a volatile memory in a computer system serving as a server or a client in that case.
- the program may be for realizing a part of the functions described above, or may be a program that can realize the functions described above in combination with a program already recorded in a computer system, It may be realized by using a programmable logic device such as an FPGA (Field Programmable Gate Array).
- FPGA Field Programmable Gate Array
- the present invention is applicable to a moving picture coding apparatus.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
フィルタ選択方法は、低い時間解像度の動画像の符号化対象フレームを、高い時間解像度の動画像のフレームから、フィルタ係数に応じて生成するステップと、高い時間解像度の動画像のフレームと低い時間解像度の動画像の符号化対象フレームとの乖離度を符号化対象フレームごとに取得し、符号化対象フレームの発生符号量と乖離度との加重和を最小化するフィルタ係数を、予め定められた範囲内に乖離度の和が収まるように、フィルタ係数の候補の集合から選択するステップとを含む。
Description
本発明は、フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラムに関する。
昨今の半導体技術の進歩を受け、高速度カメラにおける動画像のフレームレートが大きく向上している。高速度カメラにより取得された高フレームレート画像の用途は、画像再生時の高画質化と画像解析の高精度化とに分類される。
画像再生時の高画質化は、視覚系で検知可能(ディスプレイで表示可能)なフレームレートの上限に迫ることにより、被写体の滑らかな動きを表現することが目的である。このため、画像再生時の高画質化は、ディスプレイ装置が動画像を等速再生することが前提である。
一方、画像解析の高精度化は、視覚の検知限を越えた高フレームレート画像を用いることにより、画像解析の高精度化を行うことが目的である。スポーツ選手、FA・検査、自動車等の高速移動物体のスロー再生による画像解析は、代表的な応用例である。
動画像の入力システムのフレームレートの上限と動画像の出力システムのフレームレートの上限とは非対称である。すなわち、動画像の入力システムである高速度カメラのフレームレートの上限は、10000fpsを超えている。一方、動画像の出力システムであるディスプレイ装置のフレームレートの上限は、120fpsから240fpsまでである。このため、高速度カメラで撮影された動画像は、スロー再生に用いられる(特許文献1参照)。
視覚の検知限を越えた高フレームレート画像を用いることにより、動画像の符号化処理に対して親和性の高い等速再生用の画像を生成することができる。高フレームレート画像は、時間方向に高密度でサンプリングされたフレーム群を含んでいる。画像生成装置は、1000Hz等の高密度時間サンプリングされたフレーム群を用いて30Hz等の等速再生用の画像を生成すれば、等速再生用の画像の生成を高い時間分解能で制御することが可能である。
しかしながら、符号発生量の低減を目的とした動画像符号化の前処理では、画像生成装置が再生フレームレートでフレームをサンプリングすることが前提となっている。このため、従来の画像生成装置は、再生フレームレートよりも高い時間分解能ではフレームをサンプリングしていない。
高フレームレート画像のフレームを単純に間引く処理では、時間方向のエイリアシングに起因する画質劣化が問題となる。このような問題を回避するには、時間フィルタによる時間軸方向の帯域制限フィルタリングが必要である。
一方、動き補償フレーム間予測を用いる符号化器では、時間方向のエイリアシングの低減は、予測誤差の低減に直接の関係がない。また、動き補償フレーム間予測を用いる符号化器では、高密度時間サンプリングされたフレームが十分に活用されておらず、時間フィルタとしての自由度には制約がある。
すなわち、30fps又は60fps等の低フレームレート画像の場合、フィルタリングのための十分な数のサンプル(フレーム)が確保できないため、フィルタの特性を高精度に近似することは困難である。例えば、60fpsの動画像信号をフィルタリングすることによって60fpsの動画像信号から30fpsの動画像信号が生成される場合、フィルタリングの対象のフレームが重複しないという条件下では、フィルタリングの対象のフレームは2(=60/30)フレームに限定されるという制約がある。
一方、高フレームレート画像の場合、フィルタ設計の自由度は拡張される。例えば、1000fpsの動画像信号をフィルタリングすることによって、1000fpsの動画像信号から62.5fpsの動画像信号が生成される場合、フィルタリングの対象のフレームが重複しないという条件下でも、フィルタリングの対象のフレームは、2フレームよりも多い16(=1000/62.5)フレームとすることができる。このように、高フレームレート画像から低フレームレート画像を生成する場合、フィルタリング設計の自由度は高い。この自由度の高さを利用することで、符号化器は符号化効率を向上させることができる可能性がある。ここで、符号化器は、符号化処理における低フレームレート画像の発生符号量を低減した上で、低フレームレート画像の客観画質を一定以上に保つ必要がある。
しかしながら、従来のフィルタ選択装置は、動画像符号化の前処理において、時間フィルタによって高フレームレート画像から生成される低フレームレート画像の発生符号量を符号化器が低減した上で、低フレームレート画像の客観画質を符号化器が一定以上に保つように、低フレームレート画像の符号化効率を向上させる時間フィルタの係数を選択することができない場合があった。
上記事情に鑑み、本発明は、時間フィルタによって高フレームレート画像から生成される低フレームレート画像の発生符号量を符号化器が低減した上で、低フレームレート画像の客観画質を符号化器が一定以上に保つように、低フレームレート画像の符号化効率を向上させる時間フィルタの係数を選択することが可能であるフィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラムを提供することを目的としている。
本発明の一態様は、フィルタ選択装置が実行するフィルタ選択方法であって、低い時間解像度の動画像の符号化対象フレームを、高い時間解像度の動画像のフレームから、フィルタ係数に応じて生成するステップと、前記高い時間解像度の動画像のフレームと前記低い時間解像度の動画像の符号化対象フレームとの乖離度を前記符号化対象フレームごとに取得し、前記符号化対象フレームの発生符号量と前記乖離度との加重和を最小化する前記フィルタ係数を、予め定められた範囲内に前記乖離度の和が収まるように、前記フィルタ係数の候補の集合から選択するステップとを含むフィルタ選択方法である。
本発明の一態様は、上記のフィルタ選択方法であって、前記乖離度を更新するステップを更に含み、前記選択するステップでは、更新された前記乖離度に基づいて前記乖離度の和を反復して算出し、前記更新するステップでは、前記乖離度を表す構成要素の前回までの算出結果を記憶部に記録し、前記前回までの算出結果を前記記憶部から取得し、取得された前記前回までの算出結果に基づいて前記乖離度を更新する。
本発明の一態様は、低い時間解像度の動画像の符号化対象フレームを、高い時間解像度の動画像のフレームから、フィルタ係数に応じて生成するフィルタと、前記高い時間解像度の動画像のフレームと前記低い時間解像度の動画像の符号化対象フレームとの乖離度を前記符号化対象フレームごとに取得し、前記符号化対象フレームの発生符号量と前記乖離度との加重和を最小化する前記フィルタ係数を、予め定められた範囲内に前記乖離度の和が収まるように、前記フィルタ係数の候補の集合から選択する選択部とを備えるフィルタ選択装置である。
本発明の一態様は、上記のフィルタ選択装置であって、前記乖離度を更新する更新部を更に備え、前記選択部は、更新された前記乖離度に基づいて前記乖離度の和を反復して算出し、前記更新部は、前記乖離度を表す構成要素の前回までの算出結果を記憶部に記録し、前記前回までの算出結果を前記記憶部から取得し、取得された前記前回までの算出結果に基づいて前記乖離度を更新する。
本発明の一態様は、上記のフィルタ選択装置としてコンピュータを機能させるためのフィルタ選択プログラムである。
本発明により、時間フィルタによって高フレームレート画像から生成される低フレームレート画像の発生符号量を符号化器が低減した上で、低フレームレート画像の客観画質を符号化器が一定以上に保つように、低フレームレート画像の符号化効率を向上させる時間フィルタの係数を選択することが可能である。
本発明の実施形態について、図面を参照して詳細に説明する。
図1は、フィルタ選択装置1の構成例を示す図である。フィルタ選択装置1は、フィルタの係数を選択する情報処理装置である。フィルタ選択装置1は、記憶部10と、取得部11と、フィルタ12と、符号化器13と、初期化部14と、更新部15と、選択部16とを備える。
図1は、フィルタ選択装置1の構成例を示す図である。フィルタ選択装置1は、フィルタの係数を選択する情報処理装置である。フィルタ選択装置1は、記憶部10と、取得部11と、フィルタ12と、符号化器13と、初期化部14と、更新部15と、選択部16とを備える。
各機能部のうち一部又は全部は、例えば、CPU(Central Processing Unit)等のプロセッサが、記憶部に記憶されたプログラムを実行することにより実現される。各機能部のうち一部又は全部は、LSI(Large Scale Integration)やASIC(Application Specific Integrated Circuit)等のハードウェアを用いて実現されてもよい。
記憶部10は、例えばフラッシュメモリ、HDD(Hard Disk Drive)などの不揮発性の記録媒体(非一時的な記録媒体)である。記憶部10は、例えば、RAM(Random Access Memory)やレジスタなどの揮発性の記録媒体を有してもよい。記憶部10は、例えば、フィルタリング処理前の高い時間解像度の動画像のフレーム群と、データテーブルと、プログラムとを記憶する。
取得部11は、高い時間解像度の動画像のフレーム(高フレームレート画像)を、符号化対象フレームの原画像として取得する。取得部11は、高い時間解像度の動画像のフレーム群を、記憶部10に記録する。フィルタ12は、時間フィルタである。フィルタ12は、高い時間解像度の動画像のフレームを、取得部11から取得する。
フィルタ12は、選択部16によって選択された係数ベクトル及びシフト量に基づいて、高い時間解像度の動画像のフレームに対してフィルタリング処理(ダウンサンプリング)を実行する。フィルタ12は、フィルタリング処理によって生成された低い時間解像度の動画像の符号化対象フレーム(低フレームレート画像)を、符号化器13及び選択部16に出力する。
符号化器13は、低い時間解像度の動画像の符号化対象フレームに対して、動き補償予測(Motion Compensation Prediction)を用いた可逆符号化を実行する。符号化器13は、低い時間解像度の動画像の符号化対象フレームの発生符号量を表す情報を、選択部16に出力する。
初期化部14は、各フレーム(ステージ)のフィルタリング処理に関して、フィルタリング処理によるフィルタ係数が施される時間位置(フィルタリングによるダウンサンプリング時刻)のシフト量を初期化する。初期化部14は、反復回数を表すインデックスkを0に初期化する。初期化部14は、フィルタリング処理前の高い時間解像度の動画像のフレームと、フィルタリング処理後の低い時間解像度の動画像の符号化対象フレームとの乖離度に乗算される係数λの値を、所定値に初期化する。
更新部15は、閾値Dthを取得する。更新部15は、乖離度の和Dkを選択部16から取得する。更新部15は、反復回数を表すインデックスkと、乖離度Φに乗算される係数λの値とを、初期化部14から取得する。更新部15は、符号化対象フレームの発生符号量が最小化し、符号化対象フレームの客観画質が一定以上に保たれるように、乖離度の和Dkと閾値Dthとの比較結果に基づいて、乖離度Φに乗算される係数λを更新する。
更新部15は、乖離度Φに乗算される係数λの値を、選択部16に出力する。
更新部15は、乖離度Φに乗算される係数λの値を、選択部16に出力する。
選択部16は、フィルタリング処理前の高い時間解像度の動画像のフレーム群を、記憶部10から取得する。選択部16は、フィルタリング処理後の低い時間解像度の動画像の符号化対象フレームを、フィルタ12から取得する。選択部16は、低い時間解像度の動画像の符号化対象フレームの発生符号量を表す情報を、符号化器13から取得する。
選択部16は、フィルタリング処理前の高い時間解像度の動画像のフレームと、フィルタリング処理後の低い時間解像度の動画像の符号化対象フレームとの乖離度を、符号化対象フレームごとに算出する。選択部16は、低い時間解像度の動画像の符号化対象フレームの発生符号量と乖離度との加重和を算出する。以下、選択される係数ベクトルの候補を「係数候補ベクトル」という。
選択部16は、発生符号量と乖離度との加重和に基づいて、係数候補ベクトルの集合の中から、加重和を最小化する係数ベクトル(フィルタ係数)を選択する。すなわち、選択部16は、発生符号量と乖離度との加重和に基づいて、係数候補ベクトルの集合の中から、低い時間解像度の符号化対象フレームを生成するために最適な係数ベクトルを選択する。選択部16は、発生符号量と乖離度との加重和に基づいて、フィルタ係数が施される時間位置(フィルタリングによるダウンサンプリング時刻)のシフト量を選択してもよい。
このように、符号化に適した低い時間解像度の動画像の符号化対象フレームは、選択部16によって選択された係数ベクトル及びシフト量に基づく時間フィルタを用いて、高い時間解像度の動画像のフレーム群から生成される。
選択部16は、複数のフレームについて乖離度の和Dkを算出する。選択部16は、乖離度が更新された場合、乖離度の和Dkを更新する。選択部16は、乖離度の和Dkを更新部15に出力する。
次に、フィルタ選択装置1の構成例の詳細を説明する。
[表記法]
以下では、表記の簡略化のため、動画像の各フレームは一次元信号として表される。(2Δ+1)タップの時間フィルタに基づいて生成された第iフレームは、式(1)のように表される。
[表記法]
以下では、表記の簡略化のため、動画像の各フレームは一次元信号として表される。(2Δ+1)タップの時間フィルタに基づいて生成された第iフレームは、式(1)のように表される。
iは、ダウンサンプリング後のフレームを指定するインデックスを表す。インデックスは、非負の整数値をとる。δtは、フィルタ12に入力される動画像のフレームの間隔を表す。各フレームは、時刻t=jδt(j=0,1,…)においてサンプリングされる。f(x,t)(x=0,…,X-1)は、第tフレームの空間位置xにおける画素値である。
wi[j]は、動き補償予測に用いられる参照フレームに対するフィルタ係数を表す。wi[j]は、式(2)の関係を満たす。
wi[j]は、動き補償予測に用いられる参照フレームに対するフィルタ係数を表す。wi[j]は、式(2)の関係を満たす。
以下、係数ベクトルは、文章中の数式ではWiのように大文字で表記される。式(1)において、左辺のWiは、フィルタ係数を要素とする係数ベクトルWi=(wi[-Δ],…,wi[Δ])を表す。piは、フィルタ係数が施される時間位置を補正するパラメータを表す。
Mは、フィルタ12から出力される動画像の符号化対象フレームのフレームレートを決定するパラメータ(ダウンサンプリング比)を表す。式(1)において、フィルタ12から出力される動画像の符号化対象フレームのフレームレートは、1/(Mδt)である。
乖離度の和Dkは、ΣM-1 i=0Φ[Wi,pi]である。なお、(2Δ+1≦M)が満たされている。
乖離度の和Dkは、ΣM-1 i=0Φ[Wi,pi]である。なお、(2Δ+1≦M)が満たされている。
式(1)に示された時間フィルタの特殊形として、フィルタ係数を一定値w[i]=1/(2Δ+1)とするフィルタを、「平均フィルタ」という。平均フィルタから出力される第iフレームは、式(3)のように表される。
N種類の係数候補ベクトルは、γn=(γn[-Δ],…,γn[Δ]),(n=0,…,N-1)と表記される。また、係数候補ベクトルから選択された係数ベクトルのフィルタ係数が施される時間位置は、P通りの時間位置から選択可能である。選択部16は、低い時間解像度の動画像の符号化対象フレームを生成するために最適な係数ベクトルを、(N×P)通りのN種類の係数候補ベクトルの中から、高い時間解像度の動画像のフレームごとに選択する。
以下では、係数候補ベクトルの集合を「辞書」という。表記の簡略化のため、N種類の係数候補ベクトルから構成される辞書を、「ΓN=(γ0,…,γN-1)」と表記する。
「フィルタ係数の最適化の規準]
選択部16は、フィルタ12から出力されたフレーム(時間フィルタによって生成されたフレーム)における発生符号量を、時間フィルタの設計における最適化の規準として用いる。発生符号量を表す情報は、動き補償予測を用いた可逆符号化を実行する符号化器13から得られる。
選択部16は、フィルタ12から出力されたフレーム(時間フィルタによって生成されたフレーム)における発生符号量を、時間フィルタの設計における最適化の規準として用いる。発生符号量を表す情報は、動き補償予測を用いた可逆符号化を実行する符号化器13から得られる。
X個の画素から構成されているフレームをK個に分割して、分割された区間ごとに動き補償フレーム間予測を行う場合について説明する。以下では、数式において文字の上に記載される記号(例えば、^)は、その文字の直前に記載される。
取得部11は、フレーム^f(x,iMδt,Wi,pi)を、サイズ(X/K)の区間B[k](k=0,1,…,K-1)に分割する。符号化器13がフレーム^f(x,(i-1)Mδt,Wi-1,pi-1)を参照フレームとして、各区間B[k](k=0,1,…,K-1)に対して動き補償(変位量di=(di[0],…,di[K-1])を実行する場合、そのフレーム内の動き補償フレーム間予測誤差は、式(4)のように表現される。
この動き補償フレーム間予測誤差を符号化とする符号化器13から得られる発生符号量Ψは、式(5)のように表される。
ここで、Rhは、符号化器13が生成するヘッダー情報の符号量を表す。Rd(di[0],…,di[K-1])は、推定変位量di[0],…,di[K-1]に関する符号量を表す。Re(ei(0,Wi,Wi-1,pi,pi-1),…,ei(X-1,Wi,Wi-1,pi,pi-1))は、動き補償フレーム間予測誤差に関する符号量を表す。
可逆符号化を実行する符号化器13から得られる発生符号量Ψが用いられているため、動き補償フレーム間予測誤差は、符号化対象フレーム及び参照フレームのみに依存する。
したがって、式(5)における発生符号量Ψは、変位量と、ヘッダ情報と、第iフレームに対する係数ベクトルWiと、第iフレームに対してフィルタ係数が施される時間位置の補正パラメータであるシフト量piと、第(i-1)フレームに対する係数ベクトルWi-1と、第(i-1)フレーム対してフィルタ係数が施される時間位置の補正パラメータであるシフト量pi-1とにより定まる。
したがって、式(5)における発生符号量Ψは、変位量と、ヘッダ情報と、第iフレームに対する係数ベクトルWiと、第iフレームに対してフィルタ係数が施される時間位置の補正パラメータであるシフト量piと、第(i-1)フレームに対する係数ベクトルWi-1と、第(i-1)フレーム対してフィルタ係数が施される時間位置の補正パラメータであるシフト量pi-1とにより定まる。
選択部16は、サンプリングされた高い時間解像度の動画像のフレーム(原画像)と、フィルタ12から出力された動画像の符号化対象フレーム(フィルタリングにより生成された動画像のフレーム)との乖離度Φとして、式(6)の値を算出する。
選択部16は、式(5)に示された発生符号量Ψと式(6)に示された乖離度との加重和として、式(7)の値を算出する。式(7)の値は、フィルタ12のフィルタ設計の評価尺度を表す。
ここで、λは、加重和の係数として、乖離度Φ[Wi,pi]に乗算される係数である。λの値は、初期化部14によって初期化される。λの値は、更新部15によって更新される。選択部16は、初期化又は更新されたλの値を取得する。選択部16は、式(7)に示された評価尺度が最小化されるように、フィルタ12を設計する。発生符号量Ψ[Wi,Wi-1,pi,pi-1]に基づいてフィルタ12が設計されることで、フィルタリングにより生成された動画像の発生符号量を低減する効果が期待できる。また、乖離度Φ[Wi,pi]に基づいてフィルタ12が設計されることで、フィルタリングにより生成された動画像とサンプリングされた動画像とが乖離することを抑止する効果が期待できる。
[フィルタ係数の最適化]
選択部16は、発生符号量を最小化するフレームをフィルタ12が生成するように、式(5)の発生符号量をコスト関数とする最小化問題を解く。選択部16は、式(8)を満たす(J/M)組の係数ベクトル及びシフト量を算出する。Jは、高い時間解像度の動画像のフレーム数を表す。
選択部16は、発生符号量を最小化するフレームをフィルタ12が生成するように、式(5)の発生符号量をコスト関数とする最小化問題を解く。選択部16は、式(8)を満たす(J/M)組の係数ベクトル及びシフト量を算出する。Jは、高い時間解像度の動画像のフレーム数を表す。
(N×P)種類の係数ベクトルを係数候補ベクトルとする場合、係数ベクトル及びシフト量の組合せは、NPJ/M通りとなる。最適な係数ベクトルを選択部16が選択するためには、指数オーダの演算量が必要である。このため、係数ベクトル及びシフト量の最適な組み合わせ(W*
0,…,W*
J/M-1,p*
0,…,p*
J/M-1)を総当りで探索することは、演算量の観点から現実的ではない。
Ξ[Wi,Wi-1,pi,pi-1]がWi,pi及びWi-1,pi-1のみに依存することに着目すれば、式(8)は、単純マルコフ過程における最適化問題として定式化可能である。単純マルコフ過程における最適化問題は、動的計画法によって、多項式オーダの演算量で最適解を算出することが可能である。以下では、動的計画法を用いた解法を示す。
選択部16は、Wi,pi(i=1,…,J/M-1)に関して、式(9)のSi(Wi,pi)を定義する。
Si(Wi,pi)は、第iステージにおいて、シフト量piでシフトされたフィルタ係数Wiによって第iフレームが生成され、その状態に至る経路において最適な係数ベクトルが用いられた場合における、コストの総和を表す。ここで、Wi,piがそれぞれ固定された場合、Ξ[Wi,Wi-1,pi,pi-1]がWi-1,pi-1のみに依存することに着目すると、Si(Wi,pi)は、式(10)のような漸化式として表される。
なお、第(i-1)ステージにおいて、Si-1(Wi-1,pi-1)は、同様の漸化式を用いて算出済みである。Si-1(Wi-1,pi-1)は、Si(Wi,pi)の算出時には、参照可能な値としてレジスタ等の記憶部に記憶されている。
式(10)に示されているように、選択部16は、Ξ[Wi,Wi-1,pi,pi-1]+Si-1(Wi-1,pi-1)を最小化する係数候補ベクトル及びシフト量piを辞書ΓNから選択することによって、Si(Wi,pi)を算出する。
以下、係数ベクトルWiに対する係数候補ベクトルのインデックスをniと表記する。
選択部16は、式(10)が最小値となる係数候補ベクトルのインデックス^ni-1(ni,pi)を、niごとに記憶部10に記録する。選択部16は、式(10)が最小値となるシフト量^pi-1(ni,pi)を、niごとに記憶部10に記録する。これによって、選択部16は、式(10)が最小値となる係数候補ベクトルのインデックス^ni-1(ni,pi)と、シフト量^pi-1(ni,pi)とを、後段の処理において記憶部10から取得することができる。
選択部16は、式(10)が最小値となる係数候補ベクトルのインデックス^ni-1(ni,pi)を、niごとに記憶部10に記録する。選択部16は、式(10)が最小値となるシフト量^pi-1(ni,pi)を、niごとに記憶部10に記録する。これによって、選択部16は、式(10)が最小値となる係数候補ベクトルのインデックス^ni-1(ni,pi)と、シフト量^pi-1(ni,pi)とを、後段の処理において記憶部10から取得することができる。
式(10)の漸化式が再帰的に用いられることで、式(8)の最小化問題は、式(11)のように表される。
このように、式(10)の漸化式が再帰的に用いられる方法であれば、式(8)の最適解(W*
0,…,W*
J/M-1,p*
0,…,p*
J/M-1)は、{(NP)2J/M}通りの中から最適解を探索する問題に帰着される。これによって、選択部16は、多項式オーダの演算量で、式(11)を算出することが可能である。
ΣJ/M-1
i=1[Wi,Wi-1,pi,pi-1]の最小値が算出された場合、選択部16は、最適解(W*
0,…,W*
J/M-1,p*
0,…,p*
J/M-1)を、以下に示されるバックトラック過程によって算出する。
式(12)のSJ/M-1(wJ/M-1,pJ/M-1)を最小化する係数候補ベクトル及びシフト量の組み合わせは、(W*
J/M-1,p*
J/M-1)と表される。
W*
J/M-1を表す係数候補ベクトルのインデックスは、nJ/M-1と表される。
第(J/M-1)フレームに対する係数候補ベクトルのインデックスは、nJ/M-1と表される。第(J/M-1)フレームに対するシフト量は、pJ/M-1と表される。
第(J/M-1)フレームに対する係数候補ベクトルのインデックスは、nJ/M-1と表される。第(J/M-1)フレームに対するシフト量は、pJ/M-1と表される。
第(J/M-2)フレームに対する最適な係数候補ベクトルのインデックスは、^nJ/M-2(nJ/M-1,pJ/M-1)として、記憶部10に記憶されている。第(J/M-2)フレームに対するシフト量は、^pJ/M-2(nJ/M-1,pJ/M-1)として、記憶部10に記憶されている。
選択部16は、第(J/M-2)フレームに対する係数ベクトル及びシフト量の組み合わせを、{W*
J/M-2=γ^nJ/M-2(nJ/M-1,pJ/M-1),p*
J/M-2=^pJ/M-2(nJ/M-1,pJ/M-1)}と同定する。
選択部16は、バックトラック過程において、{W*
J/M-3=γ^nJ/M-3(nJ/M-2,pJ/M-2),p*
J/M-3=^pJ/M-3(nJ/M-2,pJ/M-2)},…,{W*
0=γ^n0(n1,p1),p*
0=^p0(n1,p1)}のように、同様の参照処理を繰り返す。このように、選択部16は、係数ベクトル及びシフト量の組み合わせを選択する。
次に、乖離度Φ[Wi,pi]に乗算される係数λの更新処理における、乖離度の制約について説明する。
式(6)に示された乖離度は、式(13)のように変形可能である。
式(6)に示された乖離度は、式(13)のように変形可能である。
式(7)に示された係数λの更新処理において、更新部15は、式(12)に示された、ΣM-1
k=0Σxf(x,(iM+k)δt)2の値を算出する。更新部15は、算出結果を必要に応じて参照するため、算出結果を記憶部10に記録する。
式(7)に示された係数λの更新処理において、更新部15は、式(12)に示された、ΣM-1
k=0Σxf(x,(iM+pi+j)δt)2の値と、ΣM-1
k=0Σxf(x,(iM+k)δt)f(x,(iM+pi+j)δt)の値と、ΣM-1
k=0Σxf(x,(iM+pi+j)δt)f(x,(iM+pi+j’)δt)の値とが算出過程において初出である場合、これらの値をpi及びjについて算出する。
更新部15は、更新された乖離度に基づいて選択部16が乖離度の和Dkを反復して算出する場合、乖離度を表す構成要素の前回までの算出結果を、必要に応じて記憶部10に予め記録する。更新部15は、このような構成要素の前回までの算出結果を、必要に応じて記憶部10から取得する。更新部15は、取得された前回までの算出結果に基づいて乖離度を更新する。
図2は、フィルタ選択装置1の動作例を示すフローチャートである。初期化部14は、反復回数を表すインデックスkを0に初期化する(ステップS101)。初期化部14は、パラメータを初期化する。例えば、初期化部14は、予め定められた初期値であるλ0の値にλkの値を初期化する(ステップS102)。選択部16は、上記の[フィルタ係数の最適化]の説明に記載された選択処理を実行する。選択部16は、係数λkを用いて、各ステージにおける係数ベクトル及びシフト量を選択する(ステップS103)。選択部16は、i=0からi=M-1までについて、乖離度Φ[Wi,pi]を加算する。すなわち、選択部16は、乖離度の和Dk=ΣM-1
i=0Φ[Wi,pi]を算出する(ステップS104)。
更新部15は、乖離度の和Dkが閾値Dthよりも大きいか否かを判定する(ステップS105)。乖離度の和Dkが閾値Dthよりも大きい場合(ステップS105:YES)、更新部15は、λk-1がλkよりも小さいか否かを判定する(ステップS106)。λk-1がλkよりも小さい場合(ステップS106:YES)、更新部15は、λk+1の値を2λkとする(ステップS107)。λk-1がλk以上である場合(ステップS106:NO)、更新部15は、λk+1の値をλk/2とする(ステップS108)。
更新部15は、乖離度の和Dkが閾値Dth未満である場合(Dkが(Dth-ε)以下である場合)(ステップS105:NO)、更新部15は、λk-1がλkよりも大きいか否かを判定する(ステップS109)。λk-1がλkよりも大きい場合(ステップS109:YES)、更新部15は、λk+1の値を2λkとする。λk-1がλk以下である場合(ステップS109:NO)、更新部15は、λk+1の値をλk/2とする(ステップS111)。更新部15は、インデックスkの値を1インクリメントする(ステップS112)。
選択部16は、収束条件が満たされているか否かを判定する(ステップS113)。収束条件は、例えば、Dth-ε≦Dk<Dthが満たされるという条件である。収束条件は、例えば、指定の閾値に反復回数が到達するという条件でもよい。収束条件が満たされていない場合(ステップS113:NO)、選択部16は、ステップS103に処理を戻す。収束条件が満たされている場合(ステップS113:YES)、選択部16は、図2に示された処理を終了する。このようにして、選択部16は、辞書に記録された係数ベクトル及びシフト量の組み合わせを、時間フィルタのパラメータとして選択する。
以上のように、実施形態のフィルタ選択装置1は、フィルタ12と、選択部16とを備える。フィルタ12は、低い時間解像度の動画像の符号化対象フレームを、高い時間解像度の動画像のフレームから、フィルタ係数に応じて生成する。選択部16は、高い時間解像度の動画像のフレームと低い時間解像度の動画像の符号化対象フレームとの乖離度を符号化対象フレームごとに取得する。選択部16は、符号化対象フレームの発生符号量と乖離度との加重和を最小化するフィルタ係数を、予め定められた制約範囲内(Dth-ε≦Dk<Dth)に乖離度の和Dkが収まるように、フィルタ係数の候補の集合(辞書)から選択する。
これによって、実施形態のフィルタ選択装置1は、時間フィルタによって高フレームレート画像から生成される低フレームレート画像の発生符号量を符号化器が低減した上で、低フレームレート画像の客観画質を符号化器が一定以上に保つように、低フレームレート画像の符号化効率を向上させる時間フィルタの係数を選択することが可能である。
実施形態のフィルタ選択装置1は、フィルタリング後の動画像の主観画質を保持することが可能である。実施形態のフィルタ選択装置1は、フィルタリング後の動画像の発生符号量を低減することが可能である。実施形態のフィルタ選択装置1は、動画像のフレームに関する加重和の累積値を最小化する係数ベクトルを、選択することが可能である。
実施形態のフィルタ選択装置1は、更新部15を更に備えてもよい。更新部15は、フィルタ係数に応じた発生符号量に基づいて、他のフィルタ係数に応じた発生符号量の予測分布を取得してもよい。更新部15は、他のフィルタ係数に応じた発生符号量の予測値を、予測分布に基づいて取得してもよい。更新部15は、発生符号量の予測値が最小化する他のフィルタ係数に応じた発生符号量に基づいて、フィルタ係数の候補の集合を生成してもよい。
これによって、実施形態のフィルタ選択装置1は、係数ベクトルの辞書の動的更新を伴う時間フィルタリングにおいて、辞書設計処理及び選択処理の最適化が可能である。
以上、この発明の実施形態について図面を参照して詳述してきたが、具体的な構成はこの実施形態に限られるものではなく、この発明の要旨を逸脱しない範囲の設計等も含まれる。
上述した実施形態におけるフィルタ選択装置をコンピュータで実現するようにしてもよい。その場合、この機能を実現するためのプログラムをコンピュータ読み取り可能な記録媒体に記録して、この記録媒体に記録されたプログラムをコンピュータシステムに読み込ませ、実行することによって実現してもよい。なお、ここでいう「コンピュータシステム」とは、OSや周辺機器等のハードウェアを含むものとする。また、「コンピュータ読み取り可能な記録媒体」とは、フレキシブルディスク、光磁気ディスク、ROM、CD-ROM等の可搬媒体、コンピュータシステムに内蔵されるハードディスク等の記憶装置のことをいう。さらに「コンピュータ読み取り可能な記録媒体」とは、インターネット等のネットワークや電話回線等の通信回線を介してプログラムを送信する場合の通信線のように、短時間の間、動的にプログラムを保持するもの、その場合のサーバやクライアントとなるコンピュータシステム内部の揮発性メモリのように、一定時間プログラムを保持しているものも含んでもよい。また上記プログラムは、前述した機能の一部を実現するためのものであってもよく、さらに前述した機能をコンピュータシステムにすでに記録されているプログラムとの組み合わせで実現できるものであってもよく、FPGA(Field Programmable Gate Array)等のプログラマブルロジックデバイスを用いて実現されるものであってもよい。
本発明は、動画像の符号化装置に適用可能である。
1…フィルタ選択装置、10…記憶部、11…取得部、12…フィルタ、13…符号化器、14…初期化部、15…更新部、16…選択部
Claims (5)
- フィルタ選択装置が実行するフィルタ選択方法であって、
低い時間解像度の動画像の符号化対象フレームを、高い時間解像度の動画像のフレームから、フィルタ係数に応じて生成するステップと、
前記高い時間解像度の動画像のフレームと前記低い時間解像度の動画像の符号化対象フレームとの乖離度を前記符号化対象フレームごとに取得し、前記符号化対象フレームの発生符号量と前記乖離度との加重和を最小化する前記フィルタ係数を、予め定められた範囲内に前記乖離度の和が収まるように、前記フィルタ係数の候補の集合から選択するステップと
を含むフィルタ選択方法。 - 前記乖離度を更新するステップを更に含み、
前記選択するステップでは、更新された前記乖離度に基づいて前記乖離度の和を反復して算出し、
前記更新するステップでは、前記乖離度を表す構成要素の前回までの算出結果を記憶部に記録し、前記前回までの算出結果を前記記憶部から取得し、取得された前記前回までの算出結果に基づいて前記乖離度を更新する、請求項1に記載のフィルタ選択方法。 - 低い時間解像度の動画像の符号化対象フレームを、高い時間解像度の動画像のフレームから、フィルタ係数に応じて生成するフィルタと、
前記高い時間解像度の動画像のフレームと前記低い時間解像度の動画像の符号化対象フレームとの乖離度を前記符号化対象フレームごとに取得し、前記符号化対象フレームの発生符号量と前記乖離度との加重和を最小化する前記フィルタ係数を、予め定められた範囲内に前記乖離度の和が収まるように、前記フィルタ係数の候補の集合から選択する選択部と
を備えるフィルタ選択装置。 - 前記乖離度を更新する更新部を更に備え、
前記選択部は、更新された前記乖離度に基づいて前記乖離度の和を反復して算出し、
前記更新部は、前記乖離度を表す構成要素の前回までの算出結果を記憶部に記録し、前記前回までの算出結果を前記記憶部から取得し、取得された前記前回までの算出結果に基づいて前記乖離度を更新する、請求項3に記載のフィルタ選択装置。 - 請求項3又は請求項4に記載のフィルタ選択装置としてコンピュータを機能させるためのフィルタ選択プログラム。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2018-121738 | 2018-06-27 | ||
| JP2018121738A JP2020005081A (ja) | 2018-06-27 | 2018-06-27 | フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020003936A1 true WO2020003936A1 (ja) | 2020-01-02 |
Family
ID=68986490
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2019/022307 Ceased WO2020003936A1 (ja) | 2018-06-27 | 2019-06-05 | フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラム |
Country Status (2)
| Country | Link |
|---|---|
| JP (1) | JP2020005081A (ja) |
| WO (1) | WO2020003936A1 (ja) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2018006831A (ja) * | 2016-06-27 | 2018-01-11 | 日本電信電話株式会社 | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム |
| JP2018006830A (ja) * | 2016-06-27 | 2018-01-11 | 日本電信電話株式会社 | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム |
| JP2018006829A (ja) * | 2016-06-27 | 2018-01-11 | 日本電信電話株式会社 | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム |
-
2018
- 2018-06-27 JP JP2018121738A patent/JP2020005081A/ja active Pending
-
2019
- 2019-06-05 WO PCT/JP2019/022307 patent/WO2020003936A1/ja not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2018006831A (ja) * | 2016-06-27 | 2018-01-11 | 日本電信電話株式会社 | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム |
| JP2018006830A (ja) * | 2016-06-27 | 2018-01-11 | 日本電信電話株式会社 | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム |
| JP2018006829A (ja) * | 2016-06-27 | 2018-01-11 | 日本電信電話株式会社 | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム |
Also Published As
| Publication number | Publication date |
|---|---|
| JP2020005081A (ja) | 2020-01-09 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| ES2872724T3 (es) | Mejora de los datos visuales mediante el uso de redes neuronales actualizadas | |
| CN110741640B (zh) | 用于视频代码化中的运动补偿预测的光流估计 | |
| TW201830972A (zh) | 用於視訊寫碼之低複雜度符號預測 | |
| TWI429292B (zh) | 動作向量探索方法及裝置及其程式以及記錄有程式之記錄媒體 | |
| CN117395423B (zh) | 视频图像的处理方法、装置、电子设备和存储介质 | |
| US20240129473A1 (en) | Probability estimation in multi-symbol entropy coding | |
| CN118985129A (zh) | 用于数据处理的方法、装置和介质 | |
| JP6538619B2 (ja) | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム | |
| CN121488473A (zh) | 用于可视数据处理的方法、装置和介质 | |
| JP5102174B2 (ja) | フレームレート変換方法、フレームレート変換装置、フレームレート変換プログラムおよびそのプログラムを記録したコンピュータ読み取り可能な記録媒体 | |
| KR20240173786A (ko) | 영상 처리 장치 및 영상의 움직임 추정 방법 | |
| JP6626319B2 (ja) | 符号化装置、撮像装置、符号化方法、及びプログラム | |
| WO2020003936A1 (ja) | フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラム | |
| Jia et al. | Bit rate matching algorithm optimization in JPEG-AI verification model | |
| JP5118005B2 (ja) | フレームレート変換方法、フレームレート変換装置、フレームレート変換プログラムおよびそのプログラムを記録したコンピュータ読み取り可能な記録媒体 | |
| WO2019150411A1 (ja) | 映像符号化装置、映像符号化方法、映像復号装置、映像復号方法、及び映像符号化システム | |
| JP6595442B2 (ja) | 映像フィルタリング方法、映像フィルタリング装置及びコンピュータプログラム | |
| JP6680633B2 (ja) | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム | |
| CN119155452A (zh) | 基于内容生成的图像编解码方法、电子设备和存储介质 | |
| WO2020003933A1 (ja) | フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラム | |
| JP6611256B2 (ja) | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム | |
| JP7181492B2 (ja) | 復号装置、符号化装置、復号方法、符号化方法及びプログラム | |
| Bordin et al. | Fine color guidance in diffusion models and its application to image compression at extremely low bitrates | |
| JP4583514B2 (ja) | 動画像符号化装置 | |
| JP2002247587A (ja) | 画像符号化データの再符号化装置、再符号化方法、再符号化プログラム及び再符号化プログラムを記録した記録媒体 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19824786 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19824786 Country of ref document: EP Kind code of ref document: A1 |










