WO2020003933A1 - フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラム - Google Patents
フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラム Download PDFInfo
- Publication number
- WO2020003933A1 WO2020003933A1 PCT/JP2019/022274 JP2019022274W WO2020003933A1 WO 2020003933 A1 WO2020003933 A1 WO 2020003933A1 JP 2019022274 W JP2019022274 W JP 2019022274W WO 2020003933 A1 WO2020003933 A1 WO 2020003933A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- filter
- frame
- moving image
- filter coefficient
- coefficient
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/102—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
- H04N19/117—Filters, e.g. for pre-processing or post-processing
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/136—Incoming video signal characteristics or properties
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/134—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
- H04N19/146—Data rate or code amount at the encoder output
- H04N19/149—Data rate or code amount at the encoder output by estimating the code amount by means of a model, e.g. mathematical model or statistical model
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/10—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
- H04N19/169—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
- H04N19/17—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
- H04N19/172—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a picture, frame or field
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/50—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding
- H04N19/587—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using predictive coding involving temporal sub-sampling or interpolation, e.g. decimation or subsequent interpolation of pictures in a video sequence
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N19/00—Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
- H04N19/80—Details of filtering operations specially adapted for video compression, e.g. for pixel interpolation
Definitions
- the present invention relates to a filter selection method, a filter selection device, and a filter selection program.
- High image quality during image reproduction is intended to express the smooth movement of the subject by approaching the upper limit of the frame rate that can be detected by the visual system (displayable on the display). Therefore, high image quality during image reproduction is based on the premise that the display device reproduces moving images at a constant speed.
- the purpose of improving the accuracy of image analysis is to improve the accuracy of image analysis by using a high frame rate image that exceeds the detection limit of vision.
- Image analysis by slow reproduction of a high-speed moving object such as an athlete, FA / inspection, or car is a typical application example.
- the upper limit of the frame rate of the moving image input system and the upper limit of the frame rate of the moving image output system are asymmetric. That is, the upper limit of the frame rate of the high-speed camera as the moving image input system exceeds 10,000 fps.
- the upper limit of the frame rate of the display device, which is a moving image output system is from 120 fps to 240 fps. For this reason, a moving image captured by a high-speed camera is used for slow reproduction (see Patent Document 1).
- the high frame rate image includes a group of frames sampled at high density in the time direction. If the image generation apparatus generates an image for constant speed reproduction at 30 Hz or the like by using a frame group sampled at a high density such as 1000 Hz, the generation of the image for constant speed reproduction can be controlled with high time resolution. It is possible.
- the image generation device samples a frame at a reproduction frame rate. For this reason, the conventional image generation device does not sample a frame at a time resolution higher than the reproduction frame rate.
- the degree of freedom in filter design is expanded.
- a moving image signal of 62.5 fps is generated from a moving image signal of 1000 fps by filtering a moving image signal of 1000 fps
- the frame to be filtered does not overlap even if the frames to be filtered do not overlap.
- Can be 16 ( 1000 / 62.5) frames, which is more than 2 frames.
- the degree of freedom in filtering design is high. By utilizing this high degree of freedom, the encoder may be able to improve the coding efficiency.
- the conventional filter selection device may not be able to select a temporal filter coefficient that improves the coding efficiency of a low frame rate image generated from a high frame rate image in preprocessing of video coding.
- the present invention provides a filter selection method, a filter selection device, and a filter selection method capable of selecting a coefficient of a temporal filter for improving the coding efficiency of a low frame rate image generated from a high frame rate image. It is intended to provide programs.
- One aspect of the present invention is a filter selection method executed by a filter selection device, wherein an encoding target frame of a low temporal resolution moving image is generated from a high temporal resolution moving image frame according to a filter coefficient. Step, the degree of divergence between the frame of the moving image of the high time resolution and the encoding target frame of the moving image of the low time resolution is obtained for each encoding target frame, and the generated code amount of the encoding target frame Selecting the filter coefficient that minimizes the weighted sum with the degree of divergence from a set of candidates for the filter coefficient, and according to the other filter coefficient based on the generated code amount corresponding to the filter coefficient Obtaining a predicted distribution of the generated code amount, obtaining a predicted value of the generated code amount according to the other filter coefficient based on the predicted distribution, Value is a filter selection method comprising the steps of: generating a set of candidates for the filter coefficients based on the generated code amount according to the other filter coefficients to be minimized.
- One embodiment of the present invention is a filter that generates a coding target frame of a low temporal resolution moving image from a frame of a high temporal resolution moving image according to a filter coefficient, and a frame of the high temporal resolution moving image.
- the filter that obtains the degree of divergence from the encoding target frame of the low temporal resolution moving image for each of the encoding target frames, and minimizes the weighted sum of the generated code amount of the encoding target frame and the degree of divergence.
- One embodiment of the present invention is a filter selection program for causing a computer to function as the above-described filter selection device.
- the present invention it is possible to select a coefficient of a temporal filter that improves the coding efficiency of a low frame rate image generated from a high frame rate image.
- FIG. 6 is a flowchart illustrating an operation example of the filter selection device. It is a flowchart which shows the operation example of a dictionary design process. It is a figure showing an example of operation of selection processing of a coefficient vector and a shift amount. 13 is a flowchart illustrating an operation example of a selection process of a coefficient vector and a shift amount.
- FIG. 1 is a diagram illustrating a configuration example of the filter selection device 1.
- the filter selection device 1 is an information processing device that selects a coefficient of a filter.
- the filter selection device 1 includes a storage unit 10, an acquisition unit 11, a filter 12, an encoder 13, an initialization unit 14, a set generation unit 15, and a selection unit 16.
- Part or all of the functional units are realized by, for example, a processor such as a CPU (Central Processing Unit) executing a program stored in a storage unit.
- a processor such as a CPU (Central Processing Unit) executing a program stored in a storage unit.
- Some or all of the functional units may be realized using hardware such as an LSI (Large Scale Integration) or an ASIC (Application Specific Integrated Circuit).
- the storage unit 10 is a non-volatile recording medium (non-temporary recording medium) such as a flash memory and an HDD (Hard Disk Drive).
- the storage unit 10 may include a volatile recording medium such as a RAM (Random Access Memory) and a register.
- the storage unit 10 stores, for example, a frame group of a moving image with a high time resolution before filtering processing, a data table, and a program.
- the acquisition unit 11 acquires a high-resolution video frame (high frame rate image) as an original image of the encoding target frame.
- the acquisition unit 11 records a frame group of a moving image having a high time resolution in the storage unit 10.
- Filter 12 is a temporal filter. The filter 12 acquires a frame of a moving image with a high time resolution from the acquiring unit 11.
- the filter 12 performs a filtering process (downsampling) on a frame of a moving image having a high temporal resolution based on the coefficient vector and the shift amount selected by the selecting unit 16.
- the filter 12 outputs the encoding target frame (low frame rate image) of the low temporal resolution moving image generated by the filtering process to the encoder 13 and the selection unit 16.
- the encoder 13 performs lossless encoding using motion compensation prediction (Motion Compensation Prediction) on the encoding target frame of the moving image with low temporal resolution.
- the encoder 13 outputs information indicating the generated code amount of the encoding target frame of the moving image with low temporal resolution to the set generation unit 15 and the selection unit 16.
- the initialization unit 14 initializes a shift amount of a time position (downsampling time by filtering) at which a filter coefficient is applied by the filtering process with respect to the filtering process of each frame (stage).
- the set generation unit 15 acquires, from the encoder 13, information indicating the generated code amount of the encoding target frame of the moving image with low temporal resolution.
- the set generation unit 15 generates a set of selected coefficient vector candidates (hereinafter, referred to as “coefficient candidate vectors”).
- the set generation unit 15 optimizes (updates) the generated set of coefficient candidate vectors by, for example, a method based on Bayesian optimization so that the generated code amount of the encoding target frame is minimized.
- the selection unit 16 acquires, from the storage unit 10, a frame group of a moving image having a high time resolution before the filtering process.
- the selecting unit 16 acquires, from the filter 12, a coding target frame of a moving image with a low temporal resolution after the filtering process.
- the selecting unit 16 acquires, from the encoder 13, information indicating the generated code amount of the encoding target frame of the moving image with low time resolution.
- the selection unit 16 calculates, for each encoding target frame, the degree of divergence between the high-resolution video frame before the filtering process and the encoding target frame of the low-resolution moving image after the filtering process.
- the selecting unit 16 calculates a weighted sum of the generated code amount of the encoding target frame of the moving image with low temporal resolution and the degree of deviation.
- the selection unit 16 selects a coefficient vector (filter coefficient) that minimizes the weighted sum from a set of optimized coefficient candidate vectors based on the weighted sum of the generated code amount and the degree of deviation.
- the selection unit 16 generates an encoding target frame having a low temporal resolution from the set of coefficient candidate vectors optimized by the set generation unit 15 based on the weighted sum of the generated code amount and the degree of divergence.
- the optimal coefficient vector for The selecting unit 16 may select the shift amount of the time position (down sampling time by filtering) at which the filter coefficient is applied, based on the weighted sum of the generated code amount and the divergence.
- the encoding target frame of the low temporal resolution moving image suitable for encoding is obtained by using the temporal filter based on the coefficient vector and the shift amount selected by the selecting unit 16 to obtain the high temporal resolution moving image frame. Generated from a group.
- each frame of the moving image is represented as a one-dimensional signal.
- the i-th frame generated based on the (2 ⁇ + 1) -tap temporal filter is represented by Expression (1).
- i an index designating a frame after downsampling.
- the index takes a non-negative integer value.
- [delta] t represents the distance between the frame of a moving image to be input to the filter 12.
- w i [j] represents a filter coefficient for a reference frame used for motion compensation prediction.
- w i [j] satisfies the relationship of equation (2).
- the coefficient vector the formula in the text are capitalized as W i.
- p i represents a parameter for correcting the time position at which the filter coefficient is applied.
- M represents a parameter (downsampling ratio) that determines the frame rate of the encoding target frame of the moving image output from the filter 12.
- the frame rate of the encoding target frame of the moving image output from the filter 12 is 1 / (M ⁇ t ). Note that (2 ⁇ + 1 ⁇ M) is satisfied.
- the selecting unit 16 selects a coefficient vector optimal for generating a frame to be coded for a moving image with a low temporal resolution from among (N ⁇ P) N types of coefficient candidate vectors, Select for each frame.
- a set of coefficient candidate vectors is referred to as a “dictionary”.
- the selecting unit 16 uses the generated code amount in the frame output from the filter 12 (the frame generated by the time filter) as a criterion for optimization in designing the time filter.
- Information indicating the generated code amount is obtained from the encoder 13 that executes lossless encoding using motion compensation prediction.
- R h represents a code amount of the header information coder 13 is produced.
- R d (d i [0] , ..., d i [K-1]) is the estimated displacement amount d i [0], ..., representing the code amount related to d i [K-1].
- R e (e i (0, W i, W i-1, p i, p i-1), ..., e i (X-1, W i, W i-1, p i, p i-1) ) Indicates the code amount related to the motion compensation inter-frame prediction error.
- the generated code amount ⁇ ⁇ ⁇ ⁇ in Expression (5) is a displacement amount, header information, a coefficient vector W i for the i-th frame, and a correction parameter of a time position at which a filter coefficient is applied to the i-th frame.
- the shift amount p i , the coefficient vector W i-1 for the (i-1) th frame, and the shift amount p i-1 which is a correction parameter of a time position at which a filter coefficient is applied to the (i-1) th frame Is determined by
- the selecting unit 16 determines the degree of divergence between the sampled high-resolution video frame (original image) and the encoding target frame of the video output from the filter 12 (the video frame generated by filtering).
- the value of Expression (6) is calculated as ⁇ .
- the selection unit 16 calculates the value of Expression (7) as a weighted sum of the generated code amount shown in Expression (5) and the degree of divergence shown in Expression (6).
- the value of Expression (7) represents an evaluation scale of the filter design of the filter 12.
- ⁇ is a coefficient by which the degree of deviation ⁇ [W i , p i ] is multiplied as a coefficient of the weighted sum.
- the value of ⁇ is predetermined, for example.
- the selection unit 16 designs the filter 12 such that the evaluation scale shown in Expression (7) is minimized.
- the filter 12 By designing the filter 12 based on the generated code amount ⁇ [W i , W i ⁇ 1 , p i , p i ⁇ 1 ], an effect of reducing the generated code amount of a moving image generated by filtering is expected. it can.
- the filter 12 based on the degree of divergence ⁇ [W i , p i ] an effect of suppressing the divergence between the moving image generated by the filtering and the sampled moving image can be expected.
- the selecting unit 16 solves a minimization problem in which the generated code amount in Expression (5) is used as a cost function so that the filter 12 generates a frame that minimizes the generated code amount.
- the selector 16 calculates (J / M) sets of coefficient vectors and the shift amount that satisfy the expression (8). J represents the number of frames of a moving image with a high temporal resolution.
- equation (8) gives: It can be formulated as an optimization problem in a simple Markov process.
- an optimal solution can be calculated by a dynamic programming method with an operation amount of a polynomial order.
- a solution method using the dynamic programming method will be described.
- S i-1 (W i-1 , p i-1 ) has already been calculated using the same recurrence formula.
- S i ⁇ 1 (W i ⁇ 1 , p i ⁇ 1 ) is stored in a storage unit such as a register as a referenceable value when S i (W i , p i ) is calculated.
- the selection unit 16 determines that [[W i , W i ⁇ 1 , p i , p i ⁇ 1 ] + S i ⁇ 1 (W i ⁇ 1 , p i ⁇ 1 ) S i (W i , p i ) is calculated by selecting from the dictionary ⁇ ⁇ N a coefficient candidate vector and a shift amount p i that minimize S i .
- Selection unit 16 records the index of the coefficient candidate vectors formula (10) becomes a minimum value ⁇ n i-1 (n i , p i) of the storage unit 10 for each n i.
- the selection unit 16 records the shift amount ⁇ p i-1 (n i , p i ) at which the equation (10) becomes the minimum value in the storage unit 10 for each n i . Accordingly, the selection unit 16 determines the index ⁇ n i-1 (n i , p i ) of the coefficient candidate vector that minimizes the expression (10) and the shift amount ⁇ p i-1 (n i , p i ). Can be obtained from the storage unit 10 in the subsequent processing.
- the selection unit 16 can calculate the equation (11) with the operation amount of the polynomial order.
- the selecting unit 16 selects the optimal solution (W * 0 ,..., W * J / M-1 , p * 0 ,..., P * J / M-1 ) are calculated by the backtrack process shown below.
- the index of the coefficient candidate vector representing W * J / M-1 is represented as nJ / M-1 .
- the index of the coefficient candidate vector for the (J / M-1) th frame is represented as nJ / M-1 .
- the shift amount for the (J / M-1) th frame is expressed as pJ / M-1 .
- the index of the optimal coefficient candidate vector for the (J / M-2) th frame is stored in the storage unit 10 as ⁇ n J / M-2 (n J / M ⁇ 1 , p J / M ⁇ 1 ). I have.
- the shift amount for the (J / M-2) th frame is stored in the storage unit 10 as ⁇ p J / M-2 (n J / M ⁇ 1 , p J / M ⁇ 1 ).
- the selection unit 16 selects a combination of a coefficient vector and a shift amount.
- Dictionary design refers to designing each coefficient candidate vector so as to minimize the amount of generated code of a moving image.
- the selection unit 16 selects an optimal coefficient vector from a dictionary of coefficient candidate vectors generated by the set generation unit 15.
- the objective function of the dictionary design is a function of the generated code amount ⁇ when the optimal coefficient vector and the optimal shift amount are set for the filter 12 by the selection unit 16.
- the objective function for dictionary design is expressed as in equation (13).
- the optimization problem of the dictionary design is a problem of minimizing the objective function of the dictionary design in a search space.
- the dictionary N includes N types of coefficient candidate vectors, and has (2 ⁇ + 1) coefficient candidate vectors as elements. Therefore, the dictionary ⁇ corresponds to a sample point (observation point) in the ⁇ (2 ⁇ + 1) N ⁇ -dimensional search space.
- the objective function shown in equation (13) is non-linear, non-differentiable, and non-convex. For this reason, it is impossible to analytically determine the minimum value of the objective function shown in Expression (13). Further, the cost of calculating the objective function in the search space is high.
- the search space exponentially increases as the tap length of the filter 12 increases and the number of coefficient candidate vectors increases. For this reason, a method based on a full search, such as a grid search, is not realistic because the amount of calculation is enormous.
- metaheuristic methods such as a genetic algorithm (Genetic Algorithm) and particle swarm optimization (PSO: Particle Swarm Optimization) require a large amount of sampling for an objective function. Therefore, when the cost of calculating the objective function is high, the meta-heuristic method is not practical.
- the set generation unit 15 executes dictionary design by a method based on Bayesian optimization.
- the method based on Bayesian optimization is a method suitable for a multidimensional search using limited sample points. This is because the Bayesian optimization estimates a cost function based on Bayesian estimation of a Gaussian process.
- the set generation unit 15 estimates the relationship between the dictionary and the cost function based on Bayesian optimization.
- the set generation unit 15 identifies an optimal dictionary that minimizes the cost function based on the estimation result.
- ⁇ i represents the cost function of the i-th dictionary.
- a set of cost functions ⁇ 1 ,..., ⁇ m ⁇ is denoted as ⁇ 1: m .
- I i represents the ith dictionary.
- the observation model shown in Expression (14) is used.
- ⁇ i represents noise at the time of observation.
- N (0, ⁇ 2 ) represents a Gaussian distribution with mean 0 and variance ⁇ 2 .
- h represents an unknown function assumed to be obtained from a Gaussian process as a prior distribution.
- the function value set h 1: m is a set obtained based on the multidimensional Gaussian distribution N (0, K ( ⁇ 1 : m )).
- K ( ⁇ 1 : m ) is an (m ⁇ m) matrix.
- K: The (i, j) element of (gamma 1 m) is a covariance function k ( ⁇ i, ⁇ j) .
- a “Matern5 / 2 kernel function” is used as a covariance function.
- [Selection process of search point of coefficient vector and shift amount] In Bayesian optimization, search points that are expected to minimize observations are sequentially selected from observation points.
- the set generation unit 15 sequentially selects a combination of a coefficient vector and a shift amount expected to minimize the cost function from the observation points.
- the set generation unit 15 obtains the posterior distribution of the unknown function h based on the Bayes rule.
- the set generation unit 15 analytically calculates a prediction distribution such as a Bayesian prediction distribution of the observed value ⁇ in the unknown sample ⁇ based on the obtained posterior distribution, as in Expressions (15) to (17). I do.
- k ( ⁇ ) (k ( ⁇ , ⁇ 1 ),..., K ( ⁇ , ⁇ m )) T.
- ⁇ 1: m ( ⁇ 1 , ..., ⁇ m) is T.
- T represents transposition.
- I is an (m ⁇ m) unit matrix.
- the set generation unit 15 generates an acquisition function based on the Bayesian prediction distribution shown in Expression (15).
- the acquisition function is an acquisition function of the lower confidence limit (lower confidence bound) shown in Expression (18).
- ⁇ m represents a hyper parameter that controls a trade-off between search and use.
- the set generation unit 15 selects the next search point so as to minimize the acquisition function.
- FIG. 2 is a flowchart illustrating an operation example of the filter selection device 1.
- the initialization unit 14 initializes a shift amount in each stage (each index i). For example, the initialization unit 14 sets the shift amounts in all stages to zero values (step S101).
- the set generation unit 15 executes the dictionary design processing described in the above description of [dictionary design].
- Bayesian optimization the parameter that minimizes the cost function is calculated as a continuous quantity.
- the shift amount is limited to a discrete value (frame downsampling time). For this reason, the shift amount that is a discrete value cannot be calculated by Bayesian optimization. Therefore, the set generation unit 15 optimizes the dictionary by executing the Bayesian optimization process with the shift amount fixed (step S102).
- the selection unit 16 executes the selection processing described in the above description of the “optimization of filter coefficients”.
- the selection unit 16 selects a coefficient vector and a shift amount in each stage with the optimized dictionary (set of coefficient candidate vectors) fixed (step S103).
- the selection unit 16 determines whether the convergence condition is satisfied (Step S104). If the convergence condition is not satisfied (step S104: NO), the selection unit 16 returns the process to step S102. When the convergence condition is satisfied (step S104: YES), the selection unit 16 ends the processing illustrated in FIG. In this way, the selection unit 16 selects a combination of the coefficient vector and the shift amount recorded in the dictionary as parameters of the time filter.
- FIG. 3 is a flowchart illustrating an operation example of the dictionary design process (step S102).
- the set generation unit 15 specifies K initial sample points and obtains the observation value D for each sample point. K is a parameter given from the outside (step S201).
- the set generation unit 15 calculates a Bayesian prediction distribution of the observation value D1 : m as a prediction distribution of the observation value D based on Expression (15) (Step S202).
- a lower confidence limit (predicted observed value) is calculated (step S203).
- the set generation unit 15 identifies a sample point having the minimum confidence lower limit, and determines the identified sample point as the next search point. That is, the set generation unit 15 next observes the observed value of the sample point at which the lower confidence limit is minimized (Step S204).
- the set generation unit 15 determines whether the convergence condition is satisfied.
- the convergence condition is, for example, a condition that a reduction amount of the minimum value of the observation value is lower than a specified threshold.
- the convergence condition may be, for example, a condition that the number of iterations reaches a specified threshold (step S205).
- step S205 NO
- the set generation unit 15 returns the process to step S202.
- step S205: YES the set generation unit 15 ends the processing illustrated in FIG.
- FIG. 4 is a diagram showing an operation example of the process of selecting a coefficient vector and a shift amount (step S103).
- FIG. 5 is a flowchart illustrating an operation example of the selection processing of the coefficient vector and the shift amount.
- the acquisition unit 11 reads a video signal (frame of a moving image) captured as a moving image, the number J of frames of a moving image with a high temporal resolution, and a frame rate of the moving image with a high temporal resolution (step S301). .
- the filter 12 reads the downsampling ratio M from the selection unit 16 (Step S302).
- Selection unit 16 a reference frame ⁇ f (x-d i [ k], (i-1) M ⁇ t, W i-1, p i-1) in the case of a, ⁇ f (x, iM ⁇ t , A motion vector that minimizes a motion compensation inter-frame prediction error for Wi , p i ) is obtained.
- the method of obtaining the motion vector is externally given to the selection unit 16.
- a method of obtaining a motion vector is a full search method of searching for all candidates within a search range.
- the selecting unit 16 stores the motion-compensated inter-frame prediction error when the motion vector is used in the storage unit 10 as ⁇ 2 i [W i , W i ⁇ 1 , p i , p i ⁇ 1 ] (step S308). ).
- the selection unit 16 records in the storage unit 10 Wi-1 giving S i (W i , p i ) as ⁇ W i-1 (W i , p i ).
- the selection unit 16 records p i-1 giving S i (W i , p i ) in the storage unit 10 as ⁇ p i-1 (W i , p i ) (step S311).
- the selection unit 16 stores the coefficient vector W J / M-1 that minimizes S J / M ⁇ 1 (W J / M ⁇ 1 , p J / M ⁇ 1 ) as W * J / M ⁇ 1. Record at 10.
- the selector 16 stores the shift amount pJ / M-1 that minimizes SJ / M-1 (WJ / M-1 , pJ / M-1 ) as p * J / M-1. 10 (step S312).
- the filter selection device 1 includes the filter 12, the selection unit 16, and the set generation unit 15.
- the filter 12 generates an encoding target frame of a low temporal resolution moving image from a high temporal resolution moving image frame according to a filter coefficient.
- the selecting unit 16 obtains the degree of divergence between the frame of the moving image with the high temporal resolution and the encoding target frame of the moving image with the low temporal resolution for each encoding target frame.
- the selection unit 16 selects a filter coefficient that minimizes the weighted sum of the generated code amount of the encoding target frame and the degree of deviation from a set (dictionary) of candidate filter coefficients.
- the set generation unit 15 obtains, based on the generated code amount corresponding to the filter coefficient, a predicted distribution of the generated code amount corresponding to another filter coefficient.
- the set generation unit 15 acquires a predicted value of the generated code amount according to another filter coefficient based on the predicted distribution.
- the set generation unit 15 generates a set of filter coefficient candidates based on the generated code amount corresponding to another filter coefficient whose predicted value of the generated code amount is minimized.
- the filter selection device 1 can select a coefficient of the temporal filter that improves the coding efficiency of the low frame rate image generated from the high frame rate image.
- the filter selection device 1 according to the embodiment can maintain the subjective image quality of the moving image after the filtering.
- the filter selection device 1 according to the embodiment can reduce the generated code amount of a moving image after filtering.
- the filter selection device 1 of the embodiment is capable of selecting a coefficient vector that minimizes the cumulative value of the weighted sum of frames of a moving image.
- the filter selection device 1 of the embodiment is capable of optimizing the dictionary design process and the selection process in temporal filtering involving dynamic updating of the dictionary of coefficient vectors.
- the filter selection device in the above-described embodiment may be realized by a computer.
- a program for realizing this function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read and executed by a computer system.
- the “computer system” includes an OS and hardware such as peripheral devices.
- the “computer-readable recording medium” refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, and a CD-ROM, and a storage device such as a hard disk built in a computer system.
- a “computer-readable recording medium” refers to a communication line for transmitting a program via a network such as the Internet or a communication line such as a telephone line, which dynamically holds the program for a short time.
- a program may include a program that holds a program for a certain period of time, such as a volatile memory in a computer system serving as a server or a client in that case.
- the program may be for realizing a part of the functions described above, or may be a program that can realize the functions described above in combination with a program already recorded in a computer system, It may be realized by using a programmable logic device such as an FPGA (Field Programmable Gate Array).
- FPGA Field Programmable Gate Array
- the present invention is applicable to a moving picture coding apparatus.
- # 1 filter selection device, 10: storage unit, 11 ... acquisition unit, 12 ... filter, 13 ... encoder, 14 ... initialization unit, 15 ... set generation unit, 16 ... selection unit
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Signal Processing (AREA)
- Physics & Mathematics (AREA)
- Algebra (AREA)
- General Physics & Mathematics (AREA)
- Mathematical Analysis (AREA)
- Mathematical Optimization (AREA)
- Pure & Applied Mathematics (AREA)
- Compression Or Coding Systems Of Tv Signals (AREA)
Abstract
フィルタ選択方法は、フィルタ選択装置が実行するフィルタ選択方法であって、低い時間解像度の動画像の符号化対象フレームを、高い時間解像度の動画像のフレームから、フィルタ係数に応じて生成するステップと、高い時間解像度の動画像のフレームと低い時間解像度の動画像の符号化対象フレームとの乖離度を符号化対象フレームごとに取得し、符号化対象フレームの発生符号量と乖離度との加重和を最小化するフィルタ係数を、フィルタ係数の候補の集合から選択するステップと、発生符号量の予測値が最小化する他のフィルタ係数に応じた発生符号量に基づいてフィルタ係数の候補の集合を生成するステップとを含む。
Description
本発明は、フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラムに関する。
昨今の半導体技術の進歩を受け、高速度カメラにおける動画像のフレームレートが大きく向上している。高速度カメラにより取得された高フレームレート画像の用途は、画像再生時の高画質化と画像解析の高精度化とに分類される。
画像再生時の高画質化は、視覚系で検知可能(ディスプレイで表示可能)なフレームレートの上限に迫ることにより、被写体の滑らかな動きを表現することが目的である。このため、画像再生時の高画質化は、ディスプレイ装置が動画像を等速再生することが前提である。
一方、画像解析の高精度化は、視覚の検知限を越えた高フレームレート画像を用いることにより、画像解析の高精度化を行うことが目的である。スポーツ選手、FA・検査、自動車等の高速移動物体のスロー再生による画像解析は、代表的な応用例である。
動画像の入力システムのフレームレートの上限と動画像の出力システムのフレームレートの上限とは非対称である。すなわち、動画像の入力システムである高速度カメラのフレームレートの上限は、10000fpsを超えている。一方、動画像の出力システムであるディスプレイ装置のフレームレートの上限は、120fpsから240fpsまでである。このため、高速度カメラで撮影された動画像は、スロー再生に用いられる(特許文献1参照)。
視覚の検知限を越えた高フレームレート画像を用いることにより、動画像の符号化処理に対して親和性の高い等速再生用の画像を生成することができる。高フレームレート画像は、時間方向に高密度でサンプリングされたフレーム群を含んでいる。画像生成装置は、1000Hz等の高密度時間サンプリングされたフレーム群を用いて30Hz等の等速再生用の画像を生成すれば、等速再生用の画像の生成を高い時間分解能で制御することが可能である。
しかしながら、符号発生量の低減を目的とした動画像符号化の前処理では、画像生成装置が再生フレームレートでフレームをサンプリングすることが前提となっている。このため、従来の画像生成装置は、再生フレームレートよりも高い時間分解能ではフレームをサンプリングしていない。
高フレームレート画像のフレームを単純に間引く処理では、時間方向のエイリアシングに起因する画質劣化が問題となる。このような問題を回避するには、時間フィルタによる時間軸方向の帯域制限フィルタリングが必要である。
一方、動き補償フレーム間予測を用いる符号化器では、時間方向のエイリアシングの低減は、予測誤差の低減に直接の関係がない。また、動き補償フレーム間予測を用いる符号化器では、高密度時間サンプリングされたフレームが十分に活用されておらず、時間フィルタとしての自由度には制約がある。
すなわち、30fps又は60fps等の低フレームレート画像の場合、フィルタリングのための十分な数のサンプル(フレーム)が確保できないため、フィルタの特性を高精度に近似することは困難である。例えば、60fpsの動画像信号をフィルタリングすることによって60fpsの動画像信号から30fpsの動画像信号が生成される場合、フィルタリングの対象のフレームが重複しないという条件下では、フィルタリングの対象のフレームは2(=60/30)フレームに限定されるという制約がある。
一方、高フレームレート画像の場合、フィルタ設計の自由度は拡張される。例えば、1000fpsの動画像信号をフィルタリングすることによって、1000fpsの動画像信号から62.5fpsの動画像信号が生成される場合、フィルタリングの対象のフレームが重複しないという条件下でも、フィルタリングの対象のフレームは、2フレームよりも多い16(=1000/62.5)フレームとすることができる。このように、高フレームレート画像から低フレームレート画像を生成する場合、フィルタリング設計の自由度は高い。この自由度の高さを利用することで、符号化器は符号化効率を向上させることができる可能性がある。
しかしながら、従来のフィルタ選択装置は、動画像符号化の前処理において、高フレームレート画像から生成される低フレームレート画像の符号化効率を向上させる時間フィルタの係数を選択することができない場合があった。
上記事情に鑑み、本発明は、高フレームレート画像から生成される低フレームレート画像の符号化効率を向上させる時間フィルタの係数を選択することが可能であるフィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラムを提供することを目的としている。
本発明の一態様は、フィルタ選択装置が実行するフィルタ選択方法であって、低い時間解像度の動画像の符号化対象フレームを、高い時間解像度の動画像のフレームから、フィルタ係数に応じて生成するステップと、前記高い時間解像度の動画像のフレームと前記低い時間解像度の動画像の符号化対象フレームとの乖離度を前記符号化対象フレームごとに取得し、前記符号化対象フレームの発生符号量と前記乖離度との加重和を最小化する前記フィルタ係数を、前記フィルタ係数の候補の集合から選択するステップと、前記フィルタ係数に応じた前記発生符号量に基づいて他の前記フィルタ係数に応じた前記発生符号量の予測分布を取得し、前記他のフィルタ係数に応じた前記発生符号量の予測値を前記予測分布に基づいて取得し、前記予測値が最小化する前記他のフィルタ係数に応じた前記発生符号量に基づいて前記フィルタ係数の候補の集合を生成するステップとを含むフィルタ選択方法である。
本発明の一態様は、低い時間解像度の動画像の符号化対象フレームを、高い時間解像度の動画像のフレームから、フィルタ係数に応じて生成するフィルタと、前記高い時間解像度の動画像のフレームと前記低い時間解像度の動画像の符号化対象フレームとの乖離度を前記符号化対象フレームごとに取得し、前記符号化対象フレームの発生符号量と前記乖離度との加重和を最小化する前記フィルタ係数を、前記フィルタ係数の候補の集合から選択する選択部と、前記フィルタ係数に応じた前記発生符号量に基づいて他の前記フィルタ係数に応じた前記発生符号量の予測分布を取得し、前記他のフィルタ係数に応じた前記発生符号量の予測値を前記予測分布に基づいて取得し、前記予測値が最小化する前記他のフィルタ係数に応じた前記発生符号量に基づいて前記フィルタ係数の候補の集合を生成する集合生成部とを備えるフィルタ選択装置である。
本発明の一態様は、上記のフィルタ選択装置としてコンピュータを機能させるためのフィルタ選択プログラムである。
本発明により、高フレームレート画像から生成される低フレームレート画像の符号化効率を向上させる時間フィルタの係数を選択することが可能である。
本発明の実施形態について、図面を参照して詳細に説明する。
図1は、フィルタ選択装置1の構成例を示す図である。フィルタ選択装置1は、フィルタの係数を選択する情報処理装置である。フィルタ選択装置1は、記憶部10と、取得部11と、フィルタ12と、符号化器13と、初期化部14と、集合生成部15と、選択部16とを備える。
図1は、フィルタ選択装置1の構成例を示す図である。フィルタ選択装置1は、フィルタの係数を選択する情報処理装置である。フィルタ選択装置1は、記憶部10と、取得部11と、フィルタ12と、符号化器13と、初期化部14と、集合生成部15と、選択部16とを備える。
各機能部のうち一部又は全部は、例えば、CPU(Central Processing Unit)等のプロセッサが、記憶部に記憶されたプログラムを実行することにより実現される。各機能部のうち一部又は全部は、LSI(Large Scale Integration)やASIC(Application Specific Integrated Circuit)等のハードウェアを用いて実現されてもよい。
記憶部10は、例えばフラッシュメモリ、HDD(Hard Disk Drive)などの不揮発性の記録媒体(非一時的な記録媒体)である。記憶部10は、例えば、RAM(Random Access Memory)やレジスタなどの揮発性の記録媒体を有してもよい。記憶部10は、例えば、フィルタリング処理前の高い時間解像度の動画像のフレーム群と、データテーブルと、プログラムとを記憶する。
取得部11は、高い時間解像度の動画像のフレーム(高フレームレート画像)を、符号化対象フレームの原画像として取得する。取得部11は、高い時間解像度の動画像のフレーム群を、記憶部10に記録する。フィルタ12は、時間フィルタである。フィルタ12は、高い時間解像度の動画像のフレームを、取得部11から取得する。
フィルタ12は、選択部16によって選択された係数ベクトル及びシフト量に基づいて、高い時間解像度の動画像のフレームに対してフィルタリング処理(ダウンサンプリング)を実行する。フィルタ12は、フィルタリング処理によって生成された低い時間解像度の動画像の符号化対象フレーム(低フレームレート画像)を、符号化器13及び選択部16に出力する。
符号化器13は、低い時間解像度の動画像の符号化対象フレームに対して、動き補償予測(Motion Compensation Prediction)を用いた可逆符号化を実行する。符号化器13は、低い時間解像度の動画像の符号化対象フレームの発生符号量を表す情報を、集合生成部15及び選択部16に出力する。
初期化部14は、各フレーム(ステージ)のフィルタリング処理に関して、フィルタリング処理によるフィルタ係数が施される時間位置(フィルタリングによるダウンサンプリング時刻)のシフト量を初期化する。
集合生成部15は、低い時間解像度の動画像の符号化対象フレームの発生符号量を表す情報を、符号化器13から取得する。集合生成部15は、選択される係数ベクトルの候補(以下「係数候補ベクトル」という。)の集合を生成する。集合生成部15は、生成された係数候補ベクトルの集合を、符号化対象フレームの発生符号量が最小化するように、例えばベイズ最適化に基づく手法で最適化(更新)する。
選択部16は、フィルタリング処理前の高い時間解像度の動画像のフレーム群を、記憶部10から取得する。選択部16は、フィルタリング処理後の低い時間解像度の動画像の符号化対象フレームを、フィルタ12から取得する。選択部16は、低い時間解像度の動画像の符号化対象フレームの発生符号量を表す情報を、符号化器13から取得する。
選択部16は、フィルタリング処理前の高い時間解像度の動画像のフレームと、フィルタリング処理後の低い時間解像度の動画像の符号化対象フレームとの乖離度を、符号化対象フレームごとに算出する。選択部16は、低い時間解像度の動画像の符号化対象フレームの発生符号量と乖離度との加重和を算出する。
選択部16は、発生符号量と乖離度との加重和に基づいて、最適化された係数候補ベクトルの集合の中から、加重和を最小化する係数ベクトル(フィルタ係数)を選択する。すなわち、選択部16は、発生符号量と乖離度との加重和に基づいて、集合生成部15によって最適化された係数候補ベクトルの集合の中から、低い時間解像度の符号化対象フレームを生成するために最適な係数ベクトルを選択する。選択部16は、発生符号量と乖離度との加重和に基づいて、フィルタ係数が施される時間位置(フィルタリングによるダウンサンプリング時刻)のシフト量を選択してもよい。
このように、符号化に適した低い時間解像度の動画像の符号化対象フレームは、選択部16によって選択された係数ベクトル及びシフト量に基づく時間フィルタを用いて、高い時間解像度の動画像のフレーム群から生成される。
次に、フィルタ選択装置1の構成例の詳細を説明する。
[表記法]
以下では、表記の簡略化のため、動画像の各フレームは一次元信号として表される。(2Δ+1)タップの時間フィルタに基づいて生成された第iフレームは、式(1)のように表される。
[表記法]
以下では、表記の簡略化のため、動画像の各フレームは一次元信号として表される。(2Δ+1)タップの時間フィルタに基づいて生成された第iフレームは、式(1)のように表される。
iは、ダウンサンプリング後のフレームを指定するインデックスを表す。インデックスは、非負の整数値をとる。δtは、フィルタ12に入力される動画像のフレームの間隔を表す。各フレームは、時刻t=jδt(j=0,1,…)においてサンプリングされる。f(x,t)(x=0,…,X-1)は、第tフレームの空間位置xにおける画素値である。
wi[j]は、動き補償予測に用いられる参照フレームに対するフィルタ係数を表す。wi[j]は、式(2)の関係を満たす。
wi[j]は、動き補償予測に用いられる参照フレームに対するフィルタ係数を表す。wi[j]は、式(2)の関係を満たす。
以下、係数ベクトルは、文章中の数式ではWiのように大文字で表記される。式(1)において、左辺のWiは、フィルタ係数を要素とする係数ベクトルWi=(wi[-Δ],…,wi[Δ])を表す。piは、フィルタ係数が施される時間位置を補正するパラメータを表す。
Mは、フィルタ12から出力される動画像の符号化対象フレームのフレームレートを決定するパラメータ(ダウンサンプリング比)を表す。式(1)において、フィルタ12から出力される動画像の符号化対象フレームのフレームレートは、1/(Mδt)である。
なお、(2Δ+1≦M)が満たされている。
なお、(2Δ+1≦M)が満たされている。
式(1)に示された時間フィルタの特殊形として、フィルタ係数を一定値w[i]=1/(2Δ+1)とするフィルタを、「平均フィルタ」という。平均フィルタから出力される第iフレームは、式(3)のように表される。
N種類の係数候補ベクトルは、γn=(γn[-Δ],…,γn[Δ]),(n=0,…,N-1)と表記される。また、係数候補ベクトルから選択された係数ベクトルのフィルタ係数が施される時間位置は、P通りの時間位置から選択可能である。選択部16は、低い時間解像度の動画像の符号化対象フレームを生成するために最適な係数ベクトルを、(N×P)通りのN種類の係数候補ベクトルの中から、高い時間解像度の動画像のフレームごとに選択する。
以下では、係数候補ベクトルの集合を「辞書」という。表記の簡略化のため、N種類の係数候補ベクトルから構成される辞書を、「ΓN=(γ0,…,γN-1)」と表記する。
「フィルタ係数の最適化の規準]
選択部16は、フィルタ12から出力されたフレーム(時間フィルタによって生成されたフレーム)における発生符号量を、時間フィルタの設計における最適化の規準として用いる。発生符号量を表す情報は、動き補償予測を用いた可逆符号化を実行する符号化器13から得られる。
選択部16は、フィルタ12から出力されたフレーム(時間フィルタによって生成されたフレーム)における発生符号量を、時間フィルタの設計における最適化の規準として用いる。発生符号量を表す情報は、動き補償予測を用いた可逆符号化を実行する符号化器13から得られる。
X個の画素から構成されているフレームをK個に分割して、分割された区間ごとに動き補償フレーム間予測を行う場合について説明する。以下では、数式において文字の上に記載される記号(例えば、^)は、その文字の直前に記載される。
取得部11は、フレーム^f(x,iMδt,Wi,pi)を、サイズ(X/K)の区間B[k](k=0,1,…,K-1)に分割する。符号化器13がフレーム^f(x,(i-1)Mδt,Wi-1,pi-1)を参照フレームとして、各区間B[k](k=0,1,…,K-1)に対して動き補償(変位量di=(di[0],…,di[K-1])を実行する場合、そのフレーム内の動き補償フレーム間予測誤差は、式(4)のように表現される。
この動き補償フレーム間予測誤差を符号化とする符号化器13から得られる発生符号量Ψは、式(5)のように表される。
ここで、Rhは、符号化器13が生成するヘッダー情報の符号量を表す。Rd(di[0],…,di[K-1])は、推定変位量di[0],…,di[K-1]に関する符号量を表す。Re(ei(0,Wi,Wi-1,pi,pi-1),…,ei(X-1,Wi,Wi-1,pi,pi-1))は、動き補償フレーム間予測誤差に関する符号量を表す。
可逆符号化を実行する符号化器13から得られる発生符号量Ψが用いられているため、動き補償フレーム間予測誤差は、符号化対象フレーム及び参照フレームのみに依存する。
したがって、式(5)における発生符号量Ψは、変位量と、ヘッダ情報と、第iフレームに対する係数ベクトルWiと、第iフレームに対してフィルタ係数が施される時間位置の補正パラメータであるシフト量piと、第(i-1)フレームに対する係数ベクトルWi-1と、第(i-1)フレーム対してフィルタ係数が施される時間位置の補正パラメータであるシフト量pi-1とにより定まる。
したがって、式(5)における発生符号量Ψは、変位量と、ヘッダ情報と、第iフレームに対する係数ベクトルWiと、第iフレームに対してフィルタ係数が施される時間位置の補正パラメータであるシフト量piと、第(i-1)フレームに対する係数ベクトルWi-1と、第(i-1)フレーム対してフィルタ係数が施される時間位置の補正パラメータであるシフト量pi-1とにより定まる。
選択部16は、サンプリングされた高い時間解像度の動画像のフレーム(原画像)と、フィルタ12から出力された動画像の符号化対象フレーム(フィルタリングにより生成された動画像のフレーム)との乖離度Φとして、式(6)の値を算出する。
選択部16は、式(5)に示された発生符号量Ψと式(6)に示された乖離度との加重和として、式(7)の値を算出する。式(7)の値は、フィルタ12のフィルタ設計の評価尺度を表す。
ここで、λは、加重和の係数として、乖離度Φ[Wi,pi]に乗算される係数である。λの値は、例えば予め定められる。選択部16は、式(7)に示された評価尺度が最小化されるように、フィルタ12を設計する。発生符号量Ψ[Wi,Wi-1,pi,pi-1]に基づいてフィルタ12が設計されることで、フィルタリングにより生成された動画像の発生符号量を低減する効果が期待できる。また、乖離度Φ[Wi,pi]に基づいてフィルタ12が設計されることで、フィルタリングにより生成された動画像とサンプリングされた動画像とが乖離することを抑止する効果が期待できる。
[フィルタ係数の最適化]
選択部16は、発生符号量を最小化するフレームをフィルタ12が生成するように、式(5)の発生符号量をコスト関数とする最小化問題を解く。選択部16は、式(8)を満たす(J/M)組の係数ベクトル及びシフト量を算出する。Jは、高い時間解像度の動画像のフレーム数を表す。
選択部16は、発生符号量を最小化するフレームをフィルタ12が生成するように、式(5)の発生符号量をコスト関数とする最小化問題を解く。選択部16は、式(8)を満たす(J/M)組の係数ベクトル及びシフト量を算出する。Jは、高い時間解像度の動画像のフレーム数を表す。
(N×P)種類の係数ベクトルを係数候補ベクトルとする場合、係数ベクトル及びシフト量の組合せは、NPJ/M通りとなる。最適な係数ベクトルを選択部16が選択するためには、指数オーダの演算量が必要である。このため、係数ベクトル及びシフト量の最適な組み合わせ(W*
0,…,W*
J/M-1,p*
0,…,p*
J/M-1)を総当りで探索することは、演算量の観点から現実的ではない。
Ξ[Wi,Wi-1,pi,pi-1]がWi,pi及びWi-1,pi-1のみに依存することに着目すれば、式(8)は、単純マルコフ過程における最適化問題として定式化可能である。単純マルコフ過程における最適化問題は、動的計画法によって、多項式オーダの演算量で最適解を算出することが可能である。以下では、動的計画法を用いた解法を示す。
選択部16は、Wi,pi(i=1,…,J/M-1)に関して、式(9)のSi(Wi,pi)を定義する。
Si(Wi,pi)は、第iステージにおいて、シフト量piでシフトされたフィルタ係数Wiによって第iフレームが生成され、その状態に至る経路において最適な係数ベクトルが用いられた場合における、コストの総和を表す。ここで、Wi,piがそれぞれ固定された場合、Ξ[Wi,Wi-1,pi,pi-1]がWi-1,pi-1のみに依存することに着目すると、Si(Wi,pi)は、式(10)のような漸化式として表される。
なお、第(i-1)ステージにおいて、Si-1(Wi-1,pi-1)は、同様の漸化式を用いて算出済みである。Si-1(Wi-1,pi-1)は、Si(Wi,pi)の算出時には、参照可能な値としてレジスタ等の記憶部に記憶されている。
式(10)に示されているように、選択部16は、Ξ[Wi,Wi-1,pi,pi-1]+Si-1(Wi-1,pi-1)を最小化する係数候補ベクトル及びシフト量piを辞書ΓNから選択することによって、Si(Wi,pi)を算出する。
以下、係数ベクトルWiに対する係数候補ベクトルのインデックスをniと表記する。
選択部16は、式(10)が最小値となる係数候補ベクトルのインデックス^ni-1(ni,pi)を、niごとに記憶部10に記録する。選択部16は、式(10)が最小値となるシフト量^pi-1(ni,pi)を、niごとに記憶部10に記録する。これによって、選択部16は、式(10)が最小値となる係数候補ベクトルのインデックス^ni-1(ni,pi)と、シフト量^pi-1(ni,pi)とを、後段の処理において記憶部10から取得することができる。
選択部16は、式(10)が最小値となる係数候補ベクトルのインデックス^ni-1(ni,pi)を、niごとに記憶部10に記録する。選択部16は、式(10)が最小値となるシフト量^pi-1(ni,pi)を、niごとに記憶部10に記録する。これによって、選択部16は、式(10)が最小値となる係数候補ベクトルのインデックス^ni-1(ni,pi)と、シフト量^pi-1(ni,pi)とを、後段の処理において記憶部10から取得することができる。
式(10)の漸化式が再帰的に用いられることで、式(8)の最小化問題は、式(11)のように表される。
このように、式(10)の漸化式が再帰的に用いられる方法であれば、式(8)の最適解(W*
0,…,W*
J/M-1,p*
0,…,p*
J/M-1)は、{(NP)2J/M}通りの中から最適解を探索する問題に帰着される。これによって、選択部16は、多項式オーダの演算量で、式(11)を算出することが可能である。
ΣJ/M-1
i=1[Wi,Wi-1,pi,pi-1]の最小値が算出された場合、選択部16は、最適解(W*
0,…,W*
J/M-1,p*
0,…,p*
J/M-1)を、以下に示されるバックトラック過程によって算出する。
式(12)のSJ/M-1(wJ/M-1,pJ/M-1)を最小化する係数候補ベクトル及びシフト量の組み合わせは、(W*
J/M-1,p*
J/M-1)と表される。
W*
J/M-1を表す係数候補ベクトルのインデックスは、nJ/M-1と表される。
第(J/M-1)フレームに対する係数候補ベクトルのインデックスは、nJ/M-1と表される。第(J/M-1)フレームに対するシフト量は、pJ/M-1と表される。
第(J/M-1)フレームに対する係数候補ベクトルのインデックスは、nJ/M-1と表される。第(J/M-1)フレームに対するシフト量は、pJ/M-1と表される。
第(J/M-2)フレームに対する最適な係数候補ベクトルのインデックスは、^nJ/M-2(nJ/M-1,pJ/M-1)として、記憶部10に記憶されている。第(J/M-2)フレームに対するシフト量は、^pJ/M-2(nJ/M-1,pJ/M-1)として、記憶部10に記憶されている。
選択部16は、第(J/M-2)フレームに対する係数ベクトル及びシフト量の組み合わせを、{W*
J/M-2=γ^nJ/M-2(nJ/M-1,pJ/M-1),p*
J/M-2=^pJ/M-2(nJ/M-1,pJ/M-1)}と同定する。
選択部16は、バックトラック過程において、{W*
J/M-3=γ^nJ/M-3(nJ/M-2,pJ/M-2),p*
J/M-3=^pJ/M-3(nJ/M-2,pJ/M-2)},…,{W*
0=γ^n0(n1,p1),p*
0=^p0(n1,p1)}のように、同様の参照処理を繰り返す。このように、選択部16は、係数ベクトル及びシフト量の組み合わせを選択する。
[辞書設計]
辞書設計とは、動画像の発生符号量を最小化するように各係数候補ベクトルを設計することである。選択部16は、集合生成部15によって生成された係数候補ベクトルの辞書から、最適な係数ベクトルを選択する。辞書設計の目的関数は、最適な係数ベクトルと最適なシフト量とが選択部16によってフィルタ12に対して設定された際の発生符号量Ψの関数である。辞書設計の目的関数は、式(13)のように表される。
辞書設計とは、動画像の発生符号量を最小化するように各係数候補ベクトルを設計することである。選択部16は、集合生成部15によって生成された係数候補ベクトルの辞書から、最適な係数ベクトルを選択する。辞書設計の目的関数は、最適な係数ベクトルと最適なシフト量とが選択部16によってフィルタ12に対して設定された際の発生符号量Ψの関数である。辞書設計の目的関数は、式(13)のように表される。
辞書設計の最適化問題は、辞書設計の目的関数の探索空間における最小化問題である。
辞書Γは、N種類の係数候補ベクトルを含み、(2Δ+1)個の係数候補ベクトルを要素とする。したがって、辞書Γは、{(2Δ+1)N}次元の探索空間におけるサンプル点(観測点)に対応する。
辞書Γは、N種類の係数候補ベクトルを含み、(2Δ+1)個の係数候補ベクトルを要素とする。したがって、辞書Γは、{(2Δ+1)N}次元の探索空間におけるサンプル点(観測点)に対応する。
式(13)に示されている目的関数は、非線形であり、微分不可であり、非凸である。
このため、式(13)に示されている目的関数の最小値を解析的に求めることは不可能である。また、探索空間における目的関数の算出コストは高い。
このため、式(13)に示されている目的関数の最小値を解析的に求めることは不可能である。また、探索空間における目的関数の算出コストは高い。
フィルタ12のタップ長の増加と係数候補ベクトルの本数の増加とに応じて、探索空間は指数的に増大する。このため、グリッドサーチのような全探索に基づく手法は、演算量が膨大になるため現実的ではない。また、遺伝的アルゴリズム(Genetic Algorithm)や粒子群最適化(PSO: Particle Swarm Optimization)のようなメタヒューリスティック手法は、目的関数に対する膨大なサンプリングが前提となっている。このため、目的関数の算出コストが高い場合、メタヒューリスティック手法は現実的ではない。
そこで、集合生成部15は、ベイズ最適化に基づく手法で、辞書設計を実行する。ベイズ最適化に基づく手法は、限られたサンプル点を用いた多次元探索に適した手法である。
これは、ベイズ最適化では、ガウス過程のベイズ推定に基づいてコスト関数が推定されるためである。
これは、ベイズ最適化では、ガウス過程のベイズ推定に基づいてコスト関数が推定されるためである。
集合生成部15は、ベイズ最適化に基づいて、辞書とコスト関数との関係を推定する。
集合生成部15は、推定結果に基づいて、コスト関数を最小化する最適な辞書を同定する。
集合生成部15は、推定結果に基づいて、コスト関数を最小化する最適な辞書を同定する。
Φiは、第i番目の辞書のコスト関数を表す。以下、コスト関数の集合{Φ1,…,Φm}を、Φ1:mと表記する。Γiは、第i番目の辞書を表す。以下、辞書に関する関数値の集合{h(Γ1),…,h(Γm)},Γ1:m={Γ1,…,Γm}を、h1:mと表記する。ベイズ最適化では、式(14)に示された観測モデルが用いられる。
ここで、εiは、観測時の雑音を表す。N(0,ρ2)は、平均が0であり分散がρ2であるガウス分布を表す。hは、事前分布としてのガウス過程から得られると仮定された未知関数を表す。関数値の集合h1:mは、多次元ガウス分布N(0,K(Γ1:m))に基づいて得られる集合である。K(Γ1:m)は、(m×m)行列である。K(Γ1:m)の第(i,j)要素は、共分散関数k(Γi,Γj)である。共分散関数として、「Matern5/2カーネル関数」が用いられる。したがって、式(14)は、第i番目の辞書Γiに未知関数hの雑音が重畳された場合における、観測モデルを表す。
[係数ベクトル及びシフト量の探索点の選択処理]
ベイズ最適化では、観測値を最小化することが期待される探索点が、観測点から逐次的に選択される。集合生成部15は、コスト関数を最小化することが期待される係数ベクトル及びシフト量の組み合わせを、観測点から逐次的に選択する。
ベイズ最適化では、観測値を最小化することが期待される探索点が、観測点から逐次的に選択される。集合生成部15は、コスト関数を最小化することが期待される係数ベクトル及びシフト量の組み合わせを、観測点から逐次的に選択する。
集合生成部15は、観測値D1:m={Γ1:m,Φ1:m}のように観測値を累積する。集合生成部15は、ベイズ則に基づいて、未知関数hの事後分布を求める。集合生成部15は、求められた事後分布に基づいて、未知サンプルΓにおける観測値Φのベイズ予測分布等の予測分布を、式(15)から式(17)までのように、解析的に算出する。
ここで、k(Γ)=(k(Γ,Γ1),…,k(Γ,Γm))Tである。Φ1:m=(Φ1,…,Φm)Tである。Tは転置を表す。Iは(m×m)単位行列である。
集合生成部15は、式(15)に示されたベイズ予測分布に基づいて、獲得関数を生成する。獲得関数は、式(18)に示された、信頼下限(lower confidence bound)の獲得関数である。
ここで、βmは、探索と利用とのトレードオフを制御するハイパーパラメータを表す。
集合生成部15は、獲得関数を最小化するように、次の探索点を選択する。
集合生成部15は、獲得関数を最小化するように、次の探索点を選択する。
次に、フィルタ選択装置1の動作の例を説明する。
[辞書設計の処理と係数ベクトル及びシフト量の選択処理との流れ]
図2は、フィルタ選択装置1の動作例を示すフローチャートである。初期化部14は、各ステージ(各インデックスi)におけるシフト量を初期化する。例えば、初期化部14は、全てのステージにおけるシフト量を零値とする(ステップS101)。
[辞書設計の処理と係数ベクトル及びシフト量の選択処理との流れ]
図2は、フィルタ選択装置1の動作例を示すフローチャートである。初期化部14は、各ステージ(各インデックスi)におけるシフト量を初期化する。例えば、初期化部14は、全てのステージにおけるシフト量を零値とする(ステップS101)。
集合生成部15は、上記の[辞書設計]の説明に記載された辞書設計処理を実行する。
ベイズ最適化では、コスト関数を最小化するパラメータは、連続量として算出される。これに対して、シフト量は離散値(フレームのダウンサンプリング時刻)に制限される。このため、離散値であるシフト量は、ベイズ最適化によっては算出することができない。そこで、集合生成部15は、シフト量を固定した状態でベイズ最適化の処理を実行することによって、辞書を最適化する(ステップS102)。
ベイズ最適化では、コスト関数を最小化するパラメータは、連続量として算出される。これに対して、シフト量は離散値(フレームのダウンサンプリング時刻)に制限される。このため、離散値であるシフト量は、ベイズ最適化によっては算出することができない。そこで、集合生成部15は、シフト量を固定した状態でベイズ最適化の処理を実行することによって、辞書を最適化する(ステップS102)。
選択部16は、上記の[フィルタ係数の最適化]の説明に記載された選択処理を実行する。選択部16は、最適化された辞書(係数候補ベクトルの集合)を固定した状態で、各ステージにおける係数ベクトル及びシフト量を選択する(ステップS103)。
選択部16は、収束条件が満たされているか否かを判定する(ステップS104)。収束条件が満たされていない場合(ステップS104:NO)、選択部16は、ステップS102に処理を戻す。収束条件が満たされている場合(ステップS104:YES)、選択部16は、図2に示された処理を終了する。このようにして、選択部16は、辞書に記録された係数ベクトル及びシフト量の組み合わせを、時間フィルタのパラメータとして選択する。
図3は、辞書設計処理(ステップS102)の動作例を示すフローチャートである。集合生成部15は、K個の初期サンプル点を指定し、サンプル点ごとの観測値Dを取得する。Kは、外部から与えられるパラメータである(ステップS201)。集合生成部15は、式(15)に基づいて、観測値D1:mのベイズ予測分布を、観測値Dの予測分布として算出する(ステップS202)。式(18)に基づいて、信頼下限(観測値の予測値)を算出する(ステップS203)。集合生成部15は、信頼下限が最小値となるサンプル点を同定し、同定されたサンプル点を次の探索点と定める。すなわち、集合生成部15は、信頼下限が最小化するサンプル点の観測値を、次に観測する(ステップS204)。
集合生成部15は、収束条件が満たされているか否かを判定する。収束条件は、例えば、観測値の最小値の低減量が指定の閾値を下回るという条件である。収束条件は、例えば、指定の閾値に反復回数が到達するという条件でもよい(ステップS205)。収束条件が満たされていない場合(ステップS205:NO)、集合生成部15は、ステップS202に処理を戻す。収束条件が満たされている場合(ステップS205:YES)、集合生成部15は、図3に示された処理を終了する。
図4は、係数ベクトル及びシフト量の選択処理(ステップS103)の動作例を示す図である。また、図5は、係数ベクトル及びシフト量の選択処理の動作例を示すフローチャートである。
取得部11は、動画像として撮影された映像の信号(動画像のフレーム)と、高い時間解像度の動画像のフレーム数Jと、高い時間解像度の動画像のフレームレートとを読み込む(ステップS301)。フィルタ12は、ダウンサンプリング比Mを、選択部16から読み込む(ステップS302)。
選択部16は、ダウンサンプリング後のフレームi=1からフレームi=(J/M-1)までについて、ステップS304からステップS309までの処理を繰り返す(ステップS303)。
選択部16は、フィルタ係数Wi=Ψ0からフィルタ係数Wi=ΨN-1までについて、ステップS305からステップS309までの処理を繰り返す(ステップS304)。
選択部16は、フィルタ位置を補正するパラメータpi=0からパラメータpi=P-1までについて、ステップS306からステップS309までの処理を繰り返す(ステップS305)。
選択部16は、フィルタ位置を補正するパラメータpi=0からパラメータpi=P-1までについて、ステップS306からステップS309までの処理を繰り返す(ステップS305)。
選択部16は、フィルタ係数Wi-1=Ψ0からフィルタ係数Wi=ΨN-1までについて、ステップS307からステップS309までの処理を繰り返す(ステップS306)。選択部16は、フィルタ位置を補正するパラメータpi-1=0からパラメータpi=P-1までについて、ステップS308からステップS309までの処理を繰り返す(ステップS307)。
選択部16は、参照フレームを~f(x-di[k],(i-1)Mδt,Wi-1,pi-1)とする場合における、~f(x,iMδt,Wi,pi)に対する動き補償フレーム間予測誤差を最小化する動きベクトルを求める。動きベクトルの求め方は、選択部16に外部から与えられる。例えば、動きベクトルの求め方は、探索範囲内の全ての候補を探索する全探索法である。選択部16は、動きベクトルを用いた場合における動き補償フレーム間予測誤差を、σ2
i[Wi,Wi-1,pi,pi-1]として記憶部10に記憶する(ステップS308)。
選択部16は、σ2
i[Wi,Wi-1,pi,pi-1]+Si-1(Wi-1,pi-1)の値を算出する(i=1である場合、選択部16は、σ2
i[W1,W0,p1,p0]の値を算出する。)(ステップS309)。
選択部16は、σ2
i[Wi,Wi-1,pi,pi-1]+Si-1(Wi-1,pi-1)(Wi-1=Ψ0,…,ΨN-1,pi-1=0,…,P-1)の中での最小値を、Si(Wi,pi)として記憶部10に記録する(i=1である場合、選択部16は、σ2
1[W1,W0,p1,p0](W0=Ψ0,…,ΨN-1,p0=0,…,P-1)の中での最小値を、S1(W1,p1)として記憶部10に記録する。)(ステップS310)。
選択部16は、Si(Wi,pi)を与えるWi-1を、^Wi-1(Wi,pi)として記憶部10に記録する。選択部16は、Si(Wi,pi)を与えるpi-1を、^pi-1(Wi,pi)として記憶部10に記録する(ステップS311)。
選択部16は、SJ/M-1(WJ/M-1,pJ/M-1)を最小化する係数ベクトルWJ/M-1を、W*
J/M-1として記憶部10に記録する。選択部16は、SJ/M-1(WJ/M-1,pJ/M-1)を最小化するシフト量pJ/M-1を、p*
J/M-1として記憶部10に記録する(ステップS312)。
選択部16は、ダウンサンプリング後のフレームi=(J/M-2)からフレームi=1までについて、ステップS314からステップS315までの処理を繰り返す(ステップS313)。
選択部16は、W*
i-1=^Wi-1(Wi,pi)を、記憶部10から読み出す(ステップS314)。選択部16は、p*
i-1=^pi-1(Wi,pi)を、記憶部10から読み出す(ステップS315)。選択部16は、~f(x,iMδt,W*
i,p*
i)(x=0,…,X-1)を、ダウンサンプリング後の第iフレームとして記憶部10に記録する(ステップS316)。
以上のように、実施形態のフィルタ選択装置1は、フィルタ12と、選択部16と、集合生成部15とを備える。フィルタ12は、低い時間解像度の動画像の符号化対象フレームを、高い時間解像度の動画像のフレームから、フィルタ係数に応じて生成する。選択部16は、高い時間解像度の動画像のフレームと低い時間解像度の動画像の符号化対象フレームとの乖離度を符号化対象フレームごとに取得する。選択部16は、符号化対象フレームの発生符号量と乖離度との加重和を最小化するフィルタ係数を、フィルタ係数の候補の集合(辞書)から選択する。集合生成部15は、フィルタ係数に応じた発生符号量に基づいて、他のフィルタ係数に応じた発生符号量の予測分布を取得する。集合生成部15は、他のフィルタ係数に応じた発生符号量の予測値を、予測分布に基づいて取得する。集合生成部15は、発生符号量の予測値が最小化する他のフィルタ係数に応じた発生符号量に基づいて、フィルタ係数の候補の集合を生成する。
これによって、実施形態のフィルタ選択装置1は、高フレームレート画像から生成される低フレームレート画像の符号化効率を向上させる時間フィルタの係数を選択することが可能である。
実施形態のフィルタ選択装置1は、フィルタリング後の動画像の主観画質を保持することが可能である。実施形態のフィルタ選択装置1は、フィルタリング後の動画像の発生符号量を低減することが可能である。実施形態のフィルタ選択装置1は、動画像のフレームに関する加重和の累積値を最小化する係数ベクトルを、選択することが可能である。
実施形態のフィルタ選択装置1は、係数ベクトルの辞書の動的更新を伴う時間フィルタリングにおいて、辞書設計処理及び選択処理の最適化が可能である。
以上、この発明の実施形態について図面を参照して詳述してきたが、具体的な構成はこの実施形態に限られるものではなく、この発明の要旨を逸脱しない範囲の設計等も含まれる。
上述した実施形態におけるフィルタ選択装置をコンピュータで実現するようにしてもよい。その場合、この機能を実現するためのプログラムをコンピュータ読み取り可能な記録媒体に記録して、この記録媒体に記録されたプログラムをコンピュータシステムに読み込ませ、実行することによって実現してもよい。なお、ここでいう「コンピュータシステム」とは、OSや周辺機器等のハードウェアを含むものとする。また、「コンピュータ読み取り可能な記録媒体」とは、フレキシブルディスク、光磁気ディスク、ROM、CD-ROM等の可搬媒体、コンピュータシステムに内蔵されるハードディスク等の記憶装置のことをいう。さらに「コンピュータ読み取り可能な記録媒体」とは、インターネット等のネットワークや電話回線等の通信回線を介してプログラムを送信する場合の通信線のように、短時間の間、動的にプログラムを保持するもの、その場合のサーバやクライアントとなるコンピュータシステム内部の揮発性メモリのように、一定時間プログラムを保持しているものも含んでもよい。また上記プログラムは、前述した機能の一部を実現するためのものであってもよく、さらに前述した機能をコンピュータシステムにすでに記録されているプログラムとの組み合わせで実現できるものであってもよく、FPGA(Field Programmable Gate Array)等のプログラマブルロジックデバイスを用いて実現されるものであってもよい。
本発明は、動画像の符号化装置に適用可能である。
1…フィルタ選択装置、10…記憶部、11…取得部、12…フィルタ、13…符号化器、14…初期化部、15…集合生成部、16…選択部
Claims (3)
- フィルタ選択装置が実行するフィルタ選択方法であって、
低い時間解像度の動画像の符号化対象フレームを、高い時間解像度の動画像のフレームから、フィルタ係数に応じて生成するステップと、
前記高い時間解像度の動画像のフレームと前記低い時間解像度の動画像の符号化対象フレームとの乖離度を前記符号化対象フレームごとに取得し、前記符号化対象フレームの発生符号量と前記乖離度との加重和を最小化する前記フィルタ係数を、前記フィルタ係数の候補の集合から選択するステップと、
前記フィルタ係数に応じた前記発生符号量に基づいて他の前記フィルタ係数に応じた前記発生符号量の予測分布を取得し、前記他のフィルタ係数に応じた前記発生符号量の予測値を前記予測分布に基づいて取得し、前記予測値が最小化する前記他のフィルタ係数に応じた前記発生符号量に基づいて前記フィルタ係数の候補の集合を生成するステップと
を含むフィルタ選択方法。 - 低い時間解像度の動画像の符号化対象フレームを、高い時間解像度の動画像のフレームから、フィルタ係数に応じて生成するフィルタと、
前記高い時間解像度の動画像のフレームと前記低い時間解像度の動画像の符号化対象フレームとの乖離度を前記符号化対象フレームごとに取得し、前記符号化対象フレームの発生符号量と前記乖離度との加重和を最小化する前記フィルタ係数を、前記フィルタ係数の候補の集合から選択する選択部と、
前記フィルタ係数に応じた前記発生符号量に基づいて他の前記フィルタ係数に応じた前記発生符号量の予測分布を取得し、前記他のフィルタ係数に応じた前記発生符号量の予測値を前記予測分布に基づいて取得し、前記予測値が最小化する前記他のフィルタ係数に応じた前記発生符号量に基づいて前記フィルタ係数の候補の集合を生成する集合生成部と
を備えるフィルタ選択装置。 - 請求項2に記載のフィルタ選択装置としてコンピュータを機能させるためのフィルタ選択プログラム。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2018121358A JP2020005074A (ja) | 2018-06-26 | 2018-06-26 | フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラム |
| JP2018-121358 | 2018-06-26 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020003933A1 true WO2020003933A1 (ja) | 2020-01-02 |
Family
ID=68986493
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2019/022274 Ceased WO2020003933A1 (ja) | 2018-06-26 | 2019-06-05 | フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラム |
Country Status (2)
| Country | Link |
|---|---|
| JP (1) | JP2020005074A (ja) |
| WO (1) | WO2020003933A1 (ja) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2018006831A (ja) * | 2016-06-27 | 2018-01-11 | 日本電信電話株式会社 | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム |
| JP2018006829A (ja) * | 2016-06-27 | 2018-01-11 | 日本電信電話株式会社 | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム |
| JP2018088633A (ja) * | 2016-11-29 | 2018-06-07 | 日本電信電話株式会社 | 映像フィルタリング方法、映像フィルタリング装置及びコンピュータプログラム |
-
2018
- 2018-06-26 JP JP2018121358A patent/JP2020005074A/ja active Pending
-
2019
- 2019-06-05 WO PCT/JP2019/022274 patent/WO2020003933A1/ja not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2018006831A (ja) * | 2016-06-27 | 2018-01-11 | 日本電信電話株式会社 | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム |
| JP2018006829A (ja) * | 2016-06-27 | 2018-01-11 | 日本電信電話株式会社 | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム |
| JP2018088633A (ja) * | 2016-11-29 | 2018-06-07 | 日本電信電話株式会社 | 映像フィルタリング方法、映像フィルタリング装置及びコンピュータプログラム |
Also Published As
| Publication number | Publication date |
|---|---|
| JP2020005074A (ja) | 2020-01-09 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN115210716A (zh) | 用于多帧视频帧插值的系统和方法 | |
| CN119031147A (zh) | 基于可学习任务感知机制的视频编解码加速方法及系统 | |
| CN114513670B (zh) | 一种端到端的视频压缩方法、装置和计算机可读存储介质 | |
| CN117616753A (zh) | 使用光流的视频压缩 | |
| US8391626B2 (en) | Learning of coefficients for motion deblurring by pixel classification and constraint condition weight computation | |
| WO2023179609A1 (zh) | 一种数据处理方法及装置 | |
| JP6538619B2 (ja) | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム | |
| JP5102174B2 (ja) | フレームレート変換方法、フレームレート変換装置、フレームレート変換プログラムおよびそのプログラムを記録したコンピュータ読み取り可能な記録媒体 | |
| KR20240173786A (ko) | 영상 처리 장치 및 영상의 움직임 추정 방법 | |
| JP6626319B2 (ja) | 符号化装置、撮像装置、符号化方法、及びプログラム | |
| US20230254230A1 (en) | Processing a time-varying signal | |
| JP5118005B2 (ja) | フレームレート変換方法、フレームレート変換装置、フレームレート変換プログラムおよびそのプログラムを記録したコンピュータ読み取り可能な記録媒体 | |
| WO2020003933A1 (ja) | フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラム | |
| JP6595442B2 (ja) | 映像フィルタリング方法、映像フィルタリング装置及びコンピュータプログラム | |
| CN120390126A (zh) | 一种基于视觉分层自回归的视频生成方法及系统 | |
| CN119272898A (zh) | 数据超分模型的训练方法以及数据超分处理方法 | |
| JP6680633B2 (ja) | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム | |
| JP6611256B2 (ja) | 映像フィルタリング方法、映像フィルタリング装置及び映像フィルタリングプログラム | |
| WO2020003936A1 (ja) | フィルタ選択方法、フィルタ選択装置及びフィルタ選択プログラム | |
| JP7181492B2 (ja) | 復号装置、符号化装置、復号方法、符号化方法及びプログラム | |
| JP4695015B2 (ja) | 符号量推定方法、フレームレート推定方法、符号量推定装置、フレームレート推定装置、符号量推定プログラム、フレームレート推定プログラムおよびそれらのプログラムを記録したコンピュータ読み取り可能な記録媒体 | |
| JP2007201975A (ja) | 符号量推定方法、フレームレート推定方法、符号量推定装置、フレームレート推定装置、符号量推定プログラム、フレームレート推定プログラムおよびそれらのプログラムを記録したコンピュータ読み取り可能な記録媒体 | |
| CN119697382B (zh) | 视频序列的时域采样处理方法、装置、计算机设备、可读存储介质和程序产品 | |
| Cui et al. | Adaptive memory fusion for multi-frame optical flow estimation | |
| Liu et al. | A survey on image restoration methods based on denoising diffusion probabilistic models series models |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19826262 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19826262 Country of ref document: EP Kind code of ref document: A1 |










