WO2014155471A1 - 符号化方法および符号化装置 - Google Patents

符号化方法および符号化装置 Download PDF

Info

Publication number
WO2014155471A1
WO2014155471A1 PCT/JP2013/058487 JP2013058487W WO2014155471A1 WO 2014155471 A1 WO2014155471 A1 WO 2014155471A1 JP 2013058487 W JP2013058487 W JP 2013058487W WO 2014155471 A1 WO2014155471 A1 WO 2014155471A1
Authority
WO
WIPO (PCT)
Prior art keywords
encoding
image
unit
region
analysis
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2013/058487
Other languages
English (en)
French (fr)
Inventor
岡田 光弘
稲田 圭介
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Maxell Ltd
Original Assignee
Hitachi Maxell Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hitachi Maxell Ltd filed Critical Hitachi Maxell Ltd
Priority to PCT/JP2013/058487 priority Critical patent/WO2014155471A1/ja
Priority to US14/653,483 priority patent/US10027960B2/en
Priority to JP2015507704A priority patent/JP6084682B2/ja
Priority to CN201380067188.5A priority patent/CN104871544B/zh
Publication of WO2014155471A1 publication Critical patent/WO2014155471A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/102Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or selection affected or controlled by the adaptive coding
    • H04N19/124Quantisation
    • BPERFORMING OPERATIONS; TRANSPORTING
    • B60VEHICLES IN GENERAL
    • B60WCONJOINT CONTROL OF VEHICLE SUB-UNITS OF DIFFERENT TYPE OR DIFFERENT FUNCTION; CONTROL SYSTEMS SPECIALLY ADAPTED FOR HYBRID VEHICLES; ROAD VEHICLE DRIVE CONTROL SYSTEMS FOR PURPOSES NOT RELATED TO THE CONTROL OF A PARTICULAR SUB-UNIT
    • B60W50/00Details of control systems for road vehicle drive control not related to the control of a particular sub-unit, e.g. process diagnostic or vehicle driver interfaces
    • B60W50/08Interaction between the driver and the control system
    • B60W50/14Means for informing the driver, warning the driver or prompting a driver intervention
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/50Context or environment of the image
    • G06V20/56Context or environment of the image exterior to a vehicle by using sensors mounted on the vehicle
    • GPHYSICS
    • G08SIGNALLING
    • G08GTRAFFIC CONTROL SYSTEMS
    • G08G1/00Traffic control systems for road vehicles
    • G08G1/16Anti-collision systems
    • G08G1/167Driving aids for lane monitoring, lane changing, e.g. blind spot detection
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/136Incoming video signal characteristics or properties
    • H04N19/137Motion inside a coding unit, e.g. average field, frame or block difference
    • H04N19/139Analysis of motion vectors, e.g. their magnitude, direction, variance or reliability
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/134Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the element, parameter or criterion affecting or controlling the adaptive coding
    • H04N19/136Incoming video signal characteristics or properties
    • H04N19/14Coding unit complexity, e.g. amount of activity or edge presence estimation
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/17Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/17Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
    • H04N19/174Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a slice, e.g. a line of blocks or a group of blocks
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/10Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding
    • H04N19/169Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding
    • H04N19/17Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object
    • H04N19/176Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a block, e.g. a macroblock
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N19/00Methods or arrangements for coding, decoding, compressing or decompressing digital video signals
    • H04N19/42Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by implementation details or hardware specially adapted for video compression or decompression, e.g. dedicated software implementation
    • H04N19/436Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by implementation details or hardware specially adapted for video compression or decompression, e.g. dedicated software implementation using parallelised computational arrangements

Definitions

  • the technical field relates to image coding.
  • Patent Document 1 states that “the conventional data amount control method based on feedback control can provide high encoding efficiency by entropy encoding, but the data amount cannot be reliably kept within a certain unit in units of frames. As a solution to this problem, “Highly efficient encoding by controlling the data amount so that the predetermined amount of encoded output data is within a certain value.” Is obtained by means of means for predicting the data amount in units of sections shorter than the predetermined fixed section, and the predicting means described above.
  • a method for controlling the amount of encoded output data comprising means for accumulating the difference between the predicted data amount obtained by the measuring means and the actually encoded data amount and controlling the encoding process based on the result of the accumulation "Provide”.
  • Patent Document 2 provides “a coding control device for an image signal that requires only a small amount of hardware, assigns a code amount with optimum efficiency, and obtains a decoded image with little image quality degradation”
  • Patent Reference 2 [0010] is a problem, and as means for solving the problem, “quantization parameter initial value calculation unit, macroblock line quantization parameter calculation unit, macroblock activity calculation unit, activity average value calculation unit, An image signal encoding control device including a complexity calculation unit, and classifying macroblocks into preset classes based on activities and activity average values output from the activity calculation unit and activity average value calculation unit A classifier that outputs class information and a table corresponding to the class characteristics according to the class information.
  • a conversion table unit that selects and refers to a conversion table in which information is written and converts a macroblock line quantization parameter output from the macroblock line quantization parameter calculation unit into a quantization parameter for each macroblock; (See Patent Document 2 [0011]) and the like.
  • the present application includes a plurality of means for solving the above problems.
  • An encoding step for performing the encoding of the first region, and the size of the first region analyzed in the analysis step is variable.
  • the image quality of the transmitted video can be improved while taking into account the delay time in the video transmission.
  • Animated image compression is used in many applications.
  • applications such as a TV conference system and an in-vehicle network camera system
  • image compression and image coding for transmitting a compressed image with low delay and high image quality.
  • Technology is required.
  • the amount of code generated during encoding can be made constant.
  • a buffer delay for smoothing the generated code amount can be eliminated, so that a low delay can be realized.
  • Patent Document 1 discloses an invention for controlling the amount of data so that the amount of encoded output data for each fixed interval is within a fixed value. It is effective to some extent to suppress the buffer delay. However, since the Ford back process is performed according to the generated code amount regardless of the picture pattern, it is difficult to allocate bits suitable for the picture pattern, and the image quality may be deteriorated.
  • Patent Document 2 by determining the Q parameter using the activity average value of the macroblock and the entire screen, efficient code assignment is performed, and compression coding with less image quality degradation is possible.
  • Patent Document 2 only describes the use of the activity of the entire screen, and in particular, does not describe a corresponding method when used in a low-delay image transmission system of one frame or less.
  • the activity average value of the entire screen is obtained by adding the activities of the macroblocks, the value of the previous frame is used as the activity average value of the entire screen. For this reason, when the design is suddenly changed, a large amount or a small amount of generated code may occur, and it is conceivable that the generated code amount cannot be accurately estimated.
  • H.264 is used for image coding.
  • An example using H.264 will be described.
  • FIGS. 12 shows an example of a surveillance camera system
  • FIG. 13 shows a video conference system
  • FIG. 14 shows an example of an in-vehicle camera system.
  • 1201, 1202, and 1203 are monitoring cameras installed at points A, B, and C, respectively
  • 1204 is a monitoring center that receives images captured by the monitoring cameras 1201, 1202, and 1203, and 1205 is an Internet line A wide area network (WAN). Images captured by the monitoring cameras 1201 to 1203 can be displayed on a monitor or the like in the monitoring center 1204 via the WAN 1205.
  • FIG. 12 shows an example in which there are three surveillance cameras, the number of in-vehicle cameras may be two or less or four or more.
  • the image encoding apparatus is mounted on, for example, the monitoring cameras 1201 to 1203.
  • the image encoding device performs an encoding process, which will be described later, on an input image input via the lenses of the monitoring cameras 1201 to 1203, and the encoded input image is output to the WAN 1205.
  • reference numerals 1301, 1302, and 1303 denote video conference systems installed at points A, B, and C, respectively, and 1304 denotes a WAN such as an Internet line. Images captured by the cameras of the video conference systems 1301 to 1303 can be displayed on monitors of the video conference systems 1301 to 1303 via the WAN 1304. Although FIG. 13 shows an example in which there are three video conference systems, the number of video conference systems may be two or four or more.
  • the image encoding apparatus is mounted on, for example, cameras of video conference systems 1301 to 1303.
  • the image encoding device performs an encoding process, which will be described later, on an input image input through the camera lens of the video conference systems 1301 to 1303, and the encoded input image is output to the WAN 1305.
  • reference numeral 1401 denotes an automobile
  • 1402 and 1403 denote in-vehicle cameras mounted on the automobile 1401
  • 1404 denotes a monitor for displaying images captured by the in-vehicle cameras 1402 and 1403
  • 1405 denotes a local area network (LAN) in the automobile 1401. It is. Images captured by the in-vehicle cameras 1402 and 1403 can be displayed on the monitor 1404 via the LAN 1405.
  • FIG. 14 shows an example in which two in-vehicle cameras are mounted, the number of in-vehicle cameras may be one or three or more.
  • the image encoding apparatus is mounted on, for example, on-vehicle cameras 1402 and 1403.
  • the image encoding device performs an encoding process, which will be described later, on an input image input via the lenses of the in-vehicle cameras 1402 and 1403, and the encoded input image is output to the LAN 1405.
  • FIG. 1 is an example of a configuration diagram of an image encoding device.
  • the image encoding apparatus 100 includes an input image writing unit 101, an input image complexity calculating unit 102, an input image memory 103, an encoding unit image reading unit 104, an encoding unit 105, an encoding memory 106, and an encoding unit complexity calculation.
  • QP quantization parameter
  • the input image writing unit 101 performs a process of writing the input images input in the raster scan order in the input image memory 103.
  • the input image complexity calculator 102 calculates the complexity using the input image before being written to the memory, and outputs the input image complexity.
  • the complexity is an index indicating the difficulty of the pattern of the input image in the area corresponding to the target delay time, and is given by, for example, the variance value var described in (Equation 1).
  • N represents the number of pixels in the horizontal direction to be calculated
  • M represents the number of pixels in the vertical direction to be calculated
  • x (ij) is a pixel value within the range of N ⁇ M pixels
  • X is N ⁇ M pixels. Is an average value of the pixel values within the range.
  • N and M are determined from the set target delay time.
  • the target delay time is the processing time of the encoding process from when the image is input until the stream is output, and is controlled so that the generated code amount in this delay time unit is below a certain level.
  • the target delay time is assumed to be about 1 to 50 ms for an in-vehicle network and about 100 ms for a video conference system, but the target delay time required may change depending on the situation.
  • the input image memory 103 temporarily stores the input images input in the raster scan order, and is in a coding unit (in the case of H.264, a 16 pixel ⁇ 16 pixel macroblock (hereinafter referred to as “MB”)).
  • MB 16 pixel ⁇ 16 pixel macroblock
  • This memory may be an external memory such as SDRAM or an internal memory such as SRAM.
  • the encoding unit reading unit 104 is a block that reads an MB image from the input image memory 103.
  • the MB image read by the encoding unit reading unit 104 is supplied to the encoding unit 105 and the encoding unit complexity calculation unit 107.
  • the encoding unit complexity calculation unit 107 calculates the complexity for each MB using the MB image, and outputs the encoding unit complexity. This is calculated using the same dispersion formula (Formula 1) as the input image complexity calculator 102.
  • H In the case of H.264, since MB is 16 pixels ⁇ 16 pixels, N and M are both 16.
  • the QP value calculation unit 108 outputs a QP value for each MB using the input image complexity, the encoding unit complexity, and the generated code amount generated when the encoding unit 105 actually encodes.
  • the QP indicates a quantization parameter, that is, a quantization parameter, and the calculation method of the QP value will be described later with a specific example.
  • the encoding unit 105 performs an encoding process using the MB image output from the encoding unit image reading unit 104 and the QP value output for each MB from the QP value calculation unit 108, and generates a stream.
  • the encoding memory 106 is a memory for storing a reproduced image to be used for the prediction process, and may be either SDRAM or SRAM like the input image memory. Although the input image memory 103 and the encoding memory 106 are illustrated separately in FIG. 1, it is not necessary to separate them, and one SDRAM may be used.
  • control unit 109 executes each processing block (input image writing unit 101, input image complexity calculation unit 102, encoding unit image reading unit 104, encoding unit 105, This block controls the coding unit complexity calculation unit 107 and the QP value calculation unit 108).
  • processing block input image writing unit 101, input image complexity calculation unit 102, encoding unit image reading unit 104, encoding unit 105, This block controls the coding unit complexity calculation unit 107 and the QP value calculation unit 108.
  • the control method will be described later with a specific example.
  • a configuration including the input image complexity calculation unit 102 and the coding unit complexity calculation unit 107 is also simply referred to as an analysis unit.
  • the encoding unit 105 includes a prediction unit 201, a frequency conversion / quantization unit 202, an encoding unit 203, and an inverse frequency conversion / inverse quantization unit 204.
  • the prediction unit 201 uses the MB image as an input, selects either the in-screen prediction or the inter-frame prediction, whichever is more efficient, and creates a predicted image. Thereafter, the generated predicted image and an error image obtained by subtracting the predicted image from the current image are output.
  • In-screen prediction is a method of creating a prediction image using a reproduction image of an adjacent MB stored in the encoding memory 106, and inter-frame prediction is performed using a reproduction image of a past frame stored in the encoding memory 106. This is a method of creating a predicted image.
  • the frequency transform / quantization unit 202 performs frequency transform on the error image, and then outputs a quantized coefficient obtained by quantizing the transform coefficient of each frequency component based on the quantization parameter given from the QP value calculating unit. .
  • the encoding unit 203 encodes the quantized coefficient output from the frequency transform / quantization unit 202 and outputs a stream. Further, the generated code amount used by the QP value calculation unit 108 is output.
  • the inverse frequency transform / inverse quantization unit 204 inversely quantizes the quantized coefficient to return it to the transform coefficient of each frequency component, and then performs an inverse frequency transform to generate an error image. Thereafter, a reproduction image is created by adding the prediction image output from the prediction unit 201 and stored in the encoding memory 106.
  • the precondition is that an image of 1280 pixels ⁇ 720 pixels and 60 fps (frame per second) is encoded with a target delay time of 3.33 ms.
  • FIG. 3 is a diagram in which an image of one frame (16.666 ms per frame because it is 60 fps) is divided into regions (region 1 to region 5) every target delay time 3.33 ms. All the areas have the same size, and the square blocks in the areas indicate MBs. Numbers described in the MB are assigned numbers in the processing order. In one area, the number of horizontal MBs is 80, the number of vertical MBs is 9, and the total number of MBs is 720. For each area, processing is performed so as to improve the image quality while keeping the generated code amount constant.
  • each processing block (input image writing unit 101, input image complexity calculating unit 102, encoding unit image reading unit 104, encoding unit 105, encoding unit complexity calculating unit 107, QP
  • the timing diagram which showed the processing timing of the value calculation part 108, the prediction part 201, the frequency conversion and quantization part 202, and the encoding part 203) is shown.
  • the horizontal axis indicates time, and the vertical axis indicates each processing block, so that it can be understood which region or MB is processed at which timing for each processing block.
  • the control of the processing timing is performed by the control unit 109 shown in FIG.
  • Each process of the encoding unit 105 is a pipeline process for each MB as shown in FIG.
  • pipeline processing is a technique for performing high-speed processing by dividing the encoding processing for each MB into a plurality of stages (stages) and processing the processes in each stage in parallel.
  • the input image writing process of the input image writing unit 101 and the input image complexity calculating process of the input image complexity calculating unit 102 are performed in parallel.
  • the encoding unit image reading unit 104 performs the encoding unit image reading process and outputs the encoded unit image to the encoding unit 105 and the encoding unit complexity calculation unit 107.
  • the encoding unit complexity calculation unit 107 calculates the encoding unit complexity.
  • the QP calculation process of the QP value calculation unit 108 and the prediction process of the prediction unit 201 are performed in parallel. Further, the frequency conversion / quantization unit 202 performs frequency conversion / quantization processing using the QP value calculated in the previous QP value calculation processing. Finally, the encoding unit 203 performs an encoding process and outputs a stream.
  • the target delay time can be changed by changing the region size corresponding to the target delay time and changing the start timing of the encoding process.
  • FIG. 5 is a diagram showing the internal details of the QP value calculation unit 108, and includes a base QP calculation unit 501, an MB (macroblock) QP calculation unit 502, and a QP calculation unit 503.
  • the base QP calculation unit 501 is a process executed only when straddling regions (region 1 to region 5), and outputs the base QP of the region to be processed next using the input image complexity and the code amount.
  • This base QP is given by (Equation 2) below.
  • QP ave is the average QP value of the previous region
  • bitrate is the generated code amount of the previous region
  • target_bitrate is the target code amount of the next region
  • is a coefficient
  • var next is the input image complexity of the next region
  • Var pre indicate the input image complexity of the previous area.
  • the MBQP calculation unit 502 is a process executed for each MB, and outputs MBQP from the encoding unit complexity and the input image complexity.
  • This MBQP is given by (Equation 3).
  • represents a coefficient
  • represents a limiter value
  • this process increases the MBQP of a complex picture with a large amount of generated code and reduces the MBQP of a flat picture with a small code amount to be generated. It also has the effect of smoothing the generated code amount for each.
  • the QP calculation unit 503 calculates the QP value using the formula described in (Formula 4).
  • the generated code amount in the base QP calculation unit 501 of the QP value calculation unit 108 is made constant, and the pattern in the MBQP calculation unit 502 of the QP value calculation unit 108 is determined.
  • the QP value By controlling the QP value, it is possible to achieve high image quality while keeping the generated code amount for each target delay time constant.
  • the configuration of the first embodiment is also effective when changing the target delay time during encoding or for each application.
  • the encoding unit complexity for each encoding unit in parallel with the calculation of the input image complexity in the input image complexity calculation unit 102 the encoding unit complexity for the number of MBs in the target delay time region The degree needs to be recorded in a memory (in the example of the first embodiment, a memory of 720 MB is necessary).
  • the pipeline processing delay (first embodiment) is calculated.
  • it since it is used in the next stage, it is sufficient to have a memory of 1 MB), and it is sufficient to have a fixed memory regardless of the target delay time. This is particularly effective when encoding a large image size such as a 4k8k size because a small fixed amount of memory is sufficient.
  • this configuration can be realized without changing the pipeline processing during encoding, and there is no delay in encoding processing due to the introduction of this processing.
  • the target delay time of one frame or less has been described as an example.
  • the target delay time may be set as late as the memory capacity permits.
  • the bit distribution can be performed after analyzing three input images, so that high image quality can be realized.
  • H.264 An example of H.264 is given, but it may be a moving image coding method (MPEG2, next-generation moving image coding method HEVC (H.265), etc.) having a parameter capable of changing the image quality for each coding unit. For example, the same effect can be obtained by using this configuration.
  • MPEG2 next-generation moving image coding method HEVC (H.265), etc.
  • any index indicating the complexity of the image can be used, such as the sum of difference values from adjacent pixels and the sum of edge detection filters (Sobel filter, Laplacian filter, etc.). It is not limited to the variance value.
  • the specific QP value determination method of the QP value calculation unit 108 has been described with reference to FIG. 5, but is not limited to this processing. It suffices if at least the input image complexity, coding unit complexity, and generated code amount can be used to achieve high image quality while keeping the generated code amount constant according to the target delay time.
  • the area to be input to the target delay time The complexity may be calculated for each region smaller than the size of.
  • the input image complexity is calculated for the area of 720 MB, but it is assumed that the input image complexity is calculated for each 1/3 240 MB area.
  • the code amount correction QP term shown in (Equation 4 ′) can be added every 240 MB.
  • the code amount correction QP is a value calculated every 240 MB, and is calculated based on the code amount generated immediately before 240 MB.
  • the code amount correction QP value is set to a positive value, and conversely, when the generated code amount is smaller than the code amount desired to be constant, By making the code amount QP value a negative value, it is possible to increase the accuracy of making the generated code amount constant in the target delay time.
  • the camera system in FIG. 6 includes an image transmission device 1000 and an image reception device 1100.
  • the image transmission apparatus 1000 is an in-vehicle camera, for example, and includes an imaging unit 1001 that converts light into a digital image, and an image encoding unit 1002 that encodes a digital image output from the imaging unit 1001 and generates a stream.
  • the network IF 1003 is configured to packetize the encoded stream and output it on the network.
  • the image receiving apparatus 1100 is, for example, a car navigation system, and receives a packet transmitted from the video transmitting apparatus 1000 and converts it into a stream, and generates a reproduction image by decoding the stream output from the network IF 1101
  • a voice output unit 1105 is provided to output a voice and notify the driver when the result shows a dangerous state.
  • FIG. 7 The configuration diagram of the image encoding device 1002 capable of improving the performance of the image recognition processing when performing the image recognition processing on the reproduced image is shown in FIG. 7 taking the configuration of the in-vehicle network camera system of FIG. 6 as an example.
  • the description of the components having the same functions as those already described with reference to FIG. 1 is omitted.
  • the encoding unit feature amount extraction unit 110 extracts an image feature amount from the MB image, and outputs the encoding unit image feature amount.
  • a configuration including the input image complexity calculation unit 102, the coding unit complexity calculation unit 107, and the coding unit feature amount extraction unit is also simply referred to as an analysis unit.
  • FIG. 8 shows the internal configuration of the QP value calculation unit 111. Description of portions having the same functions as those of the QP value calculation unit 108 in FIG. 5 having the same reference numerals shown in FIG.
  • a difference from FIG. 5 is a feature QP calculation unit 504 and a QP calculation unit 505.
  • the feature QP calculation unit 504 outputs a feature QP according to the size of the feature amount from the encoded unit image feature amount.
  • the QP calculation unit 505 calculates a QP value by an expression given by (Expression 5) in which not only the base QP and MBQP but also the feature QP is added.
  • the coding unit feature amount extraction unit 110 extracts white line feature amounts necessary for white line recognition. Specifically, the difference value between the adjacent pixels is calculated, and when there are consecutive difference values of the same step on the straight line, the white line feature amount (0, 1, 2) that increases the feature amount is obtained.
  • the coding unit image feature amount is output as follows: the higher the value, the higher the possibility of a white line). Further, the feature QP calculation unit 504 determines the feature quantity QP from the three levels of white line feature quantities based on the table of FIG.
  • the feature quantity QP may be determined using a mathematical expression such as a linear function or a Log function.
  • the example which used the difference value with an adjacent pixel was described for the feature-value extraction, the result of having prepared the image of the basic pattern which can search a designated picture beforehand, and performing pattern matching with the basic pattern Any index value can be used as long as it is possible to determine a target whose image quality is to be improved, such as determining the feature quantity QP using (similarity).
  • the feature amount may be calculated in an area calculated by the input image complexity calculation unit 102. .
  • FIG. 10 shows an example in which the target delay time is changed according to the application, taking the in-vehicle network camera system of FIG. 6 as an example.
  • Applications include rear obstacle detection during parking, white line departure warning during high-speed driving, and sign recognition during urban driving.
  • the application to be used changes depending on the speed of the car. Therefore, the target delay time may be changed according to the speed of the car.
  • ⁇ Obstacle detection during parking is assumed to be used at a speed of 20 km / h or less. Since the speed is low, there is little change in the image for each frame, so a low delay of 10 ms or less is not necessary, and the target delay time is 33.3 ms. Therefore, the image quality can be improved using the analysis result for one frame.
  • the white line departure warning during high speed driving is assumed to be used at a speed of 100 km / h or more. Therefore, it is considered that the image changes greatly every frame. If the delay time is large, even if image recognition processing is performed and a dangerous state is detected, it is important that the target delay time is as short as 1 ms because there is a possibility that an accident has already occurred.
  • the sign recognition when traveling in urban areas is a moderate speed of 40 km to 80 km per hour, so it does not go up to 1 ms, but it is necessary to shorten the target delay time to some extent, so it is 10 ms.
  • the image encoding device when the image encoding device is applied to the in-vehicle network camera system, the image feature focused by the image recognition processing algorithm of the image recognition unit 1104 of the image receiving device 1100 is the encoding unit feature of the image transmitting device 1000.
  • the extraction unit 110 simply extracts the feature amount, and the QP value calculation unit 111 lowers the QP value of the corresponding MB. It is possible to realize an image encoding device capable of improving the performance of recognition processing.
  • the imaging unit 1001 may be replaced with a storage device such as a recorder.
  • the transmission is performed while considering the delay in image transmission by setting the target delay time according to the application to be used. It becomes possible to improve the image quality.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Automation & Control Theory (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Human Computer Interaction (AREA)
  • Transportation (AREA)
  • Mechanical Engineering (AREA)
  • Theoretical Computer Science (AREA)
  • Compression Or Coding Systems Of Tv Signals (AREA)

Abstract

 従来技術においては、映像伝送における遅延時間を考慮しつつ、伝送する映像の画質を向上させることについては考慮されていない。 画像情報の符号化方法において、入力される画像を解析する解析ステップと、解析ステップの解析結果を用いて量子化パラメータを算出する量子化パラメータ算出ステップと、量子化パラメータ算出ステップで算出した量子化パラメータを用いて入力される画像のエンコードを行うエンコードステップと、を有し、解析ステップで解析する第1の領域の大きさが可変であること特徴とする。

Description

符号化方法および符号化装置
 技術分野は、画像符号化に関する。
 特許文献1には、「従来のフィードバック制御によるデータ量制御方式では、エントロピー符号化によって高い符号化効率が得られるが、データ量をフレーム単位などで確実に一定以内にすることができないから蓄積系メディアには適用し難いという点」等を課題とし、その解決手段として「予め定められた一定区間毎の符号化出力データ量が一定値以内になるようにデータ量を制御して高能率符号化が行われるようにした符号化出力データ量の制御方式であって、前記した予め定められた一定区間よりも短かい区間を単位にしてデータ量を予測する手段と、前記した予測手段によって得られる予測データ量に基づいて前記した予め定められた一定区間における予測データ量の合計が一定になるように符号化処理を制御する手段と、前記の予測手段によって得た予測データ量と実際に符号化されたデータ量との差を累積し、前記した累積の結果に基づいて符号化処理を制御する手段とからなる符号化出力データ量の制御方式を提供する」ことが記載されている。
 また、特許文献2には、「ハードウェアが小規模で済み、最適な効率で符号量割り当てを行い、画質劣化の少ない復号画像が得られる画像信号の符号化制御装置を提供する点」(特許文献2[0010]参照)を課題とし、その解決手段として「量子化パラメータ初期値演算部と、マクロブロックライン量子化パラメータ演算部と、マクロブロックのアクティビティ計算部と、アクティビティ平均値計算部と、複雑度演算部とを備える画像信号の符号化制御装置であって、アクティビティ計算部とアクティビティ平均値計算部とから出力されるアクティビティとアクティビティ平均値とに基づき、マクロブロックを予め設定したクラスに分類してクラス情報を出力するクラス分け部と、クラス情報に従って、クラスの特性に対応するテーブル情報が書き込まれた変換テーブルを選択して参照し、マクロブロックライン量子化パラメータ演算部から出力されるマクロブロックライン量子化パラメータを、マクロブロック毎の量子化パラメータに変換する変換テーブル部とを備えること」(特許文献2[0011]参照)等が記載されている。
特開平02-194734号公報 特開2000-270323号公報
 しかし、いずれの特許文献においても、映像伝送における遅延時間を考慮しつつ、伝送する映像の画質を向上させることについては考慮されていない。
 上記課題を解決するために、例えば特許請求の範囲に記載の構成を採用する。
  本願は上記課題を解決する手段を複数含んでいるが、その一例を挙げるならば、
 入力される画像を解析する解析ステップと、解析ステップの解析結果を用いて量子化パラメータを算出する量子化パラメータ算出ステップと、量子化パラメータ算出ステップで算出した量子化パラメータを用いて入力される画像のエンコードを行うエンコードステップと、を有し、解析ステップで解析する第1の領域の大きさが可変であることを特徴とする。
 上記手段によれば、映像伝送における遅延時間を考慮しつつ、伝送する映像の画質を向上させることができる。
画像符号化装置の構成図の例である。 エンコード部の構成図の例である。 入力画像複雑度を計算する領域の例である。 各処理ブロックの処理タイミングの一例を示すタイミング図である。 QP値計算部の構成図の例である。 車載カメラシステムの構成図の例である。 画像符号化装置の構成図の例である。 QP値計算部の構成図の例である。 特徴QPの変換テーブルの例である。 目標遅延時間の例である。 画像符号化装置の構成図の例である。 監視カメラシステムの一例である。 テレビ会議システムの一例である。 車載カメラシステムの一例である。
 動画像圧縮は多くのアプリケーションで利用されている。特に、TV会議システムや車載ネットワークカメラシステム等の用途においては、画像圧縮を利用して低遅延で画像を伝送したいというニーズがあり、低遅延かつ高画質な圧縮画像を伝送するための画像符号化技術が必要となる。
 低遅延で圧縮画像を伝送する方法として、符号化時に発生する符号量を一定にすることがあげられる。発生符号量を一定にすることで、発生した符号量を平滑化するためのバッファ遅延が無くすことができるので、低遅延化が実現できる。
 特許文献1には、一定区間ごとの符号化出力データ量が一定値以内になるようにデータ量を制御する発明が開示されており、このように符号量の変動をできるだけ抑えて、一定区間ごとの符号量が出来るだけ均一にすることは、バッファ遅延の抑制にある程度有効である。しかしながら、画像の絵柄とは関係なく、発生する符号量に応じてフォードバック処理するため画像の絵柄に適したビット配分が難しく、画質が劣化することが考えられる。
 一方、特許文献2の技術的思想では、マクロブロック及び1画面全体のアクティビティ平均値を用いてQパラメータを決めることで、効率の良い符号割り当てを行ない、画質劣化の少ない圧縮符号化を可能としている。しかしながら、特許文献2には1画面全体のアクティビティを用いる記載しかなく、特に1フレーム以下の低遅延画像伝送システムで使用する際の対応方法についての記載はない。また、マクロブロックのアクティビティを足し合わせて1画面全体のアクティビティ平均値を求める構成のため、1画面全体のアクティビティ平均値は、前フレームの値を用いることとなる。そのため、急激に絵柄が変わった場合に、発生符合量が多く発生するまたは少なく発生することがあり、発生符号量を精度よく推測できないことが考えられる。
 以下、低遅延かつ高画質な圧縮画像を伝送する実施例を、図面を用いて説明する。
 本実施例では、画像符号化を行う画像符号化装置の例を説明する。本実施例では、画像符号化にH.264を用いた例について述べる。
 まずは、図12~14を用いて本実施例の画像符号化装置が適用される画像伝送システムについて説明する。図12は監視カメラシステム、図13はテレビ会議システム、図14は車載カメラシステムの一例を示す図である。
 図12において、1201、1202、1203はそれぞれA地点、B地点、C地点に設置された監視カメラ、1204は監視カメラ1201、1202、1203で撮像された画像を受信する監視センター、1205はインターネット回線等のワイドエリアネットワーク(WAN)である。監視カメラ1201~1203で撮像された映像はWAN1205を介して監視センター1204内のモニタ等に表示することが可能である。図12においては監視カメラが3つの場合の例を示しているが、車載カメラの数は2つ以下であっても4つ以上であってもよい。
 本実施例の画像符号化装置は、例えば監視カメラ1201~1203に搭載される。画像符号化装置は、監視カメラ1201~1203のレンズを介して入力された入力画像に対して、後述する符号化処理を行い、符号化処理された入力画像はWAN1205へ出力される。
 図13において、1301、1302、1303はそれぞれA地点、B地点、C地点に設置されたテレビ会議システム、1304はインターネット回線等のWANである。テレビ会議システム1301~1303のカメラで撮像された映像はWAN1304を介してテレビ会議システム1301~1303のモニタ等に表示することが可能である。図13においてはテレビ会議システムが3つの場合の例を示しているが、テレビ会議システムの数は2つであっても4つ以上であってもよい。
 本実施例の画像符号化装置は、例えばテレビ会議システム1301~1303のカメラに搭載される。画像符号化装置は、テレビ会議システム1301~1303のカメラのレンズを介して入力された入力画像に対して、後述する符号化処理を行い、符号化処理された入力画像はWAN1305へ出力される。
 図14において、1401は自動車、1402、1403は自動車1401に搭載される車載カメラ、1404は車載カメラ1402、1403で撮像された映像を表示するモニタ、1405は自動車1401内のローカルエリアネットワーク(LAN)である。車載カメラ1402、1403で撮像された映像はLAN1405を介してモニタ1404に表示することが可能である。図14においては車載カメラが2つ搭載された例を示しているが、車載カメラの数は1つであっても3つ以上であってもよい。
 本実施例の画像符号化装置は、例えば車載カメラ1402、1403に搭載される。画像符号化装置は、車載カメラ1402、1403のレンズを介して入力された入力画像に対して、後述する符号化処理を行い、符号化処理された入力画像はLAN1405へ出力される。
 次に、本実施例の画像符号化装置について説明する。図1は、画像符号化装置の構成図の例である。画像符号化装置100は、入力画像書込み部101、入力画像複雑度計算部102、入力画像用メモリ103、符号化単位画像読込み部104、エンコード部105、エンコード用メモリ106、符号化単位複雑度計算部107、QP(量子化パラメータ)値計算部108、制御部109から構成される。
 入力画像書込み部101は、ラスタスキャン順に入力される入力画像を入力画像用メモリ103に書き込む処理を行う。
 入力画像複雑度計算部102は、メモリに書き込む前の入力画像を用いて複雑度を計算し、入力画像複雑度を出力する。ここで、複雑度とは、目標遅延時間分の領域の入力画像の絵柄の難易度を示す指標であり、例えば(式1)に記載する分散値varで与えられる。
Figure JPOXMLDOC01-appb-I000001
 ここで、Nは計算する横方向の画素数、Mは計算する縦方向の画素数を表しており、x (i.j)は上記N×M画素の範囲内の画素値、XはN×M画素の範囲内の画素値の平均値である。このNとMについては、設定された目標遅延時間より決定する。
 また、目標遅延時間とは、画像が入力されてから、ストリームが出力されるまでのエンコード処理の処理時間であり、この遅延時間単位での発生符合量が一定以下になるように制御する。目標遅延時間単位の発生符号量を一定以下にすることで、画像伝送システムを安定動作させるために必要な受信装置側のバッファリング時間を計算することができる。そのため、画像伝送システムの画像伝送時の伝送時間を保障することが可能となる。目標遅延時間は、例えば車載ネットワークであれば1ms~50ms程度、テレビ会議システムであれば100ms秒程度が想定されるが、状況に応じて要求される目標遅延時間が変化することも考えられる。
 入力画像用メモリ103は、ラスタスキャン順に入力された入力画像を一旦蓄積し、符号化単位(H.264の場合は16画素×16画素のマクロブロック(以下、「MB」と示す)。)のMB画像を連続して読み出すために使用するメモリである。このメモリはSDRAMのような外部メモリでも良いし、SRAMのような内部メモリでも良い。
 符号化単位読出し部104は、入力画像用メモリ103からMB画像を読み出すブロックである。符号化単位読出し部104で読み出されたMB画像は、エンコード部105と符号化単位複雑度計算部107へ供給される。
 符号化単位複雑度計算部107は、MB画像を用いてMB毎の複雑度を計算し、符号化単位複雑度を出力する。これは、入力画像複雑度計算部102と同一の分散式(式1)を用いて計算する。ここでH.264の場合は、MBが16画素×16画素なため、N,Mともに16になる。
 QP値計算部108は、前記入力画像複雑度と、前記符号化単位複雑度と、エンコード部105で実際にエンコードした時に発生した発生符号量を用いて、MB毎のQP値を出力する。QPとは、quantization parameter、すなわち量子化パラメータを示すものであり、QP値の算出方法については、後ほど具体例を挙げて説明する。
 エンコード部105は符号化単位画像読込み部104から出力されるMB画像とQP値計算部108からMB毎に出力されるQP値を用いてエンコード処理を行い、ストリームを生成する。
 エンコード用メモリ106は、予測処理に使用するための再生画像を蓄えておくメモリであり、入力画像用メモリと同様にSDRAM、SRAMどちらでも良い。入力画像用メモリ103とエンコード用メモリ106は、図1では分けて記載しているが分ける必要はなく、一つのSDRAMを使用するとしても良い。
 制御部109は、設定された目標遅延時間に基づいて、図1に記載の各処理ブロック(入力画像書込み部101、入力画像複雑度計算部102、符号化単位画像読込み部104、エンコード部105、符号化単位複雑度計算部107、QP値計算部108)を制御するブロックである。制御方法については後ほど具体例を挙げて説明する。
 なお、入力画像複雑度計算部102と符号化単位複雑度計算部107とを含む構成を単に解析部ともいう。
 次に、図2を用いてエンコード部105の詳細を説明する。エンコード部105は、予測部201と周波数変換・量子化部202、符号化部203、逆周波数変換・逆量子化部204から構成される。
 まず、予測部201ではMB画像を入力として、画面内予測、またはフレーム間予測のどちらか効率の良い方を選択し予測画像を作成する。その後、前記生成した予測画像と、現画像から前記予測画像を引き算した誤差画像を出力する。画面内予測は、エンコード用メモリ106で蓄えた、隣接MBの再生画像を用いて予測画像を作成する方法であり、フレーム間予測は、エンコード用メモリ106で蓄えた過去フレームの再生画像を用いて予測画像を作成する方法である。
 周波数変換・量子化部202は、誤差画像に周波数変換を施した後、QP値計算部から与えられた量子化パラメータに基づいて、各周波数成分の変換係数を量子化した量子化係数を出力する。
 符号化部203は、周波数変換・量子化部202から出力された量子化係数を符号化処理してストリームを出力する。また、QP値計算部108で使用する発生符号量を出力する。
 逆周波数変換・逆量子化部204は、量子化係数を逆量子化して各周波数成分の変換係数に戻した後、逆周波数変換して誤差画像を生成する。その後に予測部201から出力された予測画像と足し合わせて再生画像を作成し、エンコード用メモリ106に蓄える。
 次に制御部109の制御例を説明する。前提条件は、1280画素×720画素、60fps(frame per second)の画像を目標遅延時間3.33msでエンコード処理するとする。
 図3は、1フレームの画像(60fpsなので1フレームあたり16.666ms)を目標遅延時間3.33ms毎の領域(領域1~領域5)に分けた図である。すべての領域とも同じ大きさであり、領域中の四角のブロックはMBを示している。MB中に記載したナンバーは処理順番に番号を割り当てている。1つの領域に横のMB数は80個、縦のMB数は9個、合計MB数は720個となる。この領域毎に、発生符号量を一定にしつつ高画質化するように処理を行う。
 このように、目標遅延時間に基づいて領域の大きさを変化させることにより、目標遅延時間に応じた高画質化処理が可能となる。
 図4に図1、2に記載の各処理ブロック(入力画像書込み部101、入力画像複雑度計算部102、符号化単位画像読込み部104、エンコード部105、符号化単位複雑度計算部107、QP値計算部108、予測部201、周波数変換・量子化部202、符号化部203)の処理タイミングを示したタイミング図を示す。横軸は時間、縦軸は各処理ブロックを示しており、処理ブロック毎にどのタイミングでどの領域またはMBを処理しているかが分かるようになっている。この処理タイミングの制御は、図1に記載の制御部109が行う。エンコード部105の各処理は図4に記載のようにMB毎のパイプライン処理となる。ここで、パイプライン処理とは、MB毎の符号化処理を複数の段階(ステージ)に分割し、各ステージの処理を並列に処理することであり、高速処理を行なうための手法である。
 まず、入力画像書込み部101の入力画像書込み処理と、入力画像複雑度計算部102の入力画像複雑度計算処理とが並行して行われる。領域1の入力画像がすべて入力画像用メモリ103に書き込み終わったら、符号化単位画像読込み部104の符号化単位画像読込み処理を行いエンコード部105と符号化単位複雑度計算部107に出力する。ここで、読込みと並行して、符号化単位複雑度計算部107の符号化単位複雑度計算を行う。
 次に、QP値計算部108のQP計算処理と、予測部201の予測処理を並行して行う。さらに、周波数変換・量子化部202では、一つ前のQP値計算処理で計算したQP値を使用し、周波数変換・量子化処理を行う。最後に、符号化部203の符号化処理を行い、ストリームを出力する。
 ここで、実際のエンコード処理による遅延時間は、入力画像が入力されてから、符号化部203からストリームが出力されるまでとなる。よって、領域1が入力される時間3.33msに3MB分処理する時間を足し合わせた時間となるが、3MB分の処理時間は、数十マイクロ秒のオーダー(16.666ms/3600MB×3MB=0.014ms)であり十分小さいため、本明細書では入力画像が入力されてから符号化部203からストリームが出力されるまでの時間は約3.33msとみなせる。
 このように、目標遅延時間分の領域の入力画像を用いて、入力画像複雑度を計算後、実際のエンコード処理を開始することにより、目標遅延時間分の領域の入力画像の絵柄の符号化難易度が分かるため、目標遅延時間分の領域に適したビット配分を行い高画質化することができる。ここで、目標遅延時間を変更する場合は、目標遅延時間に対応して領域サイズが変更され、エンコード処理の開始タイミングを変更することで、対応可能である。
 次に図5を用いてQP値の決定方法の具体例を説明する。図5は、QP値計算部108の内部の詳細を示した図であり、ベースQP計算部501、MB(マクロブロック)QP計算部502、QP算出部503から構成される。
 ベースQP計算部501は、領域(領域1~領域5)を跨ぐときにのみ実行される処理であり、入力画像複雑度と符号量を用いて次に処理する領域のベースQPを出力する。このベースQPは下記(式2)で与えられる。
Figure JPOXMLDOC01-appb-I000002
 ここで、QPaveは、前の領域の平均QP値、bitrateは前の領域の発生符号量、target_bitrateは次の領域の目標符号量、αは係数、varnextは次の領域の入力画像複雑度、varpreは前の領域の入力画像複雑度を示している。(式2)により、前の領域の発生符号量と次の領域全体の入力画像複雑度を加味して、次の領域のベースQPを決定することが可能となるため、発生符号量を精度よく推測することができる。さらに、前回符号化した際の平均QP値と発生符号量を元にループ処理しているため、(式2)は画像に合わせて最適化され、推測する符号量と実際の発生符号量の間の誤差を小さくしている。
 次に、MBQP計算部502は、MB毎に実行される処理であり、符号化単位複雑度と入力画像複雑度からMBQPを出力する。このMBQPは(式3)で与えられる。
Figure JPOXMLDOC01-appb-I000003
 ここで、βは係数、γはリミッタ値を示しており、γよりMBQPが大きい場合はγ、-γよりMBQPが小さい場合は-γとする。(式3)を用いると、入力画像複雑度より符号化単位複雑度が大きい複雑な絵柄のMBは、MBQPを大きくして発生符合量を抑える。逆に入力画像複雑度より符号化単位複雑度が小さい平坦な絵柄のMBは、MBQPを小さくして符号量を割り当てて主観画質を良くすることができる。
 このように、平坦な絵柄の劣化に敏感な人間の視覚特性に合わせた適切なビット配分が可能となり、高画質化することができる。また、この処理は、発生する符合量の多い複雑な絵柄のMBQPを大きくして、発生する符号量の少ない平坦な絵柄のMBQPを小さくすることになるため、主観画質の向上だけではなく、MB毎の発生符号量を平滑化する効果も備えている。
 最後にQP算出部503では、(式4)に記載の式により、QP値を算出する。
Figure JPOXMLDOC01-appb-I000004
 以上、実施例1に記載の符号化装置100では、QP値計算部108のベースQP計算部501における発生符号量を一定にする処理と、QP値計算部108のMBQP計算部502における絵柄に応じたQP値の制御により、目標遅延時間毎の発生符号量を一定にしつつ、高画質化が実現可能となる。
 また、実施例1の構成は、目標遅延時間をエンコード中やアプリケーション毎に変える場合にも有効である。入力画像複雑度計算部102での入力画像複雑度の計算と並行して符号化単位毎の符号化単位複雑度を計算する場合、目標遅延時間の領域内にあるMB数分の符号化単位複雑度をメモリに記録しておく必要がある(実施例1の例では、720MB分のメモリが必要)。
 これに対して、実施例1の構成は、符号化単位画像を入力画像用メモリ103から読み出したタイミングで、符号化単位複雑度を計算しているので、パイプライン処理の遅延分(実施例1の例では、次のステージで使用するので1MB分)のメモリを持っていればよく、目標遅延時間によらないで固定のメモリを持てばよい。これは、4k8kサイズなど大きな画像サイズの画像を符号化する場合に、少ない固定量のメモリで良いため特に有効である。
 また、エンコード中のパイプライン処理を変更することなく、実現可能な構成であり、本処理を導入したことによる、エンコード処理の遅延は発生しない。
 実施例1では、1フレーム以下の目標遅延時間の例で説明したが、蓄積用途の場合、リアルタイム性が要求されないので、目標遅延時間はメモリの容量が許す限り遅く設定してもよい。例えば、3フレーム分を目標遅延時間と設定した場合、入力画像を3枚分解析してからビット配分を行うことができるため、高画質化を実現できる。
 また、ラスタキャン順に入力画像が入力される例を説明したが、ラスタスキャン順ではなく、例えば、K画素×L画素毎に一度に入力されるとしても良い。
 また、H.264についての例を挙げたが、符号化単位毎に画像の品質を変更できるパラメータを持った動画像符号化方式(MPEG2、次世代の動画像符号化方式HEVC(H.265)など)であれば、本構成を用いることで同様の効果が得られる。
 また、複雑度は分散値の例を説明したが、隣の画素との差分値の合計、エッジ検出フィルタ(Sobelフィルタ、ラプラシアンフィルタなど)の合計値など、画像の複雑度を示す指標であれば分散値に限らない。
 また、図5を用いてQP値計算部108の具体的なQP値の決定手法を説明したが、この処理に限定されるものではない。少なくとも入力画像複雑度と符号化単位複雑度と発生符号量を用いて、目標遅延時間に応じて発生符号量を一定にしつつ高画質化が実現できていれば良い。
 また、入力画像複雑度計算部102で複雑度を計算する領域の大きさは、目標遅延時間に入力される領域の大きさにした場合の例を説明したが、目標遅延時間に入力される領域の大きさより、小さい領域毎に複雑度を計算するとしても良い。例えば、上記例では、720MB分の領域で入力画像複雑度を計算していたが、1/3の240MBの領域毎に入力画像複雑度を計算するとする。この場合、240MB毎に(式4’)に示す符号量補正QPの項を加えることができる。ここで、符号量補正QPは、240MB毎に計算される値であり、直前240MBで発生した符号量を元に計算される。例えば、直前240MBで発生した符合量が一定にしたい符号量より、多かった場合は、符号量補正QP値をプラス値に、逆に発生符号量が一定にしたい符号量より、少なかった場合は、符号量QP値をマイナス値にすることで、目標遅延時間における発生符号量を一定にする精度を高めることが可能となる。
Figure JPOXMLDOC01-appb-I000005
 次に、車載ネットワークカメラシステムを例に再生画像で画像認識を行う場合に、画像認識処理の性能向上が可能な画像符号化装置の例を説明する。
 まず図6を用いて前提としているカメラシステムの構成図の例を説明する。図6のカメラシステムは画像送信装置1000と画像受信装置1100から構成される。画像送信装置1000は、例えば車載カメラであり、光をデジタルの画像に変換する撮像部1001と、撮像部1001から出力されたデジタルの画像をエンコード処理してストリームを生成する画像符号化部1002と、エンコード処理したストリームをパケット化してネットワーク上に出力するネットワークIF1003で構成される。
 また、画像受信装置1100は、例えばカーナビゲーションシステムであり、映像送信装置1000から送信されたパケットを受け取りストリームに変換するネットワークIF1101と、ネットワークIF1101から出力されたストリームを復号処理して再生画像を生成する画像復号部1102と、画像復号部1102から出力された再生画像をディスプレイなどに表示する表示部1103と、画像復号部1103から出力された再生画像に画像認識処理する画像認識部1104と画像認識した結果が危険な状態を示す結果となった場合に音声を出力して運転者に知らせる音声出力部1105を備える。
 以上、図6の車載ネットワークカメラシステムの構成を例に、再生画像で画像認識処理を行なう場合に画像認識処理の性能向上が可能な画像符号化装置1002の構成図を図7に示す。図1の画像符号化装置100のうち、既に説明した図1に示された同一の符号を付された構成と、同一の機能を有する部分については、説明を省略する。
 図1と異なる点は、符号化単位特徴量抽出部110と、QP値計算部111である。符号化単位特徴量抽出部110は、MB画像から画像特徴量を抽出し、符号化単位画像特徴量を出力する。
 なお、入力画像複雑度計算部102と符号化単位複雑度計算部107と符号化単位特徴量抽出部とを含む構成を単に解析部ともいう。
 図8にQP値計算部111の内部構成を示す。図5のQP値計算部108のうち既に説明した図5に示された同一の符号を付された構成と同一の機能を有する部分については、説明を省略する。
 図5と異なる点は、特徴QP計算部504とQP計算部505である。特徴QP計算部504は、符号化単位画像特徴量から特徴量の大きさに従い特徴QPを出力する。QP計算部505は、ベースQPとMBQPだけでなく特徴QPも加味した(式5)で与えられる式でQP値を計算する。
Figure JPOXMLDOC01-appb-I000006
 ここで、図6の車載ネットワークカメラシステムの受信装置1100で道路に引かれている白線の認識をする場合を例に、特徴QPの計算方法を説明する。まず、符号化単位特徴量抽出部110では、白線認識で必要となる白線特徴量の抽出を行う。具体的には隣の画素との差分値を計算し、直線上に連続的に同じ段差の差分値がある場合に、特徴量が大きくなるような3段階の白線特徴量(0、1、2の3段階であり、値が大きいほど白線である可能性が高いとする)という符号化単位画像特徴量を出力する。さらに、特徴QP計算部504では、図9のテーブルに基づいて、3段階の白線特徴量から特徴量QPを決定する。本実施例では、一例として、図9のテーブルに基づいた特徴量QPの決定手法の例を挙げたがこれに限定されるものではない。例えば、一次関数やLog関数等の数式を用いて特徴量QPを決定してもよい。また、特徴量の抽出には隣の画素との差分値を使用した例を記載したが、予め指定の絵柄を検索可能な基本パタンの画像を用意し、その基本パタンとパターンマッチングを行なった結果(類似度)を利用して、特徴量QPを決定するなど、画質を向上したい対象が判定できる指標値であればよい。また、本実施例では、符号化単位で特徴量を計算する例を説明したが、図11に示すように、入力画像複雑度計算部102が計算する領域で特徴量を計算する構成としても良い。
 図10は、図6の車載ネットワークカメラシステムを例に、アプリケーションに応じて目標遅延時間を変更する例を表している。アプリケーションとしては、駐車時の後方障害物検知、高速走行時の白線逸脱警報、市街地走行時の標識認識とする。車載ネットワークカメラシステムにおいては、車の速度に応じて利用されるアプリケーションが変わることが考えられるため、車の速度に応じて目標遅延時間を変更するようにしてもよい。
 駐車時の障害物検知は、時速20km以下の速度で使用されることが想定される。速度が遅いため、1フレーム毎の画像の変化は少ないので10ms以下の低遅延化は必要なく、目標遅延時間を33.3msとなるため、1フレーム分の解析結果を用いて高画質化できる。
 一方、高速走行時の白線逸脱警報では、時速100km以上の速度で使用することが想定される。よって、1フレーム毎に画像が大きく変化すると考えられる。遅延時間が大きいと画像認識処理をして危険な状態を検知したとしても、既に事故が発生していたということが起こる可能性があるため目標遅延時間は1msと少ないことが重要となる。
 また、市街地走行時の標識認識は、時速40km~80kmの中程度の速度なので、1msまでは行かないが目標遅延時間をある程度短くすることが必要となるため10msとする。
 このように、車載ネットワークカメラシステムのアプリケーションの場合、使用するアプリケーションに応じて、または車の速度に応じて目標遅延時間を決めることで、アプリケーションの画像認識性能を最大限に活かすことが可能となる。
 このように、車載ネットワークカメラシステムに画像符号化装置を適用した場合では、画像受信装置1100の画像認識部1104の画像認識処理のアルゴリズムが注目する画像特徴を、画像送信装置1000の符号化単位特徴抽出部110で簡易的に特徴量を抽出し、QP値計算部111で該当MBのQP値を下げることにより、目標遅延時間に対して符号量を一定と高画質化を実現しつつ、さらに画像認識処理の性能向上が可能な画像符号化装置を実現することが可能となる。
 また、本実施例では、車載ネットワークカメラシステムの例を説明したが、撮像部1001を例えば、レコーダなどの蓄積装置に置き換えても良い。
 車載ネットワークカメラシステム以外(例えばテレビ会議システム等)に画像符号化装置を適用した場合でも、使用されるアプリケーションに応じて目標遅延時間を設定することで、画像伝送における遅延を考慮しつつ、伝送する画像を高画質化することが可能となる。
100 画像符号化装置
101 入力画像書込み部
102 入力画像複雑度計算部
103 入力画像用メモリ
104 符号化単位画像読込み部
105 エンコード部
106 エンコード用メモリ
107 符号化単位複雑度計算部
108 QP値計算部
109 制御部
201 予測部
202 周波数変換・量子化部
203 符号化部
204 逆周波数変換・逆量子化部
501 ベースQP計算部
502 MBQP計算部
503 QP値算出部
1000 画像送信装置
1001 撮像部
1002 画像符号化部
1003 ネットワークIF
1100 画像受信装置
1101 ネットワークIF
1102 画像復号部
1103 表示部
1104 画像認識部
1105 音声出力部
1201 監視カメラ
1202 監視カメラ
1203 監視カメラ
1204 監視センター
1205 WAN
1301 テレビ会議システム
1302 テレビ会議システム
1303 テレビ会議システム
1304 WAN
1401 自動車
1402 車載カメラ
1403 車載カメラ
1404 モニタ
1405 LAN

Claims (14)

  1.  入力される画像を解析する解析ステップと、
     前記解析ステップの解析結果を用いて量子化パラメータを算出する量子化パラメータ算出ステップと、
     前記量子化パラメータ算出ステップで算出した量子化パラメータを用いて入力される画像のエンコードを行うエンコードステップと、を有し、
     前記解析ステップで解析する第1の領域の大きさが可変であることを特徴とする符号化方法。
  2.  請求項1の符号化方法であって、
     前記第1の領域の大きさは、前記符号化方法を用いて行われる映像伝送において設定される遅延時間に基づいて変化することを特徴とする符号化方法。
  3.  請求項1または2の符号化方法であって、
     前記第1の領域の大きさは入力される画像の1フレームの大きさよりも小さいことを特徴とする符号化方法。
  4.  請求項1~3のいずれかの符号化方法であって、
     前記解析ステップは、
     前記第1の領域を解析する第1の解析ステップと、
     前記第1の領域の一部である第2の領域を解析する第2の解析ステップと、を含むことを特徴とする符号化方法。
  5.  請求項4の符号化方法であって、
     前記第1の解析ステップでは前記第1の領域の画像の複雑度を計算し、
     前記第2の解析ステップでは前記第2の領域の画像の複雑度を計算することを特徴とする符号化方法。
  6.  請求項4または5の符号化方法であって、
     前記解析ステップは、前記第2の領域の画像の特徴量を抽出する第3の解析ステップを含むことを特徴とする符号化方法。
  7.  請求項4~6のいずれかの符号化方法であって、
     前記第2の領域は、前記エンコードステップにおけるエンコードの単位である符号化単位の領域であることを特徴とする符号化方法。
  8.  入力される画像を解析する解析部と、
     前記解析部の解析結果を用いて量子化パラメータを算出する量子化パラメータ算出部と、
     前記量子化パラメータ算出部で算出した量子化パラメータを用いて入力される画像のエンコードを行うエンコード部と、を有し、
     前記解析部解析する第1の領域の大きさが可変であることを特徴とする符号化装置。
  9.  請求項8の符号化装置であって、
     前記第1の領域の大きさは、前記符号化装置を用いて行われる映像伝送において設定される遅延時間に基づいて変化することを特徴とする符号化装置。
  10.  請求項8または9の符号化装置であって、
     前記第1の領域の大きさは、入力される画像の1フレームの大きさよりも小さいことを特徴とする符号化装置。
  11.  請求項8~10のいずれかの符号化装置であって、
     前記解析部は、
     前記第1の領域を解析する第1の解析部と、
     前記第1の領域の一部である第2の領域を解析する第2の解析部と、を含むことを特徴とする符号化装置。
  12.  請求項11の符号化装置であって、
     前記第1の解析部は前記第1の領域の画像の複雑度を計算し、
     前記第2の解析部は前記第2の領域の画像の複雑度を計算することを特徴とする符号化装置。
  13.  請求項11または12の符号化装置であって、
     前記解析部は、前記第2の領域の画像の特徴量を抽出する第3の解析部を含むことを特徴とする符号化装置。
  14.  請求項11~13のいずれかの符号化装置であって、
     前記第2の領域は、前記エンコード部におけるエンコードの単位である符号化単位の領域であることを特徴とする符号化装置。
PCT/JP2013/058487 2013-03-25 2013-03-25 符号化方法および符号化装置 Ceased WO2014155471A1 (ja)

Priority Applications (4)

Application Number Priority Date Filing Date Title
PCT/JP2013/058487 WO2014155471A1 (ja) 2013-03-25 2013-03-25 符号化方法および符号化装置
US14/653,483 US10027960B2 (en) 2013-03-25 2013-03-25 Coding method and coding device
JP2015507704A JP6084682B2 (ja) 2013-03-25 2013-03-25 符号化方法および符号化装置
CN201380067188.5A CN104871544B (zh) 2013-03-25 2013-03-25 编码方法以及编码装置

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2013/058487 WO2014155471A1 (ja) 2013-03-25 2013-03-25 符号化方法および符号化装置

Publications (1)

Publication Number Publication Date
WO2014155471A1 true WO2014155471A1 (ja) 2014-10-02

Family

ID=51622567

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2013/058487 Ceased WO2014155471A1 (ja) 2013-03-25 2013-03-25 符号化方法および符号化装置

Country Status (4)

Country Link
US (1) US10027960B2 (ja)
JP (1) JP6084682B2 (ja)
CN (1) CN104871544B (ja)
WO (1) WO2014155471A1 (ja)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10356408B2 (en) * 2015-11-27 2019-07-16 Canon Kabushiki Kaisha Image encoding apparatus and method of controlling the same

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH0541806A (ja) * 1991-08-06 1993-02-19 Olympus Optical Co Ltd 画像記録装置
JPH08102947A (ja) * 1994-09-29 1996-04-16 Sony Corp 画像信号符号化方法及び画像信号符号化装置
JPH11164305A (ja) * 1997-04-24 1999-06-18 Mitsubishi Electric Corp 動画像符号化方法、動画像符号化装置および動画像復号装置
JP2006519565A (ja) * 2003-03-03 2006-08-24 コーニンクレッカ フィリップス エレクトロニクス エヌ ヴィ ビデオ符号化
JP2009246540A (ja) * 2008-03-28 2009-10-22 Ibex Technology Co Ltd 符号化装置、符号化方法および符号化プログラム
JP2010141659A (ja) * 2008-12-12 2010-06-24 Sony Corp 情報処理装置および方法

Family Cites Families (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH0797753B2 (ja) 1989-01-24 1995-10-18 日本ビクター株式会社 符号化出力データ量の制御方式
TW388843B (en) * 1997-04-24 2000-05-01 Mitsubishi Electric Corp Moving image encoding method, moving image encoder and moving image decoder
JP3324551B2 (ja) 1999-03-18 2002-09-17 日本電気株式会社 画像信号の符号化制御装置
JP3893344B2 (ja) * 2002-10-03 2007-03-14 松下電器産業株式会社 画像符号化方法および画像符号化装置
JP4257655B2 (ja) * 2004-11-04 2009-04-22 日本ビクター株式会社 動画像符号化装置
JP5128389B2 (ja) * 2008-07-01 2013-01-23 株式会社日立国際電気 動画像符号化装置及び動画像符号化方法
CN105872541B (zh) * 2009-06-19 2019-05-14 三菱电机株式会社 图像编码装置、图像编码方法及图像解码装置

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH0541806A (ja) * 1991-08-06 1993-02-19 Olympus Optical Co Ltd 画像記録装置
JPH08102947A (ja) * 1994-09-29 1996-04-16 Sony Corp 画像信号符号化方法及び画像信号符号化装置
JPH11164305A (ja) * 1997-04-24 1999-06-18 Mitsubishi Electric Corp 動画像符号化方法、動画像符号化装置および動画像復号装置
JP2006519565A (ja) * 2003-03-03 2006-08-24 コーニンクレッカ フィリップス エレクトロニクス エヌ ヴィ ビデオ符号化
JP2009246540A (ja) * 2008-03-28 2009-10-22 Ibex Technology Co Ltd 符号化装置、符号化方法および符号化プログラム
JP2010141659A (ja) * 2008-12-12 2010-06-24 Sony Corp 情報処理装置および方法

Also Published As

Publication number Publication date
JPWO2014155471A1 (ja) 2017-02-16
CN104871544A (zh) 2015-08-26
JP6084682B2 (ja) 2017-02-22
US10027960B2 (en) 2018-07-17
US20150350649A1 (en) 2015-12-03
CN104871544B (zh) 2018-11-02

Similar Documents

Publication Publication Date Title
US10582196B2 (en) Generating heat maps using dynamic vision sensor events
US9706203B2 (en) Low latency video encoder
US10212456B2 (en) Deblocking filter for high dynamic range (HDR) video
CN103650504B (zh) 基于图像捕获参数对视频编码的控制
US8457205B2 (en) Apparatus and method of up-converting frame rate of decoded frame
US8493499B2 (en) Compression-quality driven image acquisition and processing system
US11102475B2 (en) Video encoding device, operating methods thereof, and vehicles equipped with a video encoding device
US12316908B2 (en) System and method to minimize repetitive flashes in streaming media content
US20130195206A1 (en) Video coding using eye tracking maps
EP3648460A1 (en) Method and apparatus for controlling encoding resolution ratio
US11924424B2 (en) Method and device for transmitting block division information in image codec for security camera
US12563238B2 (en) Encoding of pre-processed image frames
WO2012160626A1 (ja) 画像圧縮装置、画像復元装置、及びプログラム
JP6084682B2 (ja) 符号化方法および符号化装置
EP4672738A1 (en) SELECTIVE FRAME PROCESSING IN A BLOW-BASED CODING PIPELINE
Benjak et al. Deferred demosaicking: efficient first-person view drone video encoding
WO2023132163A1 (ja) 映像圧縮方法、映像圧縮装置、コンピュータプログラム、及び映像処理システム
CN111901605B (zh) 视频处理方法、装置、电子设备及存储介质
EP4637138A1 (en) Adaptive coding tool selection with content classification
US11716475B2 (en) Image processing device and method of pre-processing images of a video stream before encoding
US20250254329A1 (en) Video encoding with content adaptive resolution decision
US11089308B1 (en) Removing blocking artifacts in video encoders
Shirahase et al. Video Transfer Method Suitable for Object Detection in the Cloud
Isnardi et al. Salience-based compression: providing FMV over low-bit rate channels

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 13879706

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2015507704

Country of ref document: JP

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 14653483

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 13879706

Country of ref document: EP

Kind code of ref document: A1