WO2024005160A1 - 画像処理装置、画像処理方法、画像処理システム、およびプログラム - Google Patents
画像処理装置、画像処理方法、画像処理システム、およびプログラム Download PDFInfo
- Publication number
- WO2024005160A1 WO2024005160A1 PCT/JP2023/024253 JP2023024253W WO2024005160A1 WO 2024005160 A1 WO2024005160 A1 WO 2024005160A1 JP 2023024253 W JP2023024253 W JP 2023024253W WO 2024005160 A1 WO2024005160 A1 WO 2024005160A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- image
- face
- input image
- person
- anonymization
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/774—Generating sets of training patterns; Bootstrap methods, e.g. bagging or boosting
Definitions
- the present invention relates to an image processing device, an image processing method, an image processing system, and a program.
- Patent Document 1 discloses a technology that generates a composite face image by referring to face images of multiple people stored in a face image database, and enables annotation operations to be performed on the generated composite face image. .
- Patent Document 1 protects the privacy of multiple people by having an annotator perform an annotation operation on a composite face image synthesized from facial images of multiple people.
- characteristic information of the original image may be missing due to converting the original image to protect privacy.
- the present invention has been made in consideration of these circumstances, and provides image processing that can generate learning data effective for training machine learning models while protecting the privacy of people depicted in facial images.
- One of the purposes is to provide an apparatus, an image processing method, an image processing system, and a program. This in turn contributes to the development of sustainable transportation systems.
- An image processing device includes: an image attribute acquisition unit that acquires an image attribute that is a shooting mode of an input image; an image conversion unit that performs anonymization processing on the input image; an image determination unit that determines whether or not the input image subjected to the anonymization process satisfies predetermined requirements; If it is determined that the requirements are met, predetermined processing is performed on the input image that has been subjected to the anonymization processing, and the predetermined requirements are determined according to the acquired image attributes.
- the predetermined process is a process of saving the input image that has been subjected to the anonymization process as a target image for annotation work.
- the predetermined process includes converting the input image that has been subjected to the anonymization process to generate a behavior prediction model that predicts the behavior of a person photographed in the input image. This is a process to save as learning information.
- the predetermined process is a process of transmitting the input image that has been subjected to the anonymization process to an image server through a communication means.
- the image attribute indicates whether the input image is an image of the inside of a vehicle in which a camera that imaged the input image is mounted, or an image of the outside of the vehicle. This is information that at least indicates that.
- the anonymization process includes a process of changing the face of the person depicted in the input image to the face of another person, and the predetermined requirement is that the image attribute is If the image indicates the inside of a vehicle, it includes whether or not the line of sight direction of the person's face matches the line of sight direction of the other person's face.
- the predetermined requirement is that the line of sight direction of the person's face and the other person
- the predetermined requirement includes whether or not the direction of line of sight of the face of the person matches, and whether the direction of the face of the person's face matches the direction of the face of the face of the other person. If the attribute indicates that the image is an image of the outside of the vehicle, it includes whether or not the face direction of the person's face matches the face direction of the other person's face.
- the predetermined requirement is that the line of sight direction of the person's face and the other person This does not include whether the line of sight direction of the face matches or not.
- An image processing system includes an image attribute acquisition unit that acquires an image attribute that is a shooting mode of an input image, and an image conversion unit that performs anonymization processing on the input image. , an image determination unit that determines whether or not the input image that has been subjected to the anonymization process satisfies predetermined requirements; If it is determined that the predetermined requirements are met, predetermined processing is performed on the input image that has been subjected to the anonymization processing, and the predetermined requirements are determined according to the acquired image attributes.
- a computer acquires an image attribute that is a shooting mode of an input image, performs anonymization processing on the input image, and performs anonymization processing on the input image. It is determined whether the input image subjected to the anonymization process satisfies predetermined requirements, and if it is determined that the input image subjected to the anonymization process satisfies the predetermined requirements, the input image subjected to the anonymization process A predetermined process is performed on the image, and the predetermined requirements are determined according to the acquired image attributes.
- a program causes a computer to acquire an image attribute that is a shooting mode of an input image, performs anonymization processing on the input image, and performs anonymization processing on the input image. If it is determined that the input image subjected to the anonymization process satisfies the predetermined requirements, the input image subjected to the anonymization process is The predetermined requirements are determined according to the acquired image attributes.
- FIG. 1 is a diagram showing an overview of a system 1 including an image processing device 100 according to the present embodiment.
- 1 is a diagram illustrating an example of a functional configuration of an image processing apparatus 100 according to the present embodiment. It is a figure which shows an example of the vehicle interior image and vehicle exterior image acquired from the vehicle M1.
- 3 is a diagram for explaining processing executed by an image processing unit 130.
- FIG. FIG. 3 is a diagram for explaining processing executed by an image conversion unit 140.
- FIG. 3 is a diagram illustrating an example of time-series in-vehicle images converted by the image conversion unit 140.
- FIG. FIG. 3 is a diagram for explaining processing executed by an image determination unit 150.
- FIG. FIG. 3 is a diagram illustrating an example of an annotation work performed by an annotator.
- FIG. 5 is a diagram illustrating an example of driving support using a trained model 180.
- FIG. 3 is a diagram illustrating an example of the flow of processing executed by an image conversion unit 140.
- FIG. 3 is a diagram illustrating an example of the flow of processing executed by the image determination unit 150.
- FIG. 1 is a diagram showing an overview of a system 1 including an image processing apparatus 100 according to the present embodiment.
- the system 1 includes at least one vehicle M1 and one vehicle M2, an image processing device 100, and a terminal device 200.
- vehicle M1 and vehicle M2 are illustrated as different vehicles, but these vehicles may be the same vehicle.
- the vehicle M1 is, for example, a four-wheel drive vehicle such as a hybrid vehicle or an electric vehicle, and includes at least a camera that images the inside of the vehicle M1 and a camera that images the outside of the vehicle M1. While the vehicle M1 is traveling, the vehicle interior image and the vehicle exterior image captured by these cameras are transmitted to the image processing device 100 via a network NW such as a cellular network, a Wi-Fi network, or the Internet.
- NW such as a cellular network, a Wi-Fi network, or the Internet.
- the image processing device 100 is a server device that, upon receiving captured image data including an inside image and an outside image from the vehicle M1, performs image conversion to be described later on the received captured image data. This image conversion is a process for protecting the privacy of the person photographed in the in-vehicle image and the out-of-vehicle image.
- the image processing device 100 transmits the obtained converted image data to the terminal device 200 via the network NW.
- the terminal device 200 is a terminal device such as a desktop computer or a smartphone.
- the user of the terminal device 200 acquires the converted image data from the image processing device 100
- the user of the terminal device 200 performs an annotation operation to be described later on the acquired converted image data.
- the user of the terminal device 200 transmits the annotated image data in which the annotation is added to the converted image data to the image processing device 100.
- the image processing device 100 Upon receiving the annotated image data from the terminal device 200, the image processing device 100 uses the received annotated image data as learning data to generate a trained model, which will be described later, using an arbitrary machine learning model.
- this trained model can output the predicted behavior (trajectory) of the person in the image outside the vehicle in response to the input of the image outside the vehicle, or output the predicted behavior (trajectory) of the person depicted in the image outside the vehicle, or This is a behavior prediction model that takes into account the line of sight of the driver in the image and calls attention to pedestrians in the image outside the vehicle.
- the image data used as learning data at this time may be annotated image data in which annotations are added to the converted image data, or annotated image data obtained by reconverting the converted image data into captured image data while leaving the annotations as they are. It may be annotated image data (that is, annotated image data in which an annotation is added to captured image data).
- annotated image data in which captured image data is annotated as learning data it is possible to use learning data that is more realistic and free from the effects of image conversion.
- the image processing device 100 distributes the generated trained model to the vehicle M2 via the network NW.
- vehicle M2 is, for example, a four-wheel drive vehicle such as a hybrid vehicle or an electric vehicle, and vehicle M2 learns at least one of an interior image and an exterior image captured by a camera while driving.
- behavioral prediction data of people existing around the vehicle M2 is obtained.
- the driver of the vehicle M2 can refer to the obtained behavior prediction data and utilize it for driving the vehicle M2. The details of each process will be explained below.
- FIG. 2 is a diagram showing an example of the functional configuration of the image processing apparatus 100 according to the present embodiment.
- the image processing device 100 includes, for example, a communication unit 110, a transmission/reception control unit 120, an image processing unit 130, an image conversion unit 140, an image determination unit 150, a trained model generation unit 160, a storage unit 170, Equipped with.
- These components are realized by, for example, a hardware processor such as a CPU (Central Processing Unit) executing a program (software).
- Some or all of these components are hardware (circuit parts) such as LSI (Large Scale Integration), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array), and GPU (Graphics Processing Unit).
- the program may be stored in advance in a storage device (a storage device with a non-transitory storage medium) such as an HDD (Hard Disk Drive) or flash memory, or may be stored in a removable storage device such as a DVD or CD-ROM. It is stored in a medium (non-transitory storage medium), and may be installed by loading the storage medium into a drive device.
- the storage unit 170 is, for example, an HDD, a flash memory, a RAM (Random Access Memory), or the like.
- the storage unit 170 stores, for example, captured image data 172, converted image data 174, annotation image data 176, annotated image data 178, and a trained model 180.
- the image processing device 100 includes a trained model generation unit 160 and a storage unit 170 that stores the trained model 180.
- the model may be held by a server device different from the image processing device 100.
- the communication unit 110 is an interface that communicates with the communication device 10 of the own vehicle M via the network NW.
- the communication unit 110 includes a NIC (Network Interface Card), an antenna for wireless communication, and the like.
- the transmission/reception control unit 120 uses the communication unit 110 to transmit and receive data with the vehicles M1 and M2 and the terminal device 200. More specifically, first, the transmission/reception control unit 120 acquires, from the vehicle M1, a plurality of in-vehicle images and out-of-vehicle images captured in chronological order by a camera mounted on the vehicle M1.
- the time series in this case is, for example, images taken at predetermined intervals (for example, every second) in one driving cycle from the start to the stop of the vehicle M1.
- FIG. 3 is a diagram showing an example of an inside image and an outside image acquired from the vehicle M1.
- the left part of FIG. 3 represents the inside image acquired from vehicle M1, and the right part of FIG. 3 represents the outside image acquired from vehicle M1.
- the in-vehicle image is captured with a camera installed to capture at least the face area of the driver of the vehicle M1, and as shown in the right part of FIG.
- the image is captured with a camera installed so as to capture at least an image in front of the vehicle M1 in the direction of travel.
- the transmission/reception control unit 120 stores the inside image and the outside image acquired from the vehicle M1 in the storage unit 170 in association with the image ID as captured image data 172.
- FIG. 4 is a diagram for explaining the processing executed by the image processing unit 130.
- the image processing unit 130 performs image processing on the captured image data 172 and obtains information such as the image attribute, face attribute, direction, etc. of each image included in the captured image data 172. More specifically, when an image is input, the image processing unit 130 processes each image included in the captured image data 172 using a trained model that outputs a classification result indicating whether the image is an inside image or an outside image. Obtain the image attribute indicating whether the image is inside the car or outside the car.
- the image processing unit 130 calculates the face area, face size (area of the face area), and distance from the image capturing position to the face for all faces included in the image.
- the facial attributes of each image included in the captured image data 172 are acquired using the trained model to be output.
- the face area FA1 of the person P1 is acquired from the image inside the car
- the face area FA2 of the person P2 the face area FA3 of the person P3, and the face area FA4 of the person P4 are acquired from the image outside the car.
- the face areas FA1, FA2, FA3, and FA4 are acquired as rectangular areas, but the present invention is not limited to such a configuration.
- a model may also be used.
- the image processing unit 130 converts the captured image data 172 into captured image data 172 using a trained model that outputs at least one of the face direction and the gaze direction for all faces included in the image, for example, as a vector. Obtain the direction information of the face shown in each included image. More specifically, when the image of the captured image data 172 having the attribute of an in-vehicle image is input, the image processing unit 130 outputs the face direction and line of sight direction for all faces included in the image. Obtain direction information using the trained model.
- the image processing unit 130 uses a trained model that outputs face directions for all faces included in the image to determine the direction. Get information. This is because, compared to images outside the car, faces in images inside the car are generally closer from the shooting position, and the faces tend to be large enough to extract the line of sight. It is from.
- the face direction FD1 and line-of-sight direction ED1 of the person P1 are acquired from the image inside the car
- the face direction FD2 of the person P2 the face direction FD3 of the person P3, and the face direction of the person P4 are acquired from the image outside the car.
- FD4 has been acquired.
- the image processing unit 130 When the image processing unit 130 acquires the image attributes, face attributes, and direction information for each image of the captured image data 172, it records these image attributes, face attributes, and direction information in association with the image. Note that in the above, as an example, the image processing unit 130 acquires image attributes, face attributes, and direction information using a trained model, but the present invention is not limited to such a configuration, and image processing The unit 130 may acquire these image attributes, face attributes, and direction information using any known method.
- the image conversion unit 140 performs processing on the captured image data 172 processed by the image processing unit 130 to replace the face of the person in each image with the face of another person without changing the direction information of the person in each image. , using any software that implements such functionality.
- FIG. 5 is a diagram for explaining the processing executed by the image conversion unit 140. As shown in FIG. 5, the image conversion unit 140 replaces the faces of persons P1, P2, and P3 shown in FIG. There is. On the other hand, the face of the person P4 is covered by the mosaic MS as a result of the mosaic processing performed by the image conversion unit 140.
- the image conversion unit 140 determines whether to replace the face with another person's face or to perform mosaic processing based on the facial attributes of each face captured in each image of the captured image data 172. More specifically, the image conversion unit 140 determines whether or not the size of each face captured in each image of the captured image data 172 is equal to or larger than the first threshold Th1, and If it is determined that the face is equal to or greater than the first threshold Th1, it is determined that the face is replaced with the face of another person. On the other hand, if it is determined that the size of the face is less than the first threshold Th1, the image conversion unit 140 determines to perform mosaic processing on the face. Replacing the face of a person in a captured image with the face of another person or performing mosaic processing is an example of "anonymization processing."
- the image conversion unit 140 also determines whether the distance of each face captured in each image of the captured image data 172 is less than or equal to the second threshold Th2, and determines whether the distance of the face is less than or equal to the second threshold Th2. If it is determined that the following is true, it is determined that the face in question is to be replaced with the face of another person. On the other hand, if it is determined that the distance between the faces is greater than the second threshold Th2, the image conversion unit 140 determines to perform mosaic processing on the face. The image conversion unit 140 repeatedly executes these determination processes for the number of faces shown in the image, and replaces each face with the face of another person or performs mosaic processing according to the determination results.
- the image conversion unit 140 stores image data obtained by performing such processing on the captured image data 172 in the storage unit 170 as converted image data 174. This makes it possible to select useful data as learning data for generating a behavior prediction model, and to protect the privacy of the person depicted in each image when an annotator (described later) performs annotation work.
- the image conversion unit 140 converts the face into a different person if the size of the face is greater than or equal to the first threshold Th1 and the distance between the faces is less than or equal to the second threshold Th2.
- the face may be replaced with the face of another person. You may decide that.
- the image conversion unit 140 selects faces to be used as learning data by performing mosaic processing on faces for which direction information acquisition has failed among the faces depicted in each image of the captured image data 172. Good too.
- FIG. 6 is a diagram showing an example of a time-series in-vehicle image converted by the image conversion unit 140.
- FIG. 6 shows, as an example, an example in which time-series in-vehicle images at three time points, t, t+1, and t+2, are converted.
- These time-series in-car images are images of the same person and their faces are converted, but as shown in Figure 6, depending on the operation of the face conversion software, the face of the same person may be converted into the faces of multiple different people. It is possible. Even though the face of the same person has been converted into the faces of a plurality of different people, it is not preferable to use such converted image data as learning data as it is because it will deteriorate the accuracy of the behavior prediction model. Therefore, the image determination unit 150 determines the continuity of the time-series inside-vehicle images and outside-vehicle images by executing the process described below.
- FIG. 7 is a diagram for explaining the processing executed by the image determination unit 150.
- the image determination unit 150 first extracts feature points representing the face of the person captured in the converted image. For example, the image determination unit 150 extracts feature points representing the right eye REP, left eye LEP, nose NP, right mouth corner RMP, left mouth corner LMP, and ear EP of the person's face captured in the converted image.
- the image determination unit 150 extracts feature points of faces of people tracked as the same person from each of the time-series converted images, and collates these feature points. Note that whether or not the two people have been "tracked as the same person" can be determined by, for example, associating the same person shown in the captured images at a stage before converting the images.
- the image determination unit 150 extracts the feature points of the person shown in the converted image at time t and the feature points of the person shown in the converted image at time t+1.
- the image determining unit 150 performs matching by determining whether these two sets of extracted feature points substantially match by translation or rotation.
- the image determination unit 150 determines that the faces of the people tracked as the same person are still the faces of the same person even after conversion (that is, the faces are continuous It is determined that there is a gender.
- the image determination unit 150 determines that the faces of the people tracked as the same person are not the faces of the same person after conversion (i.e., There is no continuity in the face). In that case, the image conversion unit 140 performs the conversion process again on the face determined to have no continuity.
- the image conversion unit 140 may perform the conversion process again only on the faces that have been determined to have no continuity, or may perform the conversion process again on all human faces captured in the time-series converted images. You may do so.
- the image conversion unit 140 may perform mosaic processing on faces determined to have no continuity without performing conversion processing again, and exclude them from targets to be used as learning data.
- the image determination unit 150 may restrict performing predetermined processing on the time-series converted images, that is, exclude the time-series converted images from targets to be used as learning data. This makes it possible to prevent discontinuities from occurring due to unintended operations of the face conversion software.
- the image determination unit 150 further inputs the converted image again to the above-mentioned trained model that outputs at least one of the face direction and the line-of-sight direction, and obtains the face direction FD or the line-of-sight direction ED in the converted image.
- the image determination unit 150 determines, with respect to a person's face captured in the converted image, that the face direction FD or line-of-sight direction ED of the face is the same as the face direction FD or line-of-sight direction ED of the face captured in the captured image before conversion. Determine whether or not they substantially match.
- both the face direction FD and line-of-sight direction ED are acquired for the in-vehicle image, and the face direction FD is acquired for the out-of-vehicle image. Therefore, the image determination unit 150 determines whether or not the face direction FD and line-of-sight direction ED substantially match between the captured image before conversion and the converted image for the in-vehicle image, and the converted image for the outside of the vehicle image. It is determined whether the face direction FD substantially matches between the previous captured image and the converted image. More specifically, for example, the image determination unit 150 calculates the angular difference between the vector representing the face direction FD in the captured image before conversion and the vector representing the face direction FD in the converted image, and calculates the calculated angle. If the difference is within the threshold, it is determined that the face directions FD substantially match. The same applies to the viewing direction ED. Satisfying the continuity of faces or the consistency of direction information is an example of a "predetermined requirement.”
- the image conversion unit 140 determines that the face direction FD or the line-of-sight direction ED does not substantially match. Conversion processing is again performed on the photographed image for the face. At this time, the image conversion unit 140 may perform the conversion process again only on the faces determined to be substantially non-matching, or may perform the conversion process again on all faces included in the converted image including the faces determined to be substantially non-matching. The conversion process may be performed again. Furthermore, for example, the image conversion unit 140 may perform mosaic processing on faces that are determined to be substantially non-matching, without performing conversion processing again, and may exclude them from targets to be used as learning data.
- the continuity of the transformed images executed by the image determination unit 150 described above is determined.
- the determination process regarding the matching of the direction information and the determination process regarding the matching of the direction information may be performed not for all faces captured in the converted image but only for faces that are assumed to have a higher degree of importance.
- the image determination unit 150 selects only those faces whose size is equal to or larger than the third threshold Th3, which is larger than the first threshold Th1, in the captured image before conversion.
- these determination processes may be performed only for faces whose face distance is equal to or less than a fourth threshold Th4, which is smaller than the second threshold Th2.
- the image determination unit 150 may give more importance to the face of a person who is present in the front in the direction of travel of vehicle M1 or the face of a person whose face direction is toward the front in the direction of travel of vehicle M1.
- These determination processes may be performed assuming that the degree of failure is high.
- re-conversion processing may be performed for that face and a face that is assumed to have a high degree of importance.
- the image determination unit 150 stores the converted image data 174 whose continuity and consistency have been confirmed as annotation image data 176 in the storage unit 170. do.
- the converted image data 174 is converted together with information indicating the purpose of use, for example, information indicating that the image data is annotation image data for generating a behavior prediction model that predicts the behavior of a person depicted in the input image.
- the image data 174 may be stored in the storage unit 170 as annotation image data 176.
- the transmission/reception control unit 120 transmits the annotation image data 176 to the terminal device 200.
- the annotator who is the user of the terminal device 200 , generates annotated image data by performing an annotation work on the annotation image included in the received annotation image data 176 , and transmits the annotated image data to the image processing device 100 .
- the image processing device 100 stores the received annotated image data in the storage unit 170 as annotated image data 178.
- the converted image data 174 may be stored in the storage unit 170 as the annotation image data 176.
- the image determination unit 150 may detect, for example, that the time-series captured images (or converted images thereof) obtained at predetermined intervals (for example, every second) in one driving cycle are missing due to a camera malfunction or the like. If there are images in the time series, it is not necessary to store all of these time-series images in the storage unit 170 as the annotation image data 176.
- FIG. 8 is a diagram showing an example of annotation work performed by an annotator.
- the left part of FIG. 8 represents the annotation of the converted image of the inside of the vehicle image
- the right part of FIG. 8 represents the annotation of the converted image of the outside of the vehicle image.
- the annotator determines whether or not the driver's viewing direction ED1 shown in the converted image is appropriate in the situation shown in the converted image of the outside of the vehicle at the same time.
- information for example, 1 if appropriate, 0 if inappropriate
- the converted image of the image outside the vehicle shows that there is a pedestrian on the left hand side in the vehicle's direction of travel
- the converted image of the inside image shows that the driver is looking to the left. It is shown that.
- the annotator provides information (i.e., 1) indicating that the driver's line of sight direction ED1 is appropriate. .
- the annotator specifies, for example, a risk area RA in which a person depicted in the converted image, excluding the person who has been subjected to mosaic processing, is predicted to proceed to the converted image of the outside-of-vehicle image. Because the face of the person in the original image is converted into the face of another person through the processing by the image conversion unit 140 and the image determination unit 150, the privacy of the person is protected. At the same time, since the person's face direction and line of sight direction are maintained even after conversion, the annotator can accurately specify the risk area RA while referring to the face direction and line of sight direction of another person captured in the converted image. be able to. Thereby, it is possible to generate learning data that is effective for learning a machine learning model while protecting the privacy of the person depicted in the facial image.
- a risk area RA in which a person depicted in the converted image, excluding the person who has been subjected to mosaic processing, is predicted to proceed to the converted image of the outside-of-vehicle
- the learned model generation unit 160 uses the annotated image data 178 as learning data to generate a learned model using an arbitrary machine learning model.
- this trained model outputs the predicted behavior (trajectory) of a person depicted in the image outside the vehicle in response to an input of an image outside the vehicle, or outputs the predicted behavior (trajectory) of a person shown in the image outside the vehicle, and This is a behavior prediction model that takes into consideration the line of sight of the driver shown in the image inside the car and calls attention to pedestrians shown in the image outside the car.
- the trained model generation unit 160 stores the generated trained model in the storage unit 170 as a trained model 180.
- the transmission/reception control unit 120 distributes the generated learned model 180 to the vehicle M2 via the NetArc NW.
- the vehicle M2 uses the learned model 180 (more precisely, an application program that utilizes the learned model 180) to provide driving support to the driver of the vehicle M2.
- FIG. 9 is a diagram showing an example of driving support using the learned model 180.
- the inside image and the outside image taken by the camera mounted on the vehicle are input to the trained model 180, and the trained model 180 calculates the line of sight of the driver captured in the inside image.
- this example shows an example in which driving assistance is provided by outputting information to a human machine interface (HMI) to urge the attention of pedestrians shown in the image outside the vehicle.
- HMI human machine interface
- the HMI displays a risk area RA2 corresponding to the pedestrian P5 shown in the image outside the vehicle, and the driver's line of sight shown in the inside image is not directed toward the pedestrian P5.
- a warning message (“Be careful of distracted driving”) is output as text or audio information.
- Be careful of distracted driving is output as text or audio information.
- FIG. 10 is a diagram illustrating an example of the flow of processing executed by the image conversion unit 140.
- the process shown in FIG. 10 is executed, for example, at the timing when an in-vehicle image or an out-vehicle image is captured by a camera mounted on the vehicle M1 and processed by the image processing unit 130.
- the image conversion unit 140 acquires a captured image included in the captured image data 172 processed by the image processing unit 130 (step S100). Next, the image conversion unit 140 selects one face shown in the acquired captured image (step S102).
- the image conversion unit 140 determines whether the size of the selected face is greater than or equal to the first threshold Th1 (step S104). If it is determined that the size of the selected face is greater than or equal to the first threshold Th1, the image conversion unit 140 converts the face into the face of another person (step S106). On the other hand, if it is determined that the size of the selected face is less than the first threshold Th1, the image conversion unit 140 next determines whether the distance of the selected face is less than or equal to the second threshold Th2. (Step S108).
- step S106 converts the face into the face of another person.
- step S110 the image conversion unit 140 performs mosaic processing on the face.
- step S112 the image conversion unit 140 determines whether the processing has been performed on all faces shown in the acquired captured image.
- the image conversion unit 140 obtains an image obtained by performing the processing on all the faces as a converted image.
- the converted image data 174 is then stored in the storage unit 170 (step S114).
- the image conversion unit 140 returns the processing to step S102. This completes the processing of this flowchart.
- FIG. 11 is a diagram showing an example of the flow of processing executed by the image determination unit 150.
- the process shown in FIG. 11 is, for example, the timing at which the time-series converted images are obtained by performing the above conversion process on the time-series captured images taken in one driving cycle from the start to the stop of the vehicle M1. It is executed in
- the image determination unit 150 acquires time-series converted images (step S200). Next, the image determination unit 150 selects the face of a person who was tracked as the same person before conversion in the acquired time-series converted images (step S202).
- step S206 if it is determined that the acquired time-series converted images are in-vehicle images, the image determination unit 150 determines whether the gaze direction and face direction of these faces match the images before conversion. Determination is made (step S210). On the other hand, if it is determined that the acquired time-series converted images are not in-vehicle images, that is, they are outside-vehicle images, the image determining unit 150 determines whether the facial directions of these faces match the images before conversion. (Step S212). If it is determined in the process of step S210 or step S212 that they do not match, the image determination unit 150 advances the process to step S208.
- the image determination unit 150 determines that these faces have been converted normally, and applies all the faces shown in the time-series converted images. It is determined whether the process has been executed (step S214). If it is determined that the processing has been performed on all faces captured in the time-series converted images, the image determination unit 150 acquires these time-series converted images as annotation images and sends them to the transmission/reception control unit 120. The acquired annotation image is transmitted to the terminal device 200 (step S216). On the other hand, if it is determined that the process has not been performed on all faces captured in the time-series converted images, the image determination unit 150 returns the process to step S202. This completes the processing of this flowchart.
- the anonymization process includes a process of changing the face of a person shown in the plurality of input images to the face of another person, and the predetermined requirement is that the face of the person shown in the plurality of input images is changed to the face of a different person.
- the faces of people tracked as the same person are faces of the same person obtained through anonymization processing. That is, in this embodiment, faces that belong to the same person before the anonymization process are guaranteed to be the same person's faces even after the anonymization process, and are used as learning data. Thereby, it is possible to generate learning data that is effective for learning a machine learning model while protecting the privacy of the person depicted in the facial image.
- the predetermined requirements include face direction information of a person tracked as the same person in a plurality of input images, and face direction information of the same person in the plurality of input images subjected to the anonymization process.
- the predetermined requirements are determined according to the image attributes that are the shooting modes of the plurality of input images. That is, in this embodiment, a predetermined process, which is, for example, a process of saving learning information for generating a behavior prediction model, is executed in consideration of the shooting mode of each of the plurality of input images. Thereby, it is possible to generate learning data that is effective for learning a machine learning model while protecting the privacy of the person depicted in the facial image.
- the anonymization process is performed by the first method or by the first method. determines whether to perform anonymization processing using a different second method. That is, in this embodiment, the method of anonymization processing performed on a face is changed depending on whether it is useful for learning a machine learning model. Thereby, it is possible to generate learning data that is effective for learning a machine learning model while protecting the privacy of the person depicted in the facial image.
- the image determination unit 150 determines that the face depicted in the converted image does not meet predetermined requirements, an example will be described in which the converted image is re-converted or subjected to mosaic processing. did. However, if the image determining unit 150 determines that the predetermined requirements are not met, the image determining unit 150 does not perform the predetermined processing on the converted image, that is, it restricts the predetermined processing (the image processing such as not storing or transmitting to the server, etc.) may also be performed.
- the image processing device 100 is implemented as a server device separate from the vehicle M1.
- the image processing device 100 more specifically, a device having the functions of at least the image processing section 130, the image conversion section 140, and the image determination section 150 is mounted on the vehicle M1 as an on-vehicle device. It's okay.
- the image captured by the vehicle-mounted camera is processed by the image processing unit 130 described above, anonymized by the image conversion unit 140, and determined by the image determination unit 150. Thereafter, the in-vehicle device transmits the anonymized image, in which the continuity of the face and the consistency of the direction information have been confirmed by the image determination unit 150, to an external image server.
- the image server Upon receiving the anonymized image from vehicle M1, the image server stores the received anonymized image in the storage unit as annotation image data, and transmits the annotation image data to the annotator's terminal device 200, or 200 is permitted to access the annotation image data.
- the image server Upon receiving the annotated image data from the terminal device 200, the image server generates a learned model 180 based on the annotated image data, and distributes the generated learned model 180 to the vehicle M2. Even in this case, similar to the present embodiment, learning data effective for learning the machine learning model can be generated while protecting the privacy of the person photographed in the face image.
- the in-vehicle device since the in-vehicle device performs anonymization processing on the image and then sends the anonymized image to the image server, it is possible to more reliably protect the privacy of the person depicted in the facial image. can.
- the in-vehicle device may include only some of the functions of the image processing section 130, the image conversion section 140, and the image determination section 150, and the image server may have the remaining functions.
- the in-vehicle device may have the functions of the image processing section 130 and the image conversion section 140, and the image server may have the function of the image determination section 150, or the in-vehicle device may have the functions of the image processing section 130, and the image server The functions of the image conversion section 140 and the image determination section 150 may be provided.
- a storage medium for storing computer-readable instructions; a processor connected to the storage medium; the processor executing the computer-readable instructions to: Obtain the image attribute that is the shooting mode of the input image, Performing anonymization processing on the input image, determining whether the input image subjected to the anonymization process satisfies predetermined requirements; If it is determined that the input image subjected to the anonymization process satisfies the predetermined requirements, performing a predetermined process on the input image subjected to the anonymization process,
- the image processing apparatus is configured such that the predetermined requirement is determined according to the acquired image attribute.
- Image processing device 110 Communication unit 120 Transmission/reception control unit 130 Image processing unit 140 Image conversion unit 150 Image determination unit 160 Learned model generation unit 170 Storage unit 172 Captured image data 174 Converted image data 176 Annotation image data 178 Annotated image data 180 Trained model
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- Computing Systems (AREA)
- Evolutionary Computation (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Software Systems (AREA)
- Databases & Information Systems (AREA)
- Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Multimedia (AREA)
- Image Analysis (AREA)
- Image Processing (AREA)
- Facsimile Image Signal Circuits (AREA)
Abstract
Description
(1):この発明の一態様に係る画像処理装置は、入力画像の撮影態様である画像属性を取得する画像属性取得部と、前記入力画像に対して匿名化処理を行う画像変換部と、前記匿名化処理が施された前記入力画像が所定要件を満たすか否かを判定する画像判定部と、を備え、前記画像判定部は、前記匿名化処理が施された前記入力画像が前記所定要件を満たすと判定した場合、前記匿名化処理が施された前記入力画像に所定処理を施し、前記所定要件は、取得された前記画像属性に応じて定められるものである。
図1は、本実施形態に係る画像処理装置100を含むシステム1の概要を示す図である。図1に示す通り、システム1は、それぞれが少なくとも一台以上の車両M1および車両M2と、画像処理装置100と、端末装置200とを含む。説明の便宜上、車両M1および車両M2とを異なる車両として図示しているが、これらの車両は同一であっても良い。
図2は、本実施形態に係る画像処理装置100の機能構成の一例を示す図である。画像処理装置100は、例えば、通信部110と、送受信制御部120と、画像処理部130と、画像変換部140と、画像判定部150と、学習済みモデル生成部160と、記憶部170と、を備える。これらの構成要素は、例えば、CPU(Central Processing Unit)などのハードウェアプロセッサがプログラム(ソフトウェア)を実行することにより実現される。これらの構成要素のうち一部または全部は、LSI(Large Scale Integration)やASIC(Application Specific Integrated Circuit)、FPGA(Field-Programmable Gate Array)、GPU(Graphics Processing Unit)などのハードウェア(回路部;circuitryを含む)によって実現されてもよいし、ソフトウェアとハードウェアの協働によって実現されてもよい。プログラムは、予めHDD(Hard Disk Drive)やフラッシュメモリなどの記憶装置(非一過性の記憶媒体を備える記憶装置)に格納されていてもよいし、DVDやCD-ROMなどの着脱可能な記憶媒体(非一過性の記憶媒体)に格納されており、記憶媒体がドライブ装置に装着されることでインストールされてもよい。記憶部170は、例えば、HDDやフラッシュメモリ、RAM(Random Access Memory)等である。記憶部170は、例えば、撮像画像データ172と、変換画像データ174と、アノテーション用画像データ176と、アノテーション付画像データ178と、学習済みモデル180とを記憶する。なお、説明の便宜上、画像処理装置100は、学習済みモデル生成部160と、学習済みモデル180を記憶する記憶部170とを備えているが、学習済みモデルを生成する機能と、生成した学習済みモデルとは、画像処理装置100とは異なるサーバ装置が保有してもよい。
上述した通り、本実施形態では、画像判定部150によって、変換画像に写される顔が所定要件を満たさないと判定された場合、当該変換画像を再変換したり、モザイク処理を施す例について説明した。しかし、画像判定部150によって所定要件が満たされていないと判定された場合には、画像判定部150は変換画像に対して所定処理を施さない、すなわち、所定処理を施すことを制限する(画像保管、サーバへの送信等をしない)といった処理を行ってもよい。
コンピュータによって読み込み可能な命令(computer-readable instructions)を格納する記憶媒体(storage medium)と、
前記記憶媒体に接続されたプロセッサと、を備え、
前記プロセッサは、前記コンピュータによって読み込み可能な命令を実行することにより(the processor executing the computer-readable instructions to:)
入力画像の撮影態様である画像属性を取得し、
前記入力画像に対して匿名化処理を行い、
前記匿名化処理が施された前記入力画像が所定要件を満たすか否かを判定し、
前記匿名化処理が施された前記入力画像が前記所定要件を満たすと判定した場合、前記匿名化処理が施された前記入力画像に所定処理を施し、
前記所定要件は、取得された前記画像属性に応じて定められるものであるように構成されている、画像処理装置。
110 通信部
120 送受信制御部
130 画像処理部
140 画像変換部
150 画像判定部
160 学習済みモデル生成部
170 記憶部
172 撮像画像データ
174 変換画像データ
176 アノテーション用画像データ
178 アノテーション付画像データ
180 学習済みモデル
Claims (13)
- 入力画像の撮影態様である画像属性を取得する画像属性取得部と、
前記入力画像に対して匿名化処理を行う画像変換部と、
前記匿名化処理が施された前記入力画像が所定要件を満たすか否かを判定する画像判定部と、を備え、
前記画像判定部は、前記匿名化処理が施された前記入力画像が前記所定要件を満たすと判定した場合、前記匿名化処理が施された前記入力画像に所定処理を施し、
前記所定要件は、取得された前記画像属性に応じて定められるものである、
画像処理装置。 - 前記所定処理は、前記匿名化処理が施された前記入力画像をアノテーション作業の対象画像として保存する処理である、
請求項1に記載の画像処理装置。 - 前記所定処理は、前記匿名化処理が施された前記入力画像を、入力画像に写される人物の行動を予測する行動予測モデルを生成するための学習用情報として保存する処理である、
請求項1に記載の画像処理装置。 - 前記所定処理は、前記匿名化処理が施された前記入力画像を、通信手段を通じて画像サーバに送信する処理である、
請求項1に記載の画像処理装置。 - 前記画像属性は、前記入力画像が、前記入力画像を撮像したカメラが搭載された車両の内部を撮像した画像か、前記車両の外部を撮像した画像かを少なくとも示す情報である、
請求項1に記載の画像処理装置。 - 前記匿名化処理は、前記入力画像に写される人物の顔を別人物の顔に変更する処理を含み、
前記所定要件は、前記画像属性が前記車両の内部を撮像した画像であることを示すものである場合、前記人物の顔の視線方向と前記別人物の顔の視線方向とが一致するか否かを含むものである、
請求項5に記載の画像処理装置。 - 前記所定要件は、前記画像属性が前記車両の内部を撮像した画像であることを示すものである場合、前記人物の顔の視線方向と前記別人物の顔の視線方向とが一致するか否かと、前記人物の顔の顔方向と前記別人物の顔の顔方向とが一致するか否かと、を含むものであり、
前記所定要件は、前記画像属性が前記車両の外部を撮像した画像であることを示すものである場合、前記人物の顔の顔方向と前記別人物の顔の顔方向とが一致するか否かを含むものである、
請求項6に記載の画像処理装置。 - 前記所定要件は、前記画像属性が前記車両の外部を撮像した画像であることを示すものである場合、前記人物の顔の視線方向と前記別人物の顔の視線方向とが一致するか否かを含まないものである、
請求項6に記載の画像処理装置。 - 前記画像変換部は、前記画像判定部によって、前記匿名化処理が施された前記入力画像が前記所定要件を満たさないと判定した場合、前記入力画像に前記匿名化処理を再度、施す、
請求項1に記載の画像処理装置。 - 前記画像変換部は、前記画像判定部によって、前記匿名化処理が施された前記入力画像が前記所定要件を満たさないと判定した場合、前記匿名化処理が施された前記入力画像に前記所定処理を施さない、
請求項1に記載の画像処理装置。 - 入力画像の撮影態様である画像属性を取得する画像属性取得部と、
前記入力画像に対して匿名化処理を行う画像変換部と、
前記匿名化処理が施された前記入力画像が所定要件を満たすか否かを判定する画像判定部と、を備え、
前記画像判定部は、前記匿名化処理が施された前記入力画像が前記所定要件を満たすと判定した場合、前記匿名化処理が施された前記入力画像に所定処理を施し、
前記所定要件は、取得された前記画像属性に応じて定められるものである、
画像処理システム。 - コンピュータが、
入力画像の撮影態様である画像属性を取得し、
前記入力画像に対して匿名化処理を行い、
前記匿名化処理が施された前記入力画像が所定要件を満たすか否かを判定し、
前記匿名化処理が施された前記入力画像が前記所定要件を満たすと判定した場合、前記匿名化処理が施された前記入力画像に所定処理を施し、
前記所定要件は、取得された前記画像属性に応じて定められるものである、
画像処理方法。 - コンピュータに、
入力画像の撮影態様である画像属性を取得させ、
前記入力画像に対して匿名化処理を行わせ、
前記匿名化処理が施された前記入力画像が所定要件を満たすか否かを判定させ、
前記匿名化処理が施された前記入力画像が前記所定要件を満たすと判定した場合、前記匿名化処理が施された前記入力画像に所定処理を施させ、
前記所定要件は、取得された前記画像属性に応じて定められるものである、
プログラム。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202380049341.5A CN119422165A (zh) | 2022-06-30 | 2023-06-29 | 图像处理装置、图像处理方法、图像处理系统及程序 |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2022106676A JP7805260B2 (ja) | 2022-06-30 | 2022-06-30 | 画像処理装置、画像処理方法、画像処理システム、およびプログラム |
| JP2022-106676 | 2022-06-30 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024005160A1 true WO2024005160A1 (ja) | 2024-01-04 |
Family
ID=89382512
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2023/024253 Ceased WO2024005160A1 (ja) | 2022-06-30 | 2023-06-29 | 画像処理装置、画像処理方法、画像処理システム、およびプログラム |
Country Status (3)
| Country | Link |
|---|---|
| JP (1) | JP7805260B2 (ja) |
| CN (1) | CN119422165A (ja) |
| WO (1) | WO2024005160A1 (ja) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2014085796A (ja) * | 2012-10-23 | 2014-05-12 | Sony Corp | 情報処理装置およびプログラム |
| JP2017187850A (ja) * | 2016-04-01 | 2017-10-12 | 株式会社リコー | 画像処理システム、情報処理装置、プログラム |
| WO2019146357A1 (ja) * | 2018-01-24 | 2019-08-01 | 富士フイルム株式会社 | 医療画像処理装置、方法及びプログラム並びに診断支援装置、方法及びプログラム |
| JP2020061081A (ja) * | 2018-10-12 | 2020-04-16 | キヤノン株式会社 | 画像処理装置および画像処理方法 |
| JP2020170496A (ja) * | 2019-04-04 | 2020-10-15 | ▲広▼州大学 | 顔認識用の年齢プライバシー保護方法及びシステム |
-
2022
- 2022-06-30 JP JP2022106676A patent/JP7805260B2/ja active Active
-
2023
- 2023-06-29 CN CN202380049341.5A patent/CN119422165A/zh active Pending
- 2023-06-29 WO PCT/JP2023/024253 patent/WO2024005160A1/ja not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2014085796A (ja) * | 2012-10-23 | 2014-05-12 | Sony Corp | 情報処理装置およびプログラム |
| JP2017187850A (ja) * | 2016-04-01 | 2017-10-12 | 株式会社リコー | 画像処理システム、情報処理装置、プログラム |
| WO2019146357A1 (ja) * | 2018-01-24 | 2019-08-01 | 富士フイルム株式会社 | 医療画像処理装置、方法及びプログラム並びに診断支援装置、方法及びプログラム |
| JP2020061081A (ja) * | 2018-10-12 | 2020-04-16 | キヤノン株式会社 | 画像処理装置および画像処理方法 |
| JP2020170496A (ja) * | 2019-04-04 | 2020-10-15 | ▲広▼州大学 | 顔認識用の年齢プライバシー保護方法及びシステム |
Also Published As
| Publication number | Publication date |
|---|---|
| JP7805260B2 (ja) | 2026-01-23 |
| JP2024006102A (ja) | 2024-01-17 |
| CN119422165A (zh) | 2025-02-11 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11606516B2 (en) | Image processing device, image processing method, and image processing system | |
| US11042999B2 (en) | Advanced driver assist systems and methods of detecting objects in the same | |
| US11481913B2 (en) | LiDAR point selection using image segmentation | |
| EP3340205B1 (en) | Information processing device, information processing method, and program | |
| US20210287035A1 (en) | Image and lidar segmentation for lidar-camera calibration | |
| CN114677754B (zh) | 行为识别方法、装置、电子设备及计算机可读存储介质 | |
| JP6601506B2 (ja) | 画像処理装置、物体認識装置、機器制御システム、画像処理方法、画像処理プログラム及び車両 | |
| JP2009070344A (ja) | 画像認識装置、画像認識方法および電子制御装置 | |
| CN111731304B (zh) | 车辆控制装置、车辆控制方法及存储介质 | |
| WO2019111529A1 (ja) | 画像処理装置および画像処理方法 | |
| CN115408710A (zh) | 一种图像脱敏方法和相关装置 | |
| WO2024005073A1 (ja) | 画像処理装置、画像処理方法、画像処理システム、およびプログラム | |
| US11308324B2 (en) | Object detecting system for detecting object by using hierarchical pyramid and object detecting method thereof | |
| US20230109494A1 (en) | Methods and devices for building a training dataset | |
| CN112241963A (zh) | 基于车载视频的车道线识别方法、系统和电子设备 | |
| Dai et al. | [Retracted] An Intelligent Security Classification Model of Driver’s Driving Behavior Based on V2X in IoT Networks | |
| WO2024005160A1 (ja) | 画像処理装置、画像処理方法、画像処理システム、およびプログラム | |
| WO2024005074A1 (ja) | 画像処理装置、画像処理方法、画像処理システム、およびプログラム | |
| WO2024005051A1 (ja) | 画像処理装置、画像処理方法、画像処理システム、およびプログラム | |
| WO2024180706A1 (ja) | 画像処理装置、画像処理方法、およびプログラム | |
| EP4374335A1 (en) | Electronic device and method | |
| EP4561085A1 (en) | Information processing device and information processing method | |
| WO2024180709A1 (ja) | 画像処理装置、画像処理方法、およびプログラム | |
| WO2024180759A1 (ja) | 画像処理装置、画像処理方法、およびプログラム | |
| CN120340288A (zh) | 规划车辆行驶路径的方法、装置、存储介质及电子设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23831599 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202380049341.5 Country of ref document: CN |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| WWP | Wipo information: published in national office |
Ref document number: 202380049341.5 Country of ref document: CN |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 23831599 Country of ref document: EP Kind code of ref document: A1 |