EP4690125A1 - Method for generating at least one synthetic image - Google Patents

Method for generating at least one synthetic image

Info

Publication number
EP4690125A1
EP4690125A1 EP24716763.8A EP24716763A EP4690125A1 EP 4690125 A1 EP4690125 A1 EP 4690125A1 EP 24716763 A EP24716763 A EP 24716763A EP 4690125 A1 EP4690125 A1 EP 4690125A1
Authority
EP
European Patent Office
Prior art keywords
image
image section
contour
point
processing operation
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24716763.8A
Other languages
German (de)
French (fr)
Inventor
Muhammad Zeeshan Karamat
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
36zero Vision GmbH
Original Assignee
36zero Vision GmbH
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by 36zero Vision GmbH filed Critical 36zero Vision GmbH
Publication of EP4690125A1 publication Critical patent/EP4690125A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T11/00Two-dimensional [2D] image generation
    • G06T11/60Creating or editing images; Combining images with text

Definitions

  • the invention relates to a method for generating at least one synthetic image. Additionally, the invention relates to a data processing device comprising means for carrying out the method. In addition, the invention relates to a computer program product, a computer readable medium having stored thereon the computer program product and a data carrier signal carrying the computer program product.
  • neural networks in particular of convolutional neural networks, in different kind of applications.
  • Each of said neural networks has to be trained in a training process to provide accurate results in an operation process.
  • Operation process means the situation when new data, in particular images, that is not used in the training process is inputted into the neural network, wherein the neural network outputs a result by using said inputted data.
  • training data like images are used to train the neural network.
  • the accuracy of the output of the neural network depends on the amount and the quality of information of training data that is inputted into the neural network in the training process.
  • the generalizability of use of the neural network also depends on the amount of training data that is inputted into the neural network.
  • Generalizability means the ability of the neural network to provide accurate results in different kind of applications, including also applications for which the neural network is not specifically trained for.
  • Accuracy is the number of correctly predicted data points out of all the data points. More formally, it is defined as the number of true positives and true negatives divided by the number of true positives, true negatives, false positives, and false negatives.
  • a true positive or true negative is a data point that the algorithm correctly classified as true or false, respectively.
  • a false positive or false negative is a data point that the algorithm incorrectly classified.
  • an augmentation process can be applied on the image to generate one or more synthetic images.
  • the augmentation process can comprise one or more of the following processing changing the color of the image, changing the brightness of the image, changing the contrast of the image and rotating the image.
  • the number of generated synthetic images is often not sufficient to achieve a satisfactory increase in accuracy and generalizability of the neural network that is needed for the specific application.
  • the object of the invention is to increase the number of images having high quality that are used for training a neural network such that it is accurate and that it has a high generalizability.
  • the object is solved by a method for generating at least one synthetic image, wherein the method comprises the following steps: receiving of at least one image, processing at least one selected image section in at least one processing operation, wherein the execution of the at least one processing operation depends on a probability factor and generating at least one synthetic image comprising the processed at least one image section.
  • the method in particular computer implemented method, has the advantage that a large number of synthetic images can be generated.
  • the generated images can be used in a training process for training a neural network, in particular a convolutional neural network. Due to the large number of synthetic images, it is possible to train the neural network such that it provides accurate results.
  • a further advantage is that the synthetic images have a good quality. Quality means that the image show the information that is needed for the specific application. This is advantageous as it is sometimes hard to get real data in a sufficient amount.
  • the inventive method the number of real images taken by an image acquisition device can be low. This is possible as according to the invention a high number of synthetic images can be provided having a lot of variations so that the neural network can be better trained.
  • an overfitting of the neural network can be avoided.
  • synthetic images can be generated that have many variations of pattern, shape, color, placement, etc. in the selected image part.
  • the invention also enables to train a neural network with a small data set, i.e. a low number of real images taken by the image acquisition device. This is possible as due to the inventive method a large number of synthetic images can be generated.
  • the image processing shall depend on a probability factor to avoid generating a high number of synthetic images that are not useful for training purposes but results in high computation training costs.
  • a high number of synthetic images is preferred for training a neural network.
  • processing of the at least one image section by all theoretical possible processing operations that are described below more in detail results in synthetic images that are not useful for training a neural network. This is the case as the processing can lead to a synthetic image showing an object that would not appear in reality.
  • an image of the reality is acquired by an image acquisition device.
  • the acquired image of the reality can be labeled.
  • the image acquisition device can be any component comprising optical means by means of which an image can be acquired.
  • the image acquisition device can be a camera and/or mobile phone and/or tablet and/or a microscope, etc..
  • the image acquisition device can be configured to acquire signals.
  • the signals can have a wavelength which is in a human visible range.
  • the image acquisition device can acquire light comprising a wavelength of 380nm to 780nm (nanometer).
  • the image acquisition device can also acquire signals having a wavelength that is outside the human visible range.
  • the synthetic image is not acquired by the image acquisition device but generated by the data processing device.
  • a processed synthetic image is an image which is processed by at least one processing operation.
  • the synthetic image can show an artificial state that differs from the reality acquired by the image acquisition device.
  • the synthetic image is based on the acquired image.
  • the synthetic image can show a state that differs from the reality shown in the image.
  • the image can comprise a target object or several target objects.
  • the target object can be an object having a predetermined physical property.
  • the object can be a discrete object with well-defined boundaries and spatial extension.
  • the object can be anything that is visible for a human and/or tangible and/or that can be touched.
  • the target object can be a car, chair, scratch, etc..
  • the object can be represented by one or more pixels of the image.
  • the object can be a digital pattern having a predetermined information type.
  • the information type can define whether the digital pattern is visible and/or tangible and/or can be touched.
  • the target object can cover objects being nontangible and/or non-touchable and/or non-visible in the human light range.
  • the target object can be image portions consisting of one or more pixels having information about for example a specific temperature, reflectance, radiance, etc..
  • the probability factor is predetermined. It can be stored in an electrical memory, in particular a memory of the data processing device. The probability factor can be determined from experiments before the generation of the synthetic image is started.
  • the at least one image section can be automatically selected by the data processing device. Additionally or alternatively the at least one image section can be selected by a user. In that case the selected at least one image section is transmitted to the data processing device and/or the data processing device receives the selected at least one image section. This is explained below more in detail.
  • the selection of the image section is done such that the location and/or shape of the image section is known. In particular, the data processing device knows the location and shape of each of the selected image section.
  • the image section is a part of the image.
  • image sections can be automatically selected by the data processing device. Alternatively or additionally, several image sections can be selected by a user. In that case the selected image sections are transmitted to the data processing device and/or the data processing device receives the selected image sections. In particular, the location and/or the shape of each of the image sections is transmitted to the data processing device and/or received by the data processing device.
  • Each of said several image sections can be processed by a processing operation, in particular the same processing operation. The selection of several image sections can be done when the image has more target objects. Thus, each image section can be assigned to one target object.
  • the method can be executed in a data processing device.
  • the data processing device can comprise one or more processors or can be a processor.
  • the data processing device can be a computer.
  • the image data obtained by the image acquisition device is sent to the data processing device.
  • the data processing device receives said image data from the image acquisition device and processes said image data.
  • the data processing device can generate several synthetic images. Said several synthetic images can be based on the same received image. That means, the image received from the data processing device results in several synthetic images. Additionally or alternatively the data processing device can receive several images from the image acquisition device. For each of said received images one or more synthetic images can be generated.
  • at least one selected image section can be a polygon that has a closed contour.
  • the contour of the polygon can depend on a contour of a target object. That means, the contour of the polygon is chosen such that it corresponds to the contour of the target object. Using a polygon enables to create the closed contour that is needed and/or wished by the user. Contour means the outer boundary or rim of the image section, in particular the polygon.
  • a polygon is a shape in geometry that has closed a structure.
  • the polygon can comprise three or more corners.
  • the number of corners depends on the shape of the target object. This is the main difference of using image sections being a polygon to the use of rectangular boxes as image sections that are used in known methods.
  • the shape of the rectangular box used in the prior art is always the same independent of the form of the contour of the target object.
  • Another advantage of using a polygon as image section is that the data processing device automatically knows the target object contour and the location of the target object. This results as the contour of the polygon corresponds to the contour of the target object. Thus, there is no need for any other objection detection methods for detecting the location and contour of the target object.
  • the image can have at least one target object.
  • the image can have several target objects.
  • the selection of the image section is done such that that the target object is arranged in the selected image section. That means, the image selection is located and shaped such that the target object is arranged inside the image section.
  • the image section selection can be done such that the contour of the image section corresponds to the contour of the target object. This can be achieved by using a polygon having a contour which is adapted to fit to the contour of the target object. Contour of the target object means the outer boundary or rim of the target object.
  • the image selection can be a polygon.
  • Such kind of labeling is also indicated as "weak labeling".
  • the labeling is simplified as it has not to be assured that the contour of the image section exactly matches with the contour of the target object.
  • the accuracy of the artificial neural network that is trained using the synthetic image it is preferred to minimize the number of pixels within the image section that do not belong to the target object.
  • polygons has the advantage that the labelling can be improved as the image section comprises or mainly comprises the target object.
  • the accuracy of the neural network which is used by the labelled images can be improved.
  • points characterizing the target object can be easily determined by using polygon. Using said points enables to provide more realistic synthetic images. This is possible because by using a polygon it is ensured that only or mainly the points and thus the target object is processed by a processing operation.
  • the artificial neural network to which the synthetic image is input for training only receives information about the target object.
  • the target object can be a discrete object and/or represented by one or more pixels.
  • the target object can be labeled when the image section is selected. Labeling means that the target object is named. However, it is possible to label the target object after the image section is processed in the processing operation. Additionally or alternatively, further information of the target object, in particular the instance of the target object, can be provided. Thus, at the end of the image section selection the data processing device has information about the location and/or shape and/or name and/or further information of the target object.
  • the data processing device can determine the probability factor by using a random algorithm.
  • the random algorithm is started and the algorithm result is compared with the probability coefficient. If the algorithm result is lower than the probability factor the decision factor is set to 0 and if the algorithm result is greater than the probability factor, the decision factor is set to 1.
  • the data processing device can perform the processing operation dependent on the decision factor. If the decision factor has the value 0, the data processing device does not perform the processing operation. However, if the decision factor has the value 1, the data processing device performs the processing operation. This process is repeated for each probability factor.
  • a set of decision factors is determined wherein the set of decision factors comprises one or more decision factors.
  • the processing operation is performed only in the selected image section dependent on said at least one decision factor. As the selected image section corresponds to the target object, the processing operation is applied on the target object.
  • the generation of the different shaped target object is explained more in detail.
  • several points are determined on the basis of the contour points of the contour of the image section, in particular the polygon.
  • the determined points are connected with each other.
  • the connection is made such that the connection result represents the target object. That means, that each of the determined points can be connected with one or more other determined points.
  • a point can but does not have to be connected to all other determined points.
  • a skeleton of the target object is created. Said skeleton simplifies to determine which of the determined points can be moved and/or how it is moved.
  • a data processing device comprising means for carrying out an inventive method.
  • an image acquisition device for acquiring images is provided.
  • the image acquisition device can be at least one of the following a camera, a mobile phone, microscope and a tablet.
  • the data processing device can be part of the image acquisition device.
  • the data processing device can be electrically connected to the image acquisition device.
  • the acquisition device can be configured to acquire visible light. "Visible light” means that the acquired light has a wavelength in the range of 380 to 780 nanometers.
  • the image acquisition device can be configured to acquire non-visible light.
  • a computer program product comprising instructions which, when the program is executed by the data processing device, in particular a computer, cause the data processing device, in particular the computer, to carry out the steps of the inventive method.
  • a computer-readable data carrier is provided wherein the computer-readable data carrier has stored thereon the computer program product.
  • a data carrier signal is provided wherein the data carrier signal carries the computer program product.
  • Fig. 4 a system comprising a data processing device and an image acquisition device according to a second embodiment.
  • Fig. 6 a flow chart of generating a synthetic image.
  • Fig. 7 a probability vector and matrix for decision factors.
  • Fig. 8A a flow chart of a processing operation according to a first embodiment.
  • Fig. 9A a flow chart of a processing operation according to a second embodiment.
  • Fig. 9C the image shown in fig. 9B in which contour points of the contour of the image section are shown.
  • Fig. 9D an image in which points and their connection lines are shown.
  • Fig. 10A a flow chart of a processing operation according to a third embodiment.
  • An image 2 as shown in Fig. 1 is acquired by an image acquisition device 11 shown in fig. 3 and 4.
  • a data processing device 9 shown in fig. 3 and 4 receives said image and processes at least one image section 3 in at least one processing operation.
  • the data processing device 9 and/or a user selects two image sections 3 within a region of interest 5 of the image 1 .
  • the location and/or the shape of the image sections 3 are randomly selected.
  • two image sections 3 are selected that have the same shape. However, in other non-shown embodiments, the number and shape of selected image sections can differ.
  • the data processing device 9 has information about the location and shape of said image sections 3.
  • the execution of the at least one processing operation by the data processing device 9 depends on a probability factor. In particular, only the image section is processed by the processing operation. That means, the remaining part of the image is not processed by the processing operation.
  • the data processing device 9 generates at least one synthetic image comprising the processed at least one image section.
  • the target object has a rectangular shape comprising four corners
  • the polygon has a polygon shape with four corners.
  • the image section in particular the polygon, has a circular shape.
  • the contour of the image section on the contour of target object 4 so that a small distance might exist between the contour of the target object 4 and the contour of the image section 4. Said distance should be kept as small as possible and in the ideal case should not be existent.
  • the location and the shape of the image section 3 is dependent on the location and shape of the target object 4. Additionally, further information on the target object 4 can be provided to the data processing unit.
  • the name and/or instance of the target object 4 can be provided to the data processing device 9.
  • the data processing device 9 has information about the location, shape and further information of the image section 3 and thus of the target object 4. Said information are assigned to the selected image section, respectively.
  • the data processing device 9 generates on the basis of the received image one or more synthetic images. Thereto, one or more image sections 3 are selected. In the embodiment shown in fig. 3 the data processing devices 9 automatically selects the image sections 3. For the case that the image 2 does not comprise a target object 4, the data processing device 9 randomly selects the image sections 3. However, if the image 1 has at least one target object 4, the data processing device 9 detects the target object 4 and places the contour of the image section 3, in particular creates a polygon with a corresponding contour, so that it matches with the contour of the target object 4. The generated synthetic images are used to train a neural network 10. The neural network can be trained by the data processing device 9 or another processing device not shown in the figures.
  • Fig. 5 shows a flow chart of the general method.
  • the image acquisition device 11 acquires an image 2.
  • Said image is transmitted to the data processing device 9.
  • the image section 3 is selected. This is done be creating a polygon having a contour that depends on the contour of the target object and locating the contour such that the target object is arranged in the polygon.
  • the image section selection can be done by the data processing device 9 and/or by a user. In the latter case, the information about the selected image section 3 are transmitted to the data processing device 9.
  • the data processing device 9 has information about at least the shape and location of all selected image sections.
  • a third step G3 one or more synthetic images are generated on the basis of the received image.
  • the data processing device replaces or moves the selected image section 3 in a fourth step G4.
  • the image sections that are processed in the third step G3 are placed to a new location within the image or within the region of interest 5. The result of said replacement is that further synthetic images are created.
  • a fifth step G5 the data processing device 9 performs an augmentation process on the image or on the region of interest 5.
  • the data processing device 9 changes the brightness of the at least one image generated in the third step G3 and/or the contrast of the at least one image generated in the third step G3 and/or the color of the at least one image generated in the third step G3.
  • the result of the fifth step G5 is that further synthetic images are generated.
  • a first substep S1 the data processing device 9 determines all image sections present in the received image and that have been selected in the second step G2 shown in Fig. 5. The processing operations discussed below are performed in all determined image sections.
  • the data processing device 9 sets the probability factors.
  • the probability factors indicate the likelihood that a processing operation shall be performed on the image section 3.
  • the number of probability factors corresponds to the number of processing operations by means of which the selected image section 3 can be processed.
  • the probability factors are predetermined before the method is executed and thus the image 2 is acquired.
  • the probability factors are stored in an electrical memory of the data processing device. Alternatively, the user can enter the probability factors.
  • the probability factors can have a value between 0 and 1.
  • a probability factor vector 13 is present wherein the vector comprises the probability factor for each processing operation. Each processing operation is assigned to a probability factor.
  • the probability factor vector 13 comprises seven probability factors. This means, the data processing device 9 can process the image 2 in seven different ways as the processing operations differ from each other. In non-shown embodiments the probability factor vector 13 can comprise more or less than seven probability factors.
  • a third substep S3 at least one decision factor is determined on the basis of probability factor.
  • the data processing device 9 executes a random algorithm to determine a random number. Afterwards, the determined random number is compared with the probability factor. If the resulted random number is greater than the probability factor, the data processing device 9 sets the decision factor to 1 . However, if the resulted random number is smaller than the probability factor, the data processing device 9 sets the decision factor to 0.
  • the value "1" for the decision factor means that the processing operation associated to the probability factor is executed.
  • the value "0" for the decision factor means that the processing operation associated to the probability factor is not executed by the data processing device 9.
  • a fourth substep S4 the selected image sections are processed by the processing operations that are assigned to the decision factor comprising the value "1". This processing results in a synthetic image 1.
  • a fifth substep S5 the third and fourth substep S3, S4 is repeated for a predetermined number of times. That means, new decision factors are determined resulting in a different processing of the selected image sections. This results in further synthetic images wherein the number of synthetic images corresponds to the number of times that new decision factors are determined.
  • the decision factors are determined for the predetermined number of times and afterwards the image sections are processed using the determined decision factors.
  • the same probability factors are used for generating the synthetic images. That means, the same probability factors are used for all received image 2. Additionally, the same decision factors are used for all image sections 3. That means, each of the image sections 3 of an image is processed by using the same decision factors, i.e. the same column of the matrix 14 is used for all image sections 3 of an image 2 identified in substep S1 .
  • Fig. 8A shows a flow chart of a processing operation according to a first embodiment
  • Fig. 8B shows a selected image processed by the processing operation according to fig. 8A.
  • both figures show a processing operation that is performed in the fourth substep S4 shown in fig. 6.
  • all image sections 3 of an image are processed by the processing operation that is assigned to the decision factor having the value "1".
  • this processing operation at least one crop element 15 is used to crop a part of the image section 3 and thus of the target object 4.
  • contour points 17 of the image section 3 are determined.
  • the contour points 17 can be all points that characterize the image section 3.
  • contour points 17 can be edges, maxima and minima of the contour of the image section 3, turning points of the contour of the image section 3 and/or an end of the image section 3.
  • the image section 3 shown in fig. 9C shows several contour points 17.
  • the contour points 17 can be determined by using the result of a first and/or second derivative of the contour of the image section 3.
  • Fig. 9E shows an image in which the points 6 shown in fig. 9D are processed by the processing operation according to the second embodiment. Specifically, fig. 9E shows the image section 3 after the points 6 are moved within a plane comprising the image section. The points 6 are moved relative to the state shown in fig. 9D.
  • the data processing device determines for each point 6 the number of connections to other points 6 and/or the length of the connection line 18 and/or a target object area around the respective point 6. Based on said criteria, the data processing device determines which of the considered points 6 are moved and how the point 6 is moved.
  • the data processing device determines that each point can be moved.
  • the data processing device moves the points 6 dependent on the number of connections of the considered point to other points, the length of the connection line 18 and the target object area around the considered point.
  • the end points 6 are only connected with one other connection point.
  • said end points 6 can be moved more, in particular moved in linear direction and/or rotated, than the other points 6.
  • the data processing device can ensure that the cross section area of the processed image section, i.e. the image section having the moved points is identical to the cross section area of the image section as shown in fig. 9c. .
  • the cross section area of the resulted target object differs from the cross section area of the target object shown in fig. 9B.
  • Fig. 10A shows a flow chart of a processing operation according to a third embodiment and Fig. 10B a selected image processed by the processing operation according to fig. 10A.
  • the upper part of figure 10B shows the target object 4 before it is processed by the processing operation and the lower part of figure 10B shows the target object 4 after it is processed.
  • both figures show a processing operation that is performed in the fourth substep S4 shown in fig. 6.
  • all image sections 3 of an image are processed by the processing operation that is assigned to the decision factor having the value "1". In this processing operation at least one part of the image section is replaced with another part of the image section 3.
  • a grid 7 comprising a plurality of grid elements is created within the image section 3. This is shown in the upper part of Fig. 10B.
  • a second step R2 the image section 3 is processed such that a first grid element 8a is replaced by a second grid element 8b.

Landscapes

  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Image Analysis (AREA)
  • Image Processing (AREA)

Abstract

The invention relates to a method for generating at least one synthetic image (1), wherein the method comprises the following steps receiving of at least one image (2), processing at least one selected image section (3) in at least one processing operation, wherein the execution of the at least one processing operation depends on a probability factor and generating at least one synthetic image (1) comprising the processed at least one image section (3).

Description

Method for generating at least one synthetic image
The invention relates to a method for generating at least one synthetic image. Additionally, the invention relates to a data processing device comprising means for carrying out the method. In addition, the invention relates to a computer program product, a computer readable medium having stored thereon the computer program product and a data carrier signal carrying the computer program product.
It is known to use of neural networks, in particular of convolutional neural networks, in different kind of applications. Each of said neural networks has to be trained in a training process to provide accurate results in an operation process. Operation process means the situation when new data, in particular images, that is not used in the training process is inputted into the neural network, wherein the neural network outputs a result by using said inputted data. In the training process training data like images are used to train the neural network.
However, the accuracy of the output of the neural network depends on the amount and the quality of information of training data that is inputted into the neural network in the training process. The generalizability of use of the neural network also depends on the amount of training data that is inputted into the neural network. Generalizability means the ability of the neural network to provide accurate results in different kind of applications, including also applications for which the neural network is not specifically trained for.
Accuracy is the number of correctly predicted data points out of all the data points. More formally, it is defined as the number of true positives and true negatives divided by the number of true positives, true negatives, false positives, and false negatives. A true positive or true negative is a data point that the algorithm correctly classified as true or false, respectively. A false positive or false negative, on the other hand, is a data point that the algorithm incorrectly classified.
To increase the neural network's accuracy and generalizability and to avoid overfitting of the neural network a large number of training data having good quality is necessary. It is known from the prior art to generate synthetic images by processing the total image. For example, an augmentation process can be applied on the image to generate one or more synthetic images. The augmentation process can comprise one or more of the following processing changing the color of the image, changing the brightness of the image, changing the contrast of the image and rotating the image. However, the number of generated synthetic images is often not sufficient to achieve a satisfactory increase in accuracy and generalizability of the neural network that is needed for the specific application.
The object of the invention is to increase the number of images having high quality that are used for training a neural network such that it is accurate and that it has a high generalizability.
The object is solved by a method for generating at least one synthetic image, wherein the method comprises the following steps: receiving of at least one image, processing at least one selected image section in at least one processing operation, wherein the execution of the at least one processing operation depends on a probability factor and generating at least one synthetic image comprising the processed at least one image section.
The method, in particular computer implemented method, has the advantage that a large number of synthetic images can be generated. As is described later the generated images can be used in a training process for training a neural network, in particular a convolutional neural network. Due to the large number of synthetic images, it is possible to train the neural network such that it provides accurate results. A further advantage is that the synthetic images have a good quality. Quality means that the image show the information that is needed for the specific application. This is advantageous as it is sometimes hard to get real data in a sufficient amount. In particular, by using the inventive method the number of real images taken by an image acquisition device can be low. This is possible as according to the invention a high number of synthetic images can be provided having a lot of variations so that the neural network can be better trained.
By providing more synthetic images that can be used in the training of the neural network, an overfitting of the neural network can be avoided. This is possible as synthetic images can be generated that have many variations of pattern, shape, color, placement, etc. in the selected image part. In fact, it was realized that a large number of synthetic images can be created if only an image section is processed in at least one processing operation. The invention also enables to train a neural network with a small data set, i.e. a low number of real images taken by the image acquisition device. This is possible as due to the inventive method a large number of synthetic images can be generated.
Additionally, it was realized that the image processing shall depend on a probability factor to avoid generating a high number of synthetic images that are not useful for training purposes but results in high computation training costs. Usually, a high number of synthetic images is preferred for training a neural network. However, it is realized that processing of the at least one image section by all theoretical possible processing operations that are described below more in detail results in synthetic images that are not useful for training a neural network. This is the case as the processing can lead to a synthetic image showing an object that would not appear in reality.
In the following an image of the reality is acquired by an image acquisition device. The acquired image of the reality can be labeled. The image acquisition device can be any component comprising optical means by means of which an image can be acquired. Thus, the image acquisition device can be a camera and/or mobile phone and/or tablet and/or a microscope, etc.. The image acquisition device can be configured to acquire signals. The signals can have a wavelength which is in a human visible range. Thus, the image acquisition device can acquire light comprising a wavelength of 380nm to 780nm (nanometer). Alternatively, the image acquisition device can also acquire signals having a wavelength that is outside the human visible range.
In contrary to the image, the synthetic image is not acquired by the image acquisition device but generated by the data processing device. A processed synthetic image is an image which is processed by at least one processing operation. Thus, the synthetic image can show an artificial state that differs from the reality acquired by the image acquisition device. As discussed below more in detail the synthetic image is based on the acquired image. Thus, the synthetic image can show a state that differs from the reality shown in the image.
The image can comprise a target object or several target objects. The target object can be an object having a predetermined physical property. The object can be a discrete object with well-defined boundaries and spatial extension. Thus, the object can be anything that is visible for a human and/or tangible and/or that can be touched. For example, the target object can be a car, chair, scratch, etc..
Alternatively or additionally, the object can be represented by one or more pixels of the image. Thus, the object can be a digital pattern having a predetermined information type. The information type can define whether the digital pattern is visible and/or tangible and/or can be touched. Thus, the target object can cover objects being nontangible and/or non-touchable and/or non-visible in the human light range. In that case, the target object can be image portions consisting of one or more pixels having information about for example a specific temperature, reflectance, radiance, etc.. The probability factor is predetermined. It can be stored in an electrical memory, in particular a memory of the data processing device. The probability factor can be determined from experiments before the generation of the synthetic image is started.
The at least one image section can be automatically selected by the data processing device. Additionally or alternatively the at least one image section can be selected by a user. In that case the selected at least one image section is transmitted to the data processing device and/or the data processing device receives the selected at least one image section. This is explained below more in detail. The selection of the image section is done such that the location and/or shape of the image section is known. In particular, the data processing device knows the location and shape of each of the selected image section. The image section is a part of the image.
Several image sections can be automatically selected by the data processing device. Alternatively or additionally, several image sections can be selected by a user. In that case the selected image sections are transmitted to the data processing device and/or the data processing device receives the selected image sections. In particular, the location and/or the shape of each of the image sections is transmitted to the data processing device and/or received by the data processing device. Each of said several image sections can be processed by a processing operation, in particular the same processing operation. The selection of several image sections can be done when the image has more target objects. Thus, each image section can be assigned to one target object.
The method can be executed in a data processing device. The data processing device can comprise one or more processors or can be a processor. Alternatively, the data processing device can be a computer. The image data obtained by the image acquisition device is sent to the data processing device. Thus, the data processing device receives said image data from the image acquisition device and processes said image data.
The data processing device can generate several synthetic images. Said several synthetic images can be based on the same received image. That means, the image received from the data processing device results in several synthetic images. Additionally or alternatively the data processing device can receive several images from the image acquisition device. For each of said received images one or more synthetic images can be generated. According to an embodiment at least one selected image section can be a polygon that has a closed contour. The contour of the polygon can depend on a contour of a target object. That means, the contour of the polygon is chosen such that it corresponds to the contour of the target object. Using a polygon enables to create the closed contour that is needed and/or wished by the user. Contour means the outer boundary or rim of the image section, in particular the polygon. A polygon is a shape in geometry that has closed a structure. The polygon can comprise three or more corners. The number of corners depends on the shape of the target object. This is the main difference of using image sections being a polygon to the use of rectangular boxes as image sections that are used in known methods. The shape of the rectangular box used in the prior art is always the same independent of the form of the contour of the target object.
Another advantage of using a polygon as image section is that the data processing device automatically knows the target object contour and the location of the target object. This results as the contour of the polygon corresponds to the contour of the target object. Thus, there is no need for any other objection detection methods for detecting the location and contour of the target object.
The image can have at least one target object. In particular, the image can have several target objects. The selection of the image section is done such that that the target object is arranged in the selected image section. That means, the image selection is located and shaped such that the target object is arranged inside the image section. The image section selection can be done such that the contour of the image section corresponds to the contour of the target object. This can be achieved by using a polygon having a contour which is adapted to fit to the contour of the target object. Contour of the target object means the outer boundary or rim of the target object.
As mentioned before the image selection can be a polygon. Thus, it is possible to accurately surround the target object without having too many pixels between the target object and the image section contour. Such kind of labeling is also indicated as "weak labeling". Thus, the labeling is simplified as it has not to be assured that the contour of the image section exactly matches with the contour of the target object. However, for the accuracy of the artificial neural network that is trained using the synthetic image it is preferred to minimize the number of pixels within the image section that do not belong to the target object.
Using polygons has the advantage that the labelling can be improved as the image section comprises or mainly comprises the target object. By improving labelling, the accuracy of the neural network which is used by the labelled images can be improved. As is explained below more in detail, points characterizing the target object can be easily determined by using polygon. Using said points enables to provide more realistic synthetic images. This is possible because by using a polygon it is ensured that only or mainly the points and thus the target object is processed by a processing operation. Thus, the artificial neural network to which the synthetic image is input for training only receives information about the target object. In other words, the artificial neural network does not receive or receives much less information that do not belong to target object when the image section is a polygon which contour depends on the contour of the target object, than by using a rectangular box for selecting the image section as it is done in the prior art. Additionally, by identifying the points it can be ensured that only relevant parts of the target object are processed.
As mentioned above the target object can be a discrete object and/or represented by one or more pixels. The target object can be labeled when the image section is selected. Labeling means that the target object is named. However, it is possible to label the target object after the image section is processed in the processing operation. Additionally or alternatively, further information of the target object, in particular the instance of the target object, can be provided. Thus, at the end of the image section selection the data processing device has information about the location and/or shape and/or name and/or further information of the target object.
The image and/or the at least one generated synthetic image can be used as training data for an artificial neural network, in particular a convolutional neural network. The neural network can be trained by using the image acquired by the image acquisition device and the generated synthetic images. Furthermore, the neural network can be trained to classify in an operation mode after the training mode is finalized whether an inputted image comprises the target object or not. By using many synthetic images, the trained neural network can be made more generalized. This is possible as due to generated synthetic images the neural network can learn a plurality of variations within the selected image section. The generated synthetic images help the neural network to generalize the teaching. That means, thanks to the generated synthetic images the neural network will know variations of the image sections even though it is not trained for.
The image can comprise at least one region of interest, wherein the selected image section is arranged in the region of interest. The region of interest can be considered as mask wherein only the image parts being arranged in the region of interest are of interest for the further processing and the remaining parts can be ignored. For cases, in which the image comprises a target object, the target object is arranged in the region of interest. The region of interest can be predetermined. For example, a user can predetermine the shape and location of the region of interest. For example, if the target object is a scratch to be identified in an object like a car door, the region of interest corresponds to the car door and the target object corresponds to the scratch. It is clear that the target object is not limited to scratches and the region of interest is not limited to car doors.
By using the region of interest and target object it is possible to consider a specific application of the neural network to be trained with the image and generated synthetic images. Thus, by performing the method many synthetic images showing variations of the target objects can be generated. This enables a better training of the neural network resulting in more accurate results of the neural network.
According to an embodiment several probability factors can be provided and each probability factor is assigned to a processing operations. The probability factor can be stored in an electrical memory of the data processing device. The number of probability factors can correspond to the number of processing operations. In particular, each probability factor is assigned to a processing operation. The probability factor can be a number. Thus, the probability factor indicates the likelihood of processing the processing operation assigned to the probability factor.
The probability factors can be independent on each other. The number of factors can be in the range 0 to 1. The sum of the probability factors can differ from 1 or can have the value 1 . Using a probability factor which is smaller than 1 has the advantage that the number of processing operations to be performed on an image is not high. This is advantageous as performing too many processing operations on the selected image section, in particular the target object, result in a non-realistic target object. Training the neural network with such non-realistic target objects does not result in an increased accuracy and thus is not beneficial. The same applies if the image does not comprise a target object. By using the probability factor, the natural occurrence of certain characteristics in the image section is mimicked.
A decision factor can be determined using the probability factor, wherein a decision whether to execute the processing operation is made dependent on the decision factor. In particular, for each of the probability factors a decision factor is determined, wherein a decision whether to execute the processing operation that is assigned to a probability factor is made dependent on the decision factor that is assigned to the processing operation. Thus, the data processing device can determine the decision factor easily by using the probability factor. In particular, only one probability factor is used for determining the decision factor.
The data processing device can determine the probability factor by using a random algorithm. The random algorithm is started and the algorithm result is compared with the probability coefficient. If the algorithm result is lower than the probability factor the decision factor is set to 0 and if the algorithm result is greater than the probability factor, the decision factor is set to 1. The data processing device can perform the processing operation dependent on the decision factor. If the decision factor has the value 0, the data processing device does not perform the processing operation. However, if the decision factor has the value 1, the data processing device performs the processing operation. This process is repeated for each probability factor. Thus, at the end a set of decision factors is determined wherein the set of decision factors comprises one or more decision factors. The processing operation is performed only in the selected image section dependent on said at least one decision factor. As the selected image section corresponds to the target object, the processing operation is applied on the target object.
The data processing device can repeat the determination of the at least one decision factor for a predetermined number of times, in particular and for each of said decision factors a decision whether to execute the procession operation is made. Thus, a set of decision factors is determined for each of the predetermined times.
A synthetic image is generated after at least one processing operation, in particular one or more processing operations, is applied on the selected image section. The processing operation can be configured such that only a part of the image section is processed. This increases the number of different synthetic images.
It is possible that several processing operations are applied on the same image. That means, the synthetic image is made by applying several processing operations on it. The processing operations can differ from each other. That means, different kind of processing are applied on the at least one selected image section as it is explained below more in detail.
The number of generated synthetic images corresponds to the number of times the determination of the at least one decision factor is repeated. In other words, the number of generated synthetic images corresponds to the number of determined sets of determination factors. The number of synthetic images to be generated on the basis of the received image can be predetermined by a user.
The data processing device can receive several images. The probability factors can be identical for each received image. The data processing device can determine at least one decision factor for each received image. In other words, the same method as described above can be performed for each received image. That means, one or more synthetic images can be generated for each image.
The data processing device can determine for each received image a vector comprising the probability factor or factors. This is the case when the same probability factor shall apply to each selected image selection. Additionally, the data processing device can determine for each received image a matrix for the determination factors. The number of columns of the matrix of determination factors depends on the number of synthetic images to be generated. The column corresponds to the set of determination factors discussed above.
After the data processing device determined the determination factor or factors, the at least one processing operation can be applied on the at least one selected image section and thus on the target object. The decision whether one or more processing operation or operations are applied on the at least one selected image section, in particular a part of the at least one selected image section, depend on the determined decision factor or factors. If the image comprises several image sections, the same at least one processing operation can be applied to all image sections dependent on the determined decision factor.
According to an embodiment the processing operation can comprise to apply a crop element on the selected image section in order to remove a part of the selected image section and thus of the target object. Thus, the data processing device can easily generate a synthetic image. The data processing device can locate the crop element randomly within the selected image part. Additionally, the number of crop elements and/or the shape of the at least one crop element can be randomly selected by the data processing device and/or can be predetermined by the user. Using at least one crop element has the advantage that e.g. scratches can be generated in the synthetic image.
In another processing operation the data processing device can determine of at least one point on or of the selected image section and move the point. The point movement results in that the image section is distorted. As the target object is arranged within the image section, the target object is also distorted.
The point can be moved in a plane comprising the image section. In said case the cross-section area of the target object and/or the image section can remain the same as before the movement. Alternatively, the point can be moved in a third direction, i.e. in a direction directed away from the aforementioned plane, i.e. in a direction away from the plane comprising the image section. In said case the cross section area of the target object and/or the image section changes. The data processing device can randomly move the point, in particular can randomly move the point within a predetermined range.
In order to determine the point to be moved, the processing operation can comprise determining a contour point of the polygon, in particular of the contour of the image section. The contour point can be a characterizing point of the contour. A characterizing point of the contour can be at least one of a maximum, minimum or turning point of the contour. Additionally or alternatively the characterizing point can be an edge. As the image section is selected by the user and/or data processing device the shape and location of the image section of the image section, in particular the polygon is known. By knowing the shape of the image section, the contour of the image section, in particular the polygon, is also known. Thus, it is easily possible to determine the at least one contour point of the image section, in particular the polygon.
The at least one contour point can be determined by determining the first and/or second derivative of the contour of the image section. In particular, the at least one contour point can be determined dependent on the result of the first and/or second derivative. Additionally, a clustering method and/or a filter can be applied on the derivative results to improve the detection result of contour points.
In the clustering method the location of the determined contour points can be considered. One or more contour points can be used to define at least one group. Specifically, neighboring contour points can define the group, which can have a predetermined extension, in particular an extension in two dimensions. Thus, it is possible to determine for each of the contour point dependent on the location of said contour point whether the contour point can be assigned to a group. In other words, the contour points are classified if they are assigned to one of the determined group. If this is the case, the contour points are further processed. If this is not the case, the contour points are not considered anymore in the further processing. The filter can ensure that points on a flattering part of the contour of the image section are not considered as contour points. Thus, the contour points can be easily determined. The filter can be a low pass filed like moving average.
The at least one point to be moved can be determined dependent on the location of at least two determined contour points. Specifically, it can be determined for a contour point the neighboring determined contour points that fulfill a predetermined condition. The predetermined condition can be a predetermined distance between the contour point and the neighboring characterizing point. If the distance is within a predetermined range, the neighboring contour point are assigned to the point and thus are part of the group discussed above. This process can be repeated for each of the determined contour points. At the end of the process one or more groups comprising at least one contour point is determined. In the next step the point to be moved can be determined by determining a representative point of each group. Said representative point can be a middle point that has the same distance to each of the contour points of the respective group. Said middle point corresponds to the point mentioned above that is to be moved.
Said approach has the advantage that points of the target object can be easily determined. Specifically, the determined point or points can but do not have to be arranged on the contour of the image section. The determined points can be easily moved resulting in different shaped image sections and thus different shaped target objects. Thus, a plurality of different shaped target objects can be created. This increases the number of synthetic images that can be created.
In the following, the generation of the different shaped target object is explained more in detail. As mentioned before several points are determined on the basis of the contour points of the contour of the image section, in particular the polygon. The determined points are connected with each other. However, the connection is made such that the connection result represents the target object. That means, that each of the determined points can be connected with one or more other determined points. Specifically, dependent on the target object shape a point can but does not have to be connected to all other determined points. After the points are connected to each other, a skeleton of the target object is created. Said skeleton simplifies to determine which of the determined points can be moved and/or how it is moved. The selection of the determined point to be moved and the movement of the determined point can depend on the number of connections of the considered point to other points and/or the length of the connection line between the considered point and another point and/or the target object area around the considered point. The target object area can be determined in a predetermined region around the considered point. The target object area is the area within the target object contour being arranged in the predetermined region around the considered point. Dependent on the contour of the target object, the predetermined region can also have region parts that do not belong to the target object.
Based on the determined number of connections and/or the length of the connection line and/or the target object area, the point to be moved can be selected and/or it can be determined how the point is moved. Specifically, it can be determined that a point that is connected to a lot of other points cannot be moved and thus is not selected. On the other side a point that has one or a few connections can be moved. The movement of the point can be dependent on the number of connections of the considered point with other points and/or the length of the connection line and/or the target object area. Thus, a point with one connection or few connections can be moved more than a point with more connections. "Movement" includes the movement type, i.e. a rotary and/or translatory movement, and the movement region in which the point can be moved. The movement region is greater for points with one or few connections in comparison to points which have a high number of connections. For said points the movement area within which the point can be moved is small.
After the at least one point is moved, a different image section and thus a different target object results. The different shaped image sections can be processed by at least one processing operation mentioned above or below wherein the processing depends on the respective probability factor. Thus, starting from an image showing the target object a plurality of synthetic images can be generated that show a different shaped and/or processed target object.
The data processing device can perform a further processing operation in which a grid is created within the within the selected image section, wherein the grid comprises several grid elements. The data processing device can replace a first grid element of the grid by a second grid element of the grid. The replacement of the grid elements occurs only within the selected image section. The data processing device can perform an additional processing operation wherein the processing operation comprises to change the brightness of the selected image section and/or the contrast of the selected image section and/or the color of the selected image section.
The data processing device can execute one or more of the aforementioned processing operations sequentially or parallel to each other. As it is explained above, the data processing device executes one or more of the aforementioned processing operations for each received image in the selected image section of each image.
According to an embodiment the selected image section can be moved within the image to another location within the image or within a predetermined region of interest of the image to another location within the region of interest. Thereby, one or more synthetic images can be generated. The data processing device can repeat said movement several times. Thus, several synthetic images are created in which the image selection is arranged at a different location within the image and/or within the region of interest.
The data processing device can perform an augmentation process on the received image and/or the generated synthetic image. Thereby, a new synthetic image is generated. The augmentation process can comprise to change the brightness and/or color and/or contrast of the received image or the generated synthetic image.
According to another aspect a data processing device is provided. The data processing device comprising means for carrying out an inventive method. Additionally, an image acquisition device for acquiring images is provided. The image acquisition device can be at least one of the following a camera, a mobile phone, microscope and a tablet. The data processing device can be part of the image acquisition device. Alternatively, the data processing device can be electrically connected to the image acquisition device. The acquisition device can be configured to acquire visible light. "Visible light" means that the acquired light has a wavelength in the range of 380 to 780 nanometers. Alternatively, the image acquisition device can be configured to acquire non-visible light.
According to a further aspect of the invention a computer program product is provided wherein the computer program product comprises instructions which, when the program is executed by the data processing device, in particular a computer, cause the data processing device, in particular the computer, to carry out the steps of the inventive method. Additionally, a computer-readable data carrier is provided wherein the computer-readable data carrier has stored thereon the computer program product. Also a data carrier signal is provided wherein the data carrier signal carries the computer program product.
A computer-readable carrier can be any available medium that can be accessed by a general purpose or special purpose computer system. The computer-readable medium that stores the computer program product is non-transitory computer-readable storage media. Non-transitory computer- readable storage media includes RAM, ROM, EEPROM, CD-ROM, solid state drives ("SSDs") (e.g., based on RAM), Flash memory, phase-change memory ("PCM"), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.
In the figures, the subject matter of the invention is shown schematically, with identical or similarly acting elements being mostly provided with the same reference signs. Therein shows:
Fig. 1 an image with a randomly selected image section.
Fig. 2 an image with a selected image section and a target object.
Fig. 3 a system comprising a data processing device and an image acquisition device according to a first embodiment.
Fig. 4 a system comprising a data processing device and an image acquisition device according to a second embodiment.
Fig. 5 a flow chart of the general method comprising the generation of at least one synthetic image.
Fig. 6 a flow chart of generating a synthetic image.
Fig. 7 a probability vector and matrix for decision factors.
Fig. 8A a flow chart of a processing operation according to a first embodiment.
Fig. 8B a selected image processed by the processing operation according to fig. 8A.
Fig. 9A a flow chart of a processing operation according to a second embodiment.
Fig. 9B an image showing a target object and the image selection.
Fig. 9C the image shown in fig. 9B in which contour points of the contour of the image section are shown.
Fig. 9D an image in which points and their connection lines are shown.
Fig. 9E an image in which the points shown in fig. 9D are processed by the processing operation according to the second embodiment.
Fig. 10A a flow chart of a processing operation according to a third embodiment.
Fig. 10B a selected image processed by the processing operation according to fig. 10A.
An image 2 as shown in Fig. 1 is acquired by an image acquisition device 11 shown in fig. 3 and 4. A data processing device 9 shown in fig. 3 and 4 receives said image and processes at least one image section 3 in at least one processing operation. In the image 2 shown in fig. 1 the data processing device 9 and/or a user selects two image sections 3 within a region of interest 5 of the image 1 . The location and/or the shape of the image sections 3 are randomly selected. In the embodiment shown in fig. 1, two image sections 3 are selected that have the same shape. However, in other non-shown embodiments, the number and shape of selected image sections can differ. After the selection of the image sections, the data processing device 9 has information about the location and shape of said image sections 3.
The execution of the at least one processing operation by the data processing device 9 depends on a probability factor. In particular, only the image section is processed by the processing operation. That means, the remaining part of the image is not processed by the processing operation. The data processing device 9 generates at least one synthetic image comprising the processed at least one image section.
Fig. 2 shows an image 2 with a selected image section 3 and a target object 4. Said image 2 differs from the image 2 shown in fig. 1 in that a target object 4 is arranged in a region of interest 5 of the image 2. The user and/or the data processing device 9 selected an image section 3 that surrounds the target object 4. In particular the image section 3 is selected such that it corresponds to the target object 4, in particular that the contour of the image section 3 corresponds to the contour of the object 4. Likewise, to fig. 1 the image section 3 is a polygon with a closed contour. The contour of the polygon depends on the target object, in particular on the contour of the target object. In other words, the contour of the polygon corresponds to the contour of the target object 4. As in fig. 2 the target object has a rectangular shape comprising four corners, the polygon has a polygon shape with four corners. For a non-shown case in which the target object has a circular shape, the image section, in particular the polygon, has a circular shape. However, as it is shown in fig. 2, in practice it is not always possible to place the contour of the image section on the contour of target object 4 so that a small distance might exist between the contour of the target object 4 and the contour of the image section 4. Said distance should be kept as small as possible and in the ideal case should not be existent. In this embodiment the location and the shape of the image section 3 is dependent on the location and shape of the target object 4. Additionally, further information on the target object 4 can be provided to the data processing unit. For example, the name and/or instance of the target object 4 can be provided to the data processing device 9. Thus, at the end of the image section selection process, the data processing device 9 has information about the location, shape and further information of the image section 3 and thus of the target object 4. Said information are assigned to the selected image section, respectively.
Fig. 3 shows a system 16 comprising the data processing device 9 and an image acquisition device 11 according to a first embodiment. The image acquisition device 11 acquires an image 2, e.g. an image as shown in fig. 1 and 2. The image information is transmitted to the data processing device 9. The data processing device 9 can be arranged within the image acquisition device 11 or be a component arranged distant from the image acquisition device 11.
The data processing device 9 generates on the basis of the received image one or more synthetic images. Thereto, one or more image sections 3 are selected. In the embodiment shown in fig. 3 the data processing devices 9 automatically selects the image sections 3. For the case that the image 2 does not comprise a target object 4, the data processing device 9 randomly selects the image sections 3. However, if the image 1 has at least one target object 4, the data processing device 9 detects the target object 4 and places the contour of the image section 3, in particular creates a polygon with a corresponding contour, so that it matches with the contour of the target object 4. The generated synthetic images are used to train a neural network 10. The neural network can be trained by the data processing device 9 or another processing device not shown in the figures.
Fig. 4 shows a system 16 comprising the data processing device 9 and an image acquisition device 11 according to a second embodiment. The second embodiment differs from the first embodiment in that the data processing device 9 does not automatically randomly select the image section 3 and/or does not detect the target object 4. In the second embodiment a user selects the image section 3 and/or places the contour of the image section 3 such that it matches with the contour of the target object 4. Additionally, the user can provide the data processing device 9 with further information about the target object 4, in particular the user can name the target object 4. Thereto, the system 16 comprises at least one input means by means of which the aforementioned information can be transmitted to the data processing device 9. The input means can be a keyboard and/or touchscreen and/or interfaces by means of which information about the selected image section 3 are transmitted to the data processing device 9.
Fig. 5 shows a flow chart of the general method. As mentioned before in a first step G1 the image acquisition device 11 acquires an image 2. Said image is transmitted to the data processing device 9. In a second step G2 the image section 3 is selected. This is done be creating a polygon having a contour that depends on the contour of the target object and locating the contour such that the target object is arranged in the polygon. As it is described above, the image section selection can be done by the data processing device 9 and/or by a user. In the latter case, the information about the selected image section 3 are transmitted to the data processing device 9. Thus, at the end of the second step G2 the data processing device 9 has information about at least the shape and location of all selected image sections.
In a third step G3 one or more synthetic images are generated on the basis of the received image. After the synthetic images are created, the data processing device replaces or moves the selected image section 3 in a fourth step G4. In particular, the image sections that are processed in the third step G3 are placed to a new location within the image or within the region of interest 5. The result of said replacement is that further synthetic images are created.
In a fifth step G5 the data processing device 9 performs an augmentation process on the image or on the region of interest 5. In the augmentation process the data processing device 9 changes the brightness of the at least one image generated in the third step G3 and/or the contrast of the at least one image generated in the third step G3 and/or the color of the at least one image generated in the third step G3. The result of the fifth step G5 is that further synthetic images are generated.
In fig. 5 the fifth step G5 is executed after the fourth step G4. However, in a non-shown embodiment the fourth step G4 and the fifth step G5 can be executed parallel to each other. The second to fifth step G2-G5 are executed in the data processing device 9.
In a sixth step G6 the image 2 acquired by the image acquisition device 11 and the generated synthetic images 1 are transmitted to a neural network to train the neural network. In the embodiment shown in fig. 5 the training of the neural network is not performed in the data processing device 11. However, in a non-shown embodiment the training of the neural network can be performed by the data processing device 11. Fig. 6 shows a flow chart of generating a synthetic image 2. In particular, fig. 6 shows the substeps of the third step G3 shown in fig. 4. Fig. 7 shows a probability vector and a matrix for decision factors. Fig. 7 shows the received image 2 comprising two target objects 4 and thus two selected image sections 3. The target objects 4 and thus the image sections 3 differ in their shape and location from each other. In the following, the generation of synthetic images is explained referring to figures 6 and 7.
In a first substep S1 the data processing device 9 determines all image sections present in the received image and that have been selected in the second step G2 shown in Fig. 5. The processing operations discussed below are performed in all determined image sections.
In a second substep S2 the data processing device 9 sets the probability factors. The probability factors indicate the likelihood that a processing operation shall be performed on the image section 3. The number of probability factors corresponds to the number of processing operations by means of which the selected image section 3 can be processed. The probability factors are predetermined before the method is executed and thus the image 2 is acquired. In particular, the probability factors are stored in an electrical memory of the data processing device. Alternatively, the user can enter the probability factors.
The probability factors can have a value between 0 and 1. At the end of the second substep S2 a probability factor vector 13 is present wherein the vector comprises the probability factor for each processing operation. Each processing operation is assigned to a probability factor. In the case shown in fig. 7 the probability factor vector 13 comprises seven probability factors. This means, the data processing device 9 can process the image 2 in seven different ways as the processing operations differ from each other. In non-shown embodiments the probability factor vector 13 can comprise more or less than seven probability factors.
In a third substep S3 at least one decision factor is determined on the basis of probability factor. The data processing device 9 executes a random algorithm to determine a random number. Afterwards, the determined random number is compared with the probability factor. If the resulted random number is greater than the probability factor, the data processing device 9 sets the decision factor to 1 . However, if the resulted random number is smaller than the probability factor, the data processing device 9 sets the decision factor to 0. The value "1" for the decision factor means that the processing operation associated to the probability factor is executed. The value "0" for the decision factor means that the processing operation associated to the probability factor is not executed by the data processing device 9.
In a fourth substep S4 the selected image sections are processed by the processing operations that are assigned to the decision factor comprising the value "1". This processing results in a synthetic image 1. In a fifth substep S5 the third and fourth substep S3, S4 is repeated for a predetermined number of times. That means, new decision factors are determined resulting in a different processing of the selected image sections. This results in further synthetic images wherein the number of synthetic images corresponds to the number of times that new decision factors are determined.
In a non-shown embodiment, it is possible that the decision factors are determined for the predetermined number of times and afterwards the image sections are processed using the determined decision factors.
As it is shown in fig. 7 the decision factors are stored in a matrix 14. The matrix 14 has 7 rows and 5 columns. Each column comprises decision factors that are used for processing the image sections 3 of an image 2. As the matrix 14 comprises five columns the data processing device 9 generates five synthetic images by processing the image 2 dependent on the decision factors of each column of the matrix 14.
As is evident from fig. 7, the synthetic images 1 differ from the image 2 in the selected image sections 3. That means, the data processing device 9 only processes the selected image sections 3 by applying the processing operations and does not process the remaining part of the image 2. The two synthetic images 1 shown in fig. 7 differ from each other in their shape of their processed image sections 3 and thus in the shape of the target objects 4. This results as each of the decision factors of the column of the matrix 14 is assigned to a probability factor of the probability factor vector 13. The decision factors of the columns differ from each other in their value due to the determination process of the third substep S3. Said different decision factor values result in that the image sections 3 of the images are processed by different processing operations.
In the embodiment shown in fig. 7 the same probability factors are used for generating the synthetic images. That means, the same probability factors are used for all received image 2. Additionally, the same decision factors are used for all image sections 3. That means, each of the image sections 3 of an image is processed by using the same decision factors, i.e. the same column of the matrix 14 is used for all image sections 3 of an image 2 identified in substep S1 .
Fig. 8A shows a flow chart of a processing operation according to a first embodiment and Fig. 8B shows a selected image processed by the processing operation according to fig. 8A. In particular, both figures show a processing operation that is performed in the fourth substep S4 shown in fig. 6. As it is explained above all image sections 3 of an image are processed by the processing operation that is assigned to the decision factor having the value "1". In this processing operation at least one crop element 15 is used to crop a part of the image section 3 and thus of the target object 4.
In a first step C1 the data processing device 9 determines a crop element shape and a crop element position within the image section 3. The data processing device 9 can randomly locate the crop element within the image section 3. A plurality of different shaped crop elements can be stored in a memory of the data processing device 9 so that the data processing device 9 can select one of the plurality of crop elements. Alternatively, a user can select the crop element and decide where to locate the crop element. In the present case three crop elements 15 are placed within the image section 3 in a second step C2 wherein two crop elements have a rectangular shape and one crop element has a triangular shape.
Fig. 9A shows a flow chart of a processing operation according to a second embodiment and Fig. 9B shows an image showing a target object 4 and the selected image section 3. The image section 3 is a polygon that is adapted to fit to the contour of the target object 4. In fig. 9B the contour of the polygon is distant to the target object 4. However, this is just made for illustration purposes. Ideally the contour of the image section 3, in particular of the polygon, matches the outer contour of the target object 3.
As it is explained above all image sections 3 of an image are processed by the processing operation that is assigned to the decision factor having the value "1". In this processing operation at least one point 6 shown in fig. 9D and 9E is moved as is explained below more in detail.
Thereto, in a first step P1 contour points 17 of the image section 3 are determined. The contour points 17 can be all points that characterize the image section 3. For example, contour points 17 can be edges, maxima and minima of the contour of the image section 3, turning points of the contour of the image section 3 and/or an end of the image section 3. The image section 3 shown in fig. 9C shows several contour points 17. As the shape of the contour and the location of the contour of the image section 3 is known, the contour points 17 can be determined by using the result of a first and/or second derivative of the contour of the image section 3.
In a second step P2 the determined contour points 17 are used to determine points 6 which are shown in fig. 9D. Thereto, the data processing device determines for a contour point 17 the neighboring characterizing points that are arranged within predetermined location of said contour point. Said determined contour neighboring points and the contour point are assigned to a group. In fig. 9C one group comprising four contour points 17 is symbolized with the dotted rectangular. This process can be repeated for all contour points. At the end several groups are determined.
Afterwards, for each of the groups a point 6 is determined. Thereto, a middle point can be determined that has the same distance to each of the characterizing points being part of the group. The middle point corresponds to the point 6. Fig. 9D shows several points 6 which are determined in the aforementioned way. In a third step P3 the determined contour points are connected to each other. In the fig. 9D the connection is realized by connection lines 18. The connection lines 18 are shown in dotted lines in Fig. 9D. In Fig. 9D the target object 4 is not shown for illustration purposes but only a skeleton realized by the connection lines 18 is shown. As the contour of the image section 3 corresponds to the contour of the target object 3, the determined points 6 are arranged on the target object.
Fig. 9E shows an image in which the points 6 shown in fig. 9D are processed by the processing operation according to the second embodiment. Specifically, fig. 9E shows the image section 3 after the points 6 are moved within a plane comprising the image section. The points 6 are moved relative to the state shown in fig. 9D. The data processing device determines for each point 6 the number of connections to other points 6 and/or the length of the connection line 18 and/or a target object area around the respective point 6. Based on said criteria, the data processing device determines which of the considered points 6 are moved and how the point 6 is moved.
For the case shown in fig. 9E, the data processing device determines that each point can be moved. The data processing device moves the points 6 dependent on the number of connections of the considered point to other points, the length of the connection line 18 and the target object area around the considered point. For example, the end points 6 are only connected with one other connection point. Thus, said end points 6 can be moved more, in particular moved in linear direction and/or rotated, than the other points 6. The data processing device can ensure that the cross section area of the processed image section, i.e. the image section having the moved points is identical to the cross section area of the image section as shown in fig. 9c. .
In a non-shown embodiment the points 6 in a direction directed away from said plane in third direction. In said case the cross section area of the resulted target object differs from the cross section area of the target object shown in fig. 9B.
Fig. 10A shows a flow chart of a processing operation according to a third embodiment and Fig. 10B a selected image processed by the processing operation according to fig. 10A. The upper part of figure 10B shows the target object 4 before it is processed by the processing operation and the lower part of figure 10B shows the target object 4 after it is processed. In particular, both figures show a processing operation that is performed in the fourth substep S4 shown in fig. 6. As it is explained above all image sections 3 of an image are processed by the processing operation that is assigned to the decision factor having the value "1". In this processing operation at least one part of the image section is replaced with another part of the image section 3.
Thereto, in a first step R1 a grid 7 comprising a plurality of grid elements is created within the image section 3. This is shown in the upper part of Fig. 10B. In a second step R2 the image section 3 is processed such that a first grid element 8a is replaced by a second grid element 8b.
Reference Signs
1 synthetic image
2 image
3 selected image section
4 target object
5 region of interest
6 point
7 grid
8a first grid element
8b second grid element
9 data processing device
10 Neural Network
11 image acquisition device
12 input means
13 probability factor vector
14 decision factor matrix
15 crop element
16 system
17 characterizing points
18 connection line between two points
G1 -G6 method steps
S1 -S5 method substeps
C1 -C2 method steps for crop element processing operation
P1 -P4 method steps for point movement processing operation
R1-R2 method steps for grid element replacing processing operation

Claims

Patent Claims
1 . Computer implemented method for generating at least one synthetic image (1), wherein the method comprises the following steps: receiving of at least one image (2), which has at least one target object (4), processing at least one selected image section (3) in at least one processing operation, wherein the execution of the at least one processing operation depends on a probability factor and generating at least one synthetic image (1) comprising the processed at least one image section (3) wherein the image section (3) is selected such that the target object (4) is arranged in the selected image section (3).
2. Computer implemented method according to claim 1, characterized in that the at least one selected image section (3) is a polygon having a closed contour and/or the at least one selected image section (3) is a polygon having a contour that depends on a contour of the target object (4).
3. Computer implemented method according to claim 1 or 2, characterized in that the image (2) comprises at least one region of interest (5), wherein the selected image section (3) is arranged in the region of interest (5).
4. Computer implemented method according to at least one of the claims 1 to 3, characterized in that several probability factors are provided and each probability factor is assigned to a processing operation.
5. Computer implemented method according to claim 4, characterized in that the probability factors are independent on each other.
6. Computer implemented method according to at least one of the claims claim 1 to 5, characterized in that a. a decision factor is determined using the probability factor, wherein a decision whether to execute the processing operation is made dependent on the decision factor or that b. for each of the probability factors a decision factor is determined, wherein a decision whether to execute the processing operation that is assigned to a probability factor is made dependent on the decision factor that is assigned to the processing operation. 7. Computer implemented method according to claim 6, characterized in that the determination of the decision factor is repeated for a predetermined number of times.
8. Computer implemented method according to claim 6 or 7, characterized in that a synthetic image is generated after the selected image section (3), in particular a part of the image section, is processed by at least one processing operation.
9. Computer implemented method according to at least one of the claims 1 to 8, characterized in that a. several images (2) are received wherein the probability factors are identical for each image (2) and/or in that b. at least one decision factor is determined for each image (2).
10. Computer implemented method according to at least one of the claims 1 to 9, characterized in that the processing operation comprises to apply a crop element on the selected image section (3) in order to remove a part of the selected image section (3).
11. Computer implemented method according to at least one of the claims 1 to 10, characterized in that a. the processing operation comprises determining of at least one point (6) of the selected image section (3) and moving the point (6) or in that b. the processing operation comprises determining of at least one point (6) of the selected image section (3) and moving the point (6) such that the cross section area of the selected image section (3) remains the same.
12. Computer implemented method according to at least one of the claims 1 to 11, characterized in that the processing operation comprises determining a contour point (17) of the polygon, in particular of the contour of the image section, or the processing operation comprises determining a contour point (17) of the polygon, in particular of the contour of the image section, wherein the contour point is a characterizing point of the contour of the image section.
13. Computer implemented method according to claim 12, characterized in that at least one point (6) is determined dependent on the location of at least two determined contour points (17). 14. Computer implemented method according to at least one of the claims 1 to 13, characterized in that a selection of the point (6) to be moved and/or the movement of the point (6) depends on the number of connections of said point (6) with other points and/or of the length of a connection line (18) connecting said point (6) to another point and/or a target object area around said point (6).
15. Computer implemented method according to at least one of the claims 1 to 14, characterized in that a grid (7) comprises several grid elements (8a, 8b) and the processing operation comprises to create a grid (7) within the selected image section (3) and to replace a first grid element (8a) by a second grid element (8b).
16. Computer implemented method according to at least one of the claims 1 to 15, characterized in that the processing operation comprises to change a. the brightness of the selected image section (3) and/or b. the contrast of the selected image section (3) and/or c. the color of the selected image section (3).
17. Computer implemented method according to at least one of the claims 1 to 16, characterized in that the selected image section (3) is moved within the image (2) to another location or within a predetermined region of interest (5) of the image (2) to another location.
18. Computer implemented method according to at least one of the claims 1 to 17, characterized in that an augmentation process is executed on the received image (2) and/or the generated synthetic image (1).
19. Computer implemented method according to at least one of the claims 1 to 17, characterized in that the image (2) and/or the at least one generated synthetic image (1) is used as training data for an artificial neural network.
20. Data processing device (9) comprising means for carrying out the method according to at least one of the claims 1 to 19.
21. Computer program product comprising instructions, which, when the program is executed by a data processing device (9), in particular a computer, cause the data processing device (9), in particular the computer, to carry out the method according to at least one of the claims 1 to 19. 22. Computer readable medium having stored thereon the computer program product of claim
21.
23. Data carrier signal carrying the computer program product of claim 22.
EP24716763.8A 2023-04-05 2024-04-03 Method for generating at least one synthetic image Pending EP4690125A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
LU503856A LU503856B1 (en) 2023-04-05 2023-04-05 Method for generating at least one synthetic image
PCT/EP2024/058975 WO2024208847A1 (en) 2023-04-05 2024-04-03 Method for generating at least one synthetic image

Publications (1)

Publication Number Publication Date
EP4690125A1 true EP4690125A1 (en) 2026-02-11

Family

ID=86330221

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24716763.8A Pending EP4690125A1 (en) 2023-04-05 2024-04-03 Method for generating at least one synthetic image

Country Status (3)

Country Link
EP (1) EP4690125A1 (en)
LU (1) LU503856B1 (en)
WO (1) WO2024208847A1 (en)

Also Published As

Publication number Publication date
LU503856B1 (en) 2024-10-07
WO2024208847A1 (en) 2024-10-10

Similar Documents

Publication Publication Date Title
CN112036447B (en) Zero-shot object detection system and fusion method of learnable semantics and fixed semantics
CN113192040A (en) Fabric flaw detection method based on YOLO v4 improved algorithm
KR101640998B1 (en) Image processing apparatus and image processing method
CN113076804B (en) Target detection method, device and system based on YOLOv4 improved algorithm
CN114820579A (en) Semantic segmentation based image composite defect detection method and system
CN120226039A (en) System and method for joint detection, localization, segmentation and classification of anomalies in images
CN110991435A (en) A method and device for locating key information of express waybill based on deep learning
CN109977997A (en) Image object detection and dividing method based on convolutional neural networks fast robust
CN112581462A (en) Method and device for detecting appearance defects of industrial products and storage medium
CN110781882A (en) License plate positioning and identifying method based on YOLO model
CN113688709A (en) A safety helmet wearing intelligent detection method, system, terminal and medium
CN111091101B (en) High-precision pedestrian detection method, system and device based on one-step method
CN115775220A (en) Method and system for detecting anomalies in images using multiple machine learning programs
CN113744280B (en) Image processing method, device, equipment and medium
CN117315387A (en) An industrial defect image generation method
US20250037255A1 (en) Method for training defective-spot detection model, method for detecting defective-spot, and method for restoring defective-spot
CN118521651A (en) Image processing method, device, system and storage medium based on artificial intelligence
CN112150398A (en) Image synthesis method, device and equipment
CN116452947B (en) A cross-domain fault detection method based on multi-scale fusion and deformable convolution
CN116977791A (en) Multi-task model training method, task prediction method, device, computer equipment and media
EP4690125A1 (en) Method for generating at least one synthetic image
CN121213468A (en) A method and system for chromosome karyotype analysis
KR102621884B1 (en) Deep learning-based image analysis method for selecting defective products and a system therefor
CN109902751A (en) A dial digital character recognition method combining convolutional neural network and half-word template matching
CN114494212A (en) Aluminum material surface defect detection method, device and storage medium

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20251101

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR