"Scanner for insect damage" Cross-Reference to Related Applications [0001] The present application claims priority from Australian Provisional Patent Application No 2023901869 filed on 13 June 2023, the contents of which are incorporated herein by reference in their entirety. Technical Field [0002] This disclosure relates to detecting insect damage in a plant product. Background [0003] Plant products, such as cherries and blueberries, for example, are prone to damage from insects, such as fruit flies. It is important to detect such damage in order to either take the damaged product out of the supply chain or to prevent further spread of the insects. However, it is difficult for the human eye to detect insect damage at an early stage because the early markers on the product are very small. Summary [0004] This disclosure provides a method for detecting insect damage of a plant product reliably by illuminating the plant product with a narrow-band light source. This improves the detection rate significantly over other methods. [0005] A method for detecting insect damage in a plant product comprises: illuminating the plant product with a narrow-spectrum light to reduce reflections from features of the plant product that distract detection of insect damage; capturing an image of the plant product illuminated with the narrow spectrum light; applying a trained object detection machine learning model to the image of the plant product, wherein
the machine learning model is trained on training images, the training images being associated with labels indicative of insect damage, and the machine learning model is configured to detect objects in the image and the objects relate to areas of the plant product where insect damage is present; in response to the machine learning model detecting an object, outputting an indication that insect damage is detected in the plant product. [0006] It is an advantage that the method uses narrow-spectrum light that reduces reflections from features that distract the detection of the insect damage. As a result, the machine learning model has a reduced false positive rate. It is a further advantage that the machine learning model is for object detection. This way, the spatial features of the insect damage are used instead of spectral features for example. It has been found that those spatial features enable a more accurate detection. [0007] In some embodiments, the method further comprises training the machine learning model by: using a part-trained machine learning model to generate a putative bounding box in a training image; presenting the putative bounding box and the training image to a user on a user interface; receiving from the user an indication of insect damage within the putative bounding box; using the received indication as a label for the training image to train the machine learning model. [0008] In some embodiments, the machine learning model is trained to detect spatial features indicative of entomological characteristics of the insect damage. [0009] In some embodiments, the machine learning model is trained to detect spatial features surrounding a puncture within the oviposition site.
[0010] In some embodiments, the machine learning model is trained to detect spatial features comprising coloration of the oviposition site as a result of injection by the insect during oviposition into the puncture. [0011] In some embodiments, the spatial features comprise an annular oviposition site. [0012] In some embodiments, the narrow-spectrum light is monochromatic light. [0013] In some embodiments, the narrow-band light is near-infrared light. [0014] In some embodiments, the near-infrared light has a wavelength of 730 nm or 765 nm or 785 nm or 850nm. [0015] In some embodiments, capturing the image comprises using a monochrome camera to capture a greyscale image. [0016] In some embodiments, the machine learning model comprises a neural network comprising one or more convolutional layers applied to the image. [0017] In some embodiments, the machine learning model is configured to predict bounding boxes around insect damage sites. [0018] In some embodiments, the machine learning model is trained to distinguish between a spatial pattern of insect damage and other visual features of the plant product. [0019] In some embodiments, the plant product is a fruit. [0020] In some embodiments, the fruit has a skin and the insect damage comprises oviposition into a puncture of the skin of the fruit.
[0021] In some embodiments, the insect damage to the fruit comprises damage by bactrocera tryoni. [0022] In some embodiments, the method further comprises capturing a video of the plant product; applying the trained object detection machine learning model to each of multiple frames of the video; and outputting the indication in response to the machine learning model detecting an object in at least one of the multiple frames. [0023] In some embodiments, the method further comprises rotating the plant product while capturing the video to capture frames from multiple different viewing angles. [0024] In some embodiments, the trained object detection machine learning model is configured to detect objects faster than the frames of the video are captured. [0025] In some embodiments, the trained machine learning model is configured to provide real-time classification. [0026] In some embodiments, the plant product is a cherry and the narrow-spectrum illumination has a wavelength of 730 nm. [0027] A computer-implemented method for detecting insect damage in a plant product comprises a processor performing the steps of applying a trained object detection machine learning model to an image of the plant product, the image being captured under a narrow-spectrum illumination to reduce reflections from features of the plant product that distract detection of insect damage, wherein the machine learning model is trained on training images, the training images being associated with labels indicative of insect damage, and the machine learning model is configured to detect objects in the image and the objects relate to areas of the plant product where insect damage is present; and in response to the machine learning model detecting an object, outputting an indication that insect damage is detected in the plant product. [0028] A scanner for detecting insect damage in a plant product comprises:
a light source configured to illuminate the plant product with a narrow- spectrum light to reduce reflections from features of the plant product that distract detection of insect damage; camera configured to capture an image of the plant product illuminated with the narrow spectrum light; and a processor configured to: apply a trained object detection machine learning model to the image of the plant product, wherein the machine learning model is trained on training images, the training images being associated with labels indicative of insect damage, and the machine learning model is configured to detect objects in the image and the objects relate to areas of the plant product where insect damage is present; in response to the machine learning model detecting an object, outputting an indication that insect damage is detected in the plant product. Brief Description of Drawings [0029] Figure 1 illustrates a scanner for detecting insect damage in a plant product, according to an embodiment. [0030] Figure 2 illustrates a method for detecting insect damage in a plant product according to an embodiment, according to an embodiment. [0031] Figure 3 shows a broadband illumination spectrum and a narrow-band illumination spectrum, according to an embodiment. [0032] Figure 4a shows three putative bounding boxes around respective spatial features and Figure 4b shows the entire image with the three putative bounding boxes. [0033] Figure 5 shows the architecture of YOLOv1, according to an embodiment. [0034] Figure 6 is a cropping algorithm, according to an embodiment.
[0035] Figure 7 illustrates how to obtain the range of the top-left coordinate (x p,y p) of a patch given a bounding box (x,y,w,h), according to an embodiment. [0036] Figure 8 shows the architecture of YOLOv3, according to an embodiment. [0037] Figure 9 illustrates label prediction for a cherry image, according to an embodiment. Description of Embodiments [0038] Figure 1 illustrates a scanner 100 for detecting insect damage in a plant product 101. The plant product may be any product that is derived from a plant and may include fruit, sprouts, nuts, seeds, legumes, roots, leaves, flowers and other plant products. In a particular embodiment, the plant product is a fruit from the Rosaceae family of plants, such as apples, pears, quinces, apricots, plums, cherries, peaches, raspberries, blackberries, loquats, strawberries, rose hips, hawthorns, and almonds and other fruit. [0039] Insect damage may be caused by biting, stinging, puncturing or other damage. In a particular embodiment, the insect damage is the damage of puncturing the skin of the fruit for oviposition, where the insect creates an opening in the skin to lay its eggs into the fruit. [0040] Different insects may cause damage to the plant product, which is then detected, including Aphid, Codling moth, Scales, Brown marmorated stink bug, Apple maggot, Spider mite, Japanese beetle, Peachtree borer, Plum curculio, Oriental fruit moth, Tortrix moths, Thrips, San Jose scale, Mealybug, Leafhoppers, Woodboring beetle, Weevil, Spotted wing drosophila, Citrus leafminer, Mediterranean fruit fly, Red spider mite, Earwigs, Light brown apple moth, Tarnished plant bug, or others. [0041] In one embodiment, the insect damage is a puncture of the skin of a fruit by a Queensland fruit fly, or Bactrocera tryoni, for oviposition. Typically, the damage is so
small that it cannot be reliably detecting by the naked human eye. Therefore, this disclosure provides a method for a more robust detection of insect damage that does not require slow examination through a microscope by a human inspector. System for detecting insect damage [0042] Scanner 100 comprises a light source 102. Light source 102 is configured to illuminate the plant product 101 with a narrow-spectrum light. This reduces reflections from features of the plant product that distract detection of insect damage. More specifically, the narrow-spectrum light reduces reflections from pigmentation, mechanical damages or bruising of the plant product, which distracts the detection of the insect damage especially when using a trained neural network. [0043] This disclosure provides below a number of wavelengths that have been shown to be advantageous in improving the detection accuracy by reducing reflections from features of the plant product that distract the detection of the insect damage. In other words, those distracting features have a reflectivity at a certain wavelength and if the illumination does not contain that wavelength, those features do not reflect any light. This is similar to illuminating a green surface with red light, the result would be a black image under perfect conditions. [0044] Scanner 100 further comprises a camera 103 configured to capture an image of the plant product 101 illuminated with the narrow spectrum light 102. The camera 103 may be a consumer grade three channel RGB camera and may have a typical resolution of several megapixels, such as a common single lens reflex (SLR). In that example, only one of the three channels may be used for further processing. For near infrared illumination, only the red channel may be used. In other embodiments, the camera is a wide-spectrum monochromatic (greyscale) camera. Since the illumination is narrow- band, the reflected light is also narrow-band and therefore, camera 103 only captures a narrow band despite being a wide-spectrum camera. This means, the distracting features are suppressed despite using a relatively low-grade camera, which makes scanner 100 relatively easy to set-up and to maintain. Further, no colour correction, constancy or colour balance is required, which can be a problem with multispectral
cameras in non-ideal lighting conditions. Even further, the camera 103 may be a complementary metal oxide semiconductor (CMOS) or charge coupled device (CCD) camera with or without a colour filter for each pixel. In the case where the camera 103 is monochromatic, there is a further advantage that the individual pixels do not need to be filtered for individual colours, which results in a higher light sensitivity and no need for de-Bayering. [0045] Camera 103 is communicatively coupled (such as by universal serial bus (USB), Wifi, or other wired or wireless connection) to a computer system 104. Computer system 104 comprises a processor 105, program memory 106, and data memory 107. Program memory 106 is a non-transitory computer readable medium, such as a solid state drive (SSD), hard disk drive (HDD) or read only memory (ROM). The data memory 107 may also be non-volatile to persistently store data, such as image data, parameters of trained machine learning models, configuration data and other data. [0046] The program memory 106 stores program data, such as source code, scripts or compiled binary code, that, when executed by processor 105 causes the processor to perform the methods disclosed herein. This way, the processor 105 is configured, by the program data, to perform the methods disclosed herein. In other examples, the computer system 104 is configured to send the image data to a cloud server or other computer system for processing, such as classification or for training of a machine learning model. [0047] Oviposition-damaged and undamaged cherries were photographed using a digital auto-montage imaging system consisting of a Nikon SMZ25 microscope and a Nikon DS-Fi2 camera. Additionally, two panels of 730nm light-emitting diode (LED) were used for illumination. Each LED containing 18 × 1W diode with 60◦ lens attached was placed on both sides, 80mm away from the cherries. The 730nm of the light spectrum was chosen based on research (such as hyperspectral image analysis), which showed that oviposition sites were more visible on cherries between a 700-750nm spectral range. In order to use DS-Fi2, a colour camera, for near infrared imaging, the infrared filter in front of the sensor was removed and only the red channel was saved as
a monochrome image. The microscope and the camera settings used to take the images were magnification: 0.32-0.5×, exposure: 8ms, anolog gain: 1.2×. Images from four sides of an individual cherry were taken. [0048] In another embodiment, which relates to video recording as described below, the system uses the following parameters: Image library specifications: Camera: GS3-U3-41C6M-C Sensor: CMOSIS CMV4000 Channel: Monochrome Pixel count: 4.2M Pixel depth: 8 bit Resolution: 2048x2048 Lighting: 3 panels of NIR LED 730/765/785/850nm Lens focus length: 25mm Lens aperture: f/8 Exposure: <1ms Gain: 1-3 out of 10 Background: Niryo Light Duty Conveyor Belt (black faux leather) + black matt paint roller Label format: two-point rectangular boxes of infested sites (YOLO format) Method for detecting insect damage [0049] Figure 2 illustrates a method 200 for detecting insect damage in a plant product 101. According to method 200, processor 105 applies 201 a trained object detection machine learning model to an image of the plant product 101. The image is captured under a narrow-spectrum illumination 102 to reduce reflections from features of the plant product that distract detection of insect damage. That is, the wavelength (e.g. the centre of the spectral band) of the narrow-spectrum is chosen, such that the reflections are reduces that are from features of the plant product that distract detection of insect damage.
[0050] Figure 3 illustrates a light spectrum of sunlight 301, which spreads over the entire wavelength range of visible light from about 380 nm to about 750 nm. As a result, when sunlight illuminates the plant product, there will be a reflection from insect damage (if present) and a reflection from other features that distract the detection of the insect damage, such as pigmentation. Figure 3 also shows a narrow-band spectrum 302, which spreads over a much smaller range of the wavelength than sunlight spectrum 301. As a result, it has been found that the reflections from features that distract detection of insect damage are suppressed in the reflection image. Examples given herein provide a specific wavelength, such as 360 nm, in which case the light is referred to as monochromatic light because it only contains light of that wavelength. However, it is to be understood that as a physical reality, this may refer to a small range around that specific wavelength, such as 10 nm either way, so from 350-370 nm. Other narrow-band ranges include 20 nm or 5 nm either way. This also applies to other wavelengths provided herein. In one example, the wavelength is in the near-infrared spectrum from 780 nm to 2500 nm as spectrum 302 in Figure 3. Example wavelengths include 730 nm, 765 nm, 785 nm and 850 nm (including the variations discussed above). [0051] The description below provides exemplary embodiments for particular plant products. Otherwise, it is possible to test a number of different wavelengths using a hyperspectral camera with wide-band illumination, for example, or sweep across the wavelength range from near infrared to ultra-violet, train the machine learning model and determine the accuracy of the trained machine learning model. Then, the best performing wavelengths can be selected for using the corresponding machine learning model going forward. This wavelength selection process can be performed for each different type of plant product 101, such as once for cherries, once for pears, and so on. That is, a different wavelength may be selected for a different plant product. Labelling [0052] For the manual labelling of the images, a tool called Labelimg can be used, which can be found at this GitHub link: https://github.com/tzutalin/labelImg. The task
of labelling may be carried out by an entomologist. During the labelling process, two types of spots on cherries are identified: healthy spots and damage spots. To accurately identify the pattern of insect damage, the fruits may be dissected specifically at the damage spots. The entomologist then describes the pattern of damage and proceeds to draw bounding boxes around the respective spots on the images. Labels are assigned to these bounding boxes based on the type of spots identified. In essence, each bounding box serves as a region of interest, providing crucial information about the location and class of the object within the image. There are two types of spots on cherries, healthy spots and damage spots. A user draws bounding boxes around the spots and give labels to bounding boxes based on the type of spots. Those bounding boxes in the image are saved in a text file. In general, a bounding box contains the following information: coordinates of the centre point, width, height, and class of the object in the bounding box. [0053] Training a machine learning model typically relies on a large number of training data, such as training images. Those training images are labelled so that the training process can calculate a difference between the label of the image and the output of the machine learning model. The training process then adjusts the parameters of the model to minimise this difference, such as by gradient descent and back propagation. [0054] However, it is often difficult to obtain a large number of training images because they are often labelled by human experts and so the training process is long and expensive. To address this challenge, this disclosure provides an improved labelling method that can generate training images more quickly. First, a relatively small number (e.g.100) of labelled training images is used to part-train the model. It is noted, in this application, that the labelling would preferably be performed by an entomologist, who is an expert in insect behaviour. Therefore, an entomologist is qualified to distinguish insect damage from other features or other damage. [0055] Once the model is part-trained, it can be used to generate a putative bounding box in a new training image. This bounding box is referred to “putative” because it is
generated by a part-trained model. Figure 4a shows three putative bounding boxes around respective spatial features and Figure 4b shows the entire image with the three putative bounding boxes. At this stage, the model is relatively accurate in detecting objects but it is not yet accurate in classifying those objects into insect damage or other features. So the bounding boxes are around features that are insect damage and features that are not. So at this point, the bounding box indicates a detected object (i.e. spatial feature) but it is yet to be confirmed that the putative bounding box indicates insect damage. [0056] In a next step, a user interface, such as a web-interface or other interface displayed on a computer screen, or touch screen of a smartphone or tablet device (or any other device) presents the putative bounding box and the training image to an expert user. The user interface may comprise user control elements, such as buttons or input fields, to enable the user to provide feedback to the training process. In other examples, the user input is by way of input devices, such as pressing certain letters on a keyboard. That is, the user indicates by way of the user input (e.g., tapping or clicking a button on the screen or pressing a button on the keyboard), whether the putative bounding box actually contains insect damage in the image. [0057] As a result, the training process receives from the user an indication of insect damage within the putative bounding box. This indication may be binary or in another format. The training process then uses the received indication as a label for the training image to train the machine learning model. That is, the training process uses the image just like any other labelled training image. [0058] It is noted that this image had not been labelled previously and the part-trained model created a putative bounding box. For the human expert, it is now significantly faster to confirm insect damage than to create a label for a new training image without the putative bounding boxes. Therefore, more training images can be generated and are then available for training. As a result, the trained machine learning model has a higher accuracy compared to a model trained on fewer images.
[0059] As set out elsewhere herein, the machine learning model is trained to detect spatial features. Since the model is trained on entomological characteristics, such as insect damage, the spatial features, that the model is trained to detect, are indicative of entomological characteristics of the insect damage. This is an advantage because the model now embodies expert knowledge that is only available to experts trained in this field. [0060] More particularly, for detecting insect damage, the expert entomologist labels the spatial features dependent on whether the features surround a puncture that the insect created for oviposition. As a result, the machine learning model is trained to detect spatial features surrounding the puncture within the oviposition site. This accurately captures the typical appearance of a surrounding damage. [0061] For example, the skin of the product is coloured differently to the healthy skin, such as having a brown colour. Since the expert labels features showing this characteristic, the machine learning model is trained to detect spatial features comprising coloration of the oviposition site. This coloration is the result of an injection by the insect during oviposition into the puncture. The insect injects a substance into the oviposition site that includes or may support bacteria. The injected substance and/or the bacteria create a discolouration in the affected skin, which is often ring-shaped (i.e. annular). This discolouration means that light is reflected differently from the affected skin. As set out herein, the spectral reflectivity, that is the wavelength of reflectivity, is specific to the affected skin compared to other damage or other features that are not insect damage. Therefore, illuminating the product with this wavelength suppresses features that distract the detection of insect damage. Training [0062] The machine learning model is trained on training images and the training images are associated with labels indicative of insect damage as described above. That is, the training images are labelled independently from the training process, such as by human inspection of the training images, which is referred to as supervised learning. It is also possible to label each training image dependent on whether the fruit develops
further damage after the training image was captured. For example, in cases where the insect damage is too small to be recognised by the human eye, the training images can be captured for a number of plant products, then the plant products can be stored for a period of time, such as 1 or 2 weeks at room temperature, and then each plant product can be more readily examined whether any insect larvae have developed, for example. Then, the associated historical images of the infected fruit can be labelled retrospectively and used for training the machine learning model. Object detection [0063] It is noted that the machine learning model is configured to detect objects in the image and the objects relate to areas of the plant product where insect damage is present. This is in contrast to general classification models, such as neural networks that provide a classification output for a large number of general inputs, such as pixels directly. Here, the machine learning model is an object detection model that uses the image data to detect two-dimensional objects in the image data. More particularly, the machine learning model is trained to generate bounding boxes around objects. That is, a human user draws bounding boxes around objects in training images and those training images are fed into the machine learning model. The parameters are then adjusted to minimise an error in generating bounding boxes. The machine learning model is further trained to distinguish between the spatial pattern of insect damage on one hand and other visual features of the plant product, such as pigmentation, on the other hand. That is, the bounding boxes created by the human user are labelled as “damage” or “other” and the machine learning model is trained to correctly predict those labels. When applied to a test image, the machine learning model will predict a bounding box of an insect damage object. [0064] The machine learning model can equally be applied to frames of a video taken of the product. This can be useful when the product is rotated during the video so that objects on all sides of the product can be detected with a single camera. In that case, it is advantageous to use a machine learning model that can be evaluated for each frame relatively fast, so that the output is available before the next frame is generated. In other
words, the calculation of the model output is faster than the framerate of the video. It is noted that a higher false negative rate may be acceptable in this example because multiple frames are being analysed and it is sufficient to have an insect damage in a single frame to discard the product. [0065] An example machine learning model is a neural network with one or more convolutional layers applied to the image, also referred to as convolutional neural network (CNN). A CNN comprises two-dimensional filters that are stepped over the input image. The filter parameters are trained to detect objects. Because the filters are two-dimensional, the can preserve the spatial features in the image because pixels belonging to the same object are located next to each other, i.e. they are spatially connected. This can be captured by the spatial two-dimensional filters and corresponding feature maps. As a result, the object detection is very accurate. In combination with the narrow spectrum illumination, the amount of false positives is very low. This is particularly the case for damage by insects that create a set pattern when they damage the plant product. [0066] In response to the machine learning model detecting an object, such as the pattern, processor 105 outputs 202 an indication that insect damage is detected in the plant product. For example, once the output of the neural network indicates that an object is detected, which means that insect damage is present in the image, processor 105 outputs an indication accordingly. This indication may be an indication on a graphical user interface, such as an visual alert, an audible alert, a message, such as text or email, or a text document. Processor 105 may generate a report with a cumulative or average number of products being subject to insect damage. Further, the indication may be a control signal to control another device or system to process the indication. One example is a blade actuator 108 on a conveyor 109. The control signal controls the blade actuator 108 or other deviating or sorting device, to remove the plant product 101 from the conveyor 109 if the processor 105 indicates that insect damage has been detected. Other methods for sorting the plant products into those with insect damage and those without could equally be used.
[0067] The disclosed approach is extremely sensitive to detect minor level and early stage of Queensland fruit fly (Qfly) infestation. Other NIR (Near Infrared) spectroscopy relies on the detection of changing chemical composition of the infested fruit, therefore robust detection with that technology requires higher level or late stage of infection. The approach disclosed herein can offer much higher detection accuracy for fruits with only one oviposition damage, which is common under natural conditions in the orchards. This is achieved by the disclosed narrow-spectrum illumination in combination with object detection because the narrow-spectrum illumination suppresses features that would otherwise distract the neural network. Further, this disclosure provides an advantageous way for labelling images of plant products. This means that a large set of training images can be generated relatively easily. The accuracy of most machine learning models hinges on the number of training images and therefore, the disclosed method can achieve better results due to a large training set. Further, this system could be applied in the pack house or at the point of care inspection through automated fast sorting system instead of human visual inspection which is time consuming and not accurate in most of the cases. Machine learning model [0068] The detection may be accomplished by training Object Detection Machine Learning model You Only Look Once (YOLO) V3 using image libraries with damage labelled as spatial structure, which result in a damage detection model that can be used for real-time scanning detection with rotating fruits on the conveyor belt. Details on YOLOv3 can be found in J. Redmon, and A. Farhadi, ‘‘YOLOv3: An incremental improvement,’’ 2018, arXiv:1804.02767. [Online]. Available: https://arxiv.org/abs/1804.02767, which is incorporated herein in full by reference. This model uses successive 3 × 3 and 1 × 1 convolutional layers including shortcut connections. It has 53 convolutional layers and is therefore referred to as Darknet-53. The table below provides the Darknet-53 architecture from the above paper:





[0069] In another embodiment, the model comprises Mini-YOLOv3 as described in Mao, Qi-Chao, et al. "Mini-YOLOv3: real-time object detector for embedded applications." IEEE Access 7 (2019): 133529-133538., which is incorporated herein in full by reference. This model is also based on the Darknet-53 architecture comprising a feature extraction backbone network with a parameter size of only 16 percent of Darknet-53. The model further comprises a multi-scale feature pyramid network based on a U-shaped structure to improve performance. [0070] In yet another example, the model comprises YOLOX as described in Ge, Zheng, et al. "Yolox: Exceeding yolo series in 2021." arXiv preprint arXiv:2107.08430 (2021). which is included herein in full by reference. YOLOX has an anchor-free detector and conduct other advanced detection techniques, i.e., a decoupled head and the leading label assignment strategy SimOTA. This may include various different backbones, such as Darknet-53, CSPNet, Tiny and Nano detectors.
YOLOv3 details [0071] YOLOv3 is a one-stage model and has a good ability to detect small objects, which makes it suitable for the detection of insect damage under narrow-spectrum illumination . Furthermore, YOLOv3 is invariant to the input size. It enables the model to take arbitrary sizes of images as input. Figure 8 shows the architecture of YOLOv3. The new backbone of YOLOv3 has been named ”Darknet-53” which integrates residual blocks. In addition, YOLOv3 can make multi-scale predictions due to the application of the idea of feature pyramid network. Each scale focuses on predicting objects of a particular size. For example, scale one is responsible for detecting big objects, while scale three is to detect small objects. As a result, YOLOv3 has a good ability of detecting small objects. Each grid cell in each scale predicts three bounding boxes in this case. The final prediction is obtained by merging the outputs of three scales. [0072] The below image-level evaluation results of YOLOv3 on the testing set are provided noting that the images in all experiments were acquired under the narrow- spectrum illumination:

[0073] Since YOLOv3 is invariant to the input size, it is possible to use the original images to evaluate the model as well. In that case, the model performance may drop. There are 11 more false positives by using the original images to test the model when the confidence score threshold is set to 0.4. Furthermore, the specificity and F1 score dropped significantly. However, when the threshold is set to 0.5, the number of false positives has decreased to 14 from 22. Also, the specificity and F1 score increased significantly.
[0074] In one example, the method employs a machine learning approach to refine the predictions of YOLOv3. The evaluation result shows that the number of false positives decreased significantly. [0075] Figure 9 illustrates label prediction method 900 for an image of a plant product. An image 901 is fed into YOLOv3902 which predicts a large number of bounding boxes 903, then non-maximum suppression 904 is applied to remove the highly overlapped bounding boxes as well as bounding boxes with confidence scores that are smaller than the predefined threshold. Finally, the label is determined by checking the existence of infested bounding boxes 905. The proposed prediction refinement approach 910 is shown in the middle panel. It takes the predictions of YOLOv3 model 911 as the input and removes some bounding boxes 912. Then based on the remained bounding boxes 913, the feature generator 914generates a feature vector 915. Finally, the feature vector is fed into machine learning classifier 916 to get the label. The bottom panel shows the feature generator. [0076] There are some motivations behind the introduction of this prediction refinement approach. Previously, when the evaluation result is available, the method uses the same method to obtain the image level label for each image in the testing set. For example, if the method determines that there is no infected bounding box in the label file for an image, then the ground-truth image-level label for this image is healthy, otherwise it is infested. In the evaluation phase, as is shown in the top panel of Figure 9, if there is one or more infested bounding boxes in the prediction, the predicted image-level label will be infested. Under this circumstance, this is a false positive prediction. The problem may lead to too many false positives, so it is desirable to improve the way the method determines the image-level label for prediction. Hence, there is proposed an prediction refinement approach. The bottom panel of Figure 9 briefly shows the approach. [0077] To get the training data for training a Machine Learning (ML) classifier, the method uses the feature generator 914 to generate feature vectors for images. More specifically, predicted bounding boxes of an image by YOLOv3 will be converted into
a feature vector by the feature generator 914. Note that the label of the feature vector is inherited from the image. Then the training set is fed into YOLOv3. Based on the predictions of YOLOv3 on the training set, the method uses the feature generator to generate feature vectors. Eventually, these feature vectors and the corresponding labels will construct the training data for the ML classifier. The testing data will be generated in the same way. It is worth noting that in this experiment the image patches were fed into the YOLOv3 first, then all the predicted bounding boxes of patches that are from a same original image are merged into one set. Then the first step of prediction refinement approach starts by taking these bounding boxes. [0078] The feature generator 914 is a component in the prediction refinement approach. As is shown in the bottom panel in the Figure 9. It first extracts the confidence 921, class 0, and class 1 scores 922, then generate three histograms 923 and corresponding histogram vectors 924, respectively. Finally, it concatenates 925 all the vectors into one which becomes the final feature vector 926. [0079] Once the training and testing data is available, then the training of a machine learning classifier is the next step. In this example, a Support Vector Machine (SVM) is used as the classifier. [0080] Image-level evaluation results of YOLOv3 and prediction refinement approach on the testing set:

[0081] This table shows the comparison of the evaluation results of YOLOv3 and prediction refinement approach. The number of false positives has decreased significantly. The specificity and F1 scores improved, while the sensitivity decreased.
The results demonstrate that prediction refinement approach leads to a good trade-off between sensitivity and specificity. [0082] While some examples herein relate to YOLO and in particular, YOLOv3, it is noted that other object detection models are equally applicable. One such alternative is DeNet. Figure 10 illustrates the DeNet architecture, which uses directed sparse sampling (DSS) for object detection using a two-stage convolutional neural network (CNN). The first stage of the network estimates the probability of each pixel in the input image being a corner of a bounding box that contains an object of interest. The second stage sparsely samples a fixed number of bounding boxes from the corner distribution and classifies them using a feature vector extracted from the feature map of the first stage. The network is trained end-to-end using a loss function that jointly optimizes over the corner probability, the final classification, and the bounding box regression. The network is based on the ResNet-34 or ResNet-101 models, with some modifications to increase the spatial resolution of the feature map and the corner detector. [0083] The following table shows example filter parameters for DeNet models:

[0084] The following table provides parameters for DeNet, as a ResNet derived model for DSS Object Detection with a 512×512 input image. Layers in the base models above the line are initialized with a pretrained ResNet-34 or ResNet-101 ImageNet 2012 classification model.
[0085] The additional layers appended to the base models have the following definitions: • Conv: Convolves a series of 2D filters over the in- put activations. Filter weights were initialized via the normal distribution N (0, σ) with σ
2 = 2/(nfnxny) where nf is the number of filters and (nx, ny) their spatial shape. Following each convolution is batch normalization then the ReLU activation function. • Deconv: Applies a learnt deconvolution (upsampling) operation followed by ReLU activation. In this case it is akin to upscaling both spatial dimensions then applying a Conv layer. • Corner: Estimates a corner distribution via the soft-max function and produces a sampling feature map. • Sparse: Identifies sampling bounding boxes from corner distribution and produces a fixed size sampling feature from the sampling feature maps. • Classifier: Maps activations to the desired probability distribution via the softmax function and generates bounding box targets. [0086] For DeNet-34 we use a ResNet-34 base model and Fs = 96 to produce a feature vector of 4706 values and a total of 32M parameters. The DeNet-101 model uses a ResNet-101 base model and increased the number of filters by approximately 1.5× for the appended layers. These changes produce a sparse feature vector of 6274 values and a total of 69M parameters.
Label Quality Assessment and Improvement [0087] The experiment has decreased the false positive rate significantly. However, to further improve the model performance, the method performs label quality assessment. The label quality is another factor that may cause low model performance. There are two types of bounding boxes in the label files. A trained YOLOv3 model is used to make predictions for all the folders with image patches. The prediction and ground truth are plotted together as long as the number of the infected bounding boxes are different. For example, in an image patch, there are three infested bounding boxes in the label. If the number of infested bounding boxes is 2 in the prediction, then the prediction and ground truth will be plotted. This approach can effectively identify the majority of the bounding boxes with low quality. [0088] After inspection of the low quality labels, it was found that some bounding boxes are too small so that they do not cover the whole spots. Additionally, it was realised that there are uncertainty whether some spots are infested or healthy. To solve this problem, all the labels were inspected, and all the bounding boxes with low quality relabelled. Also, the third class (uncertain) of bounding boxes was introduced for the future use. The method treats the third class bounding boxes as healthy bounding boxes because the percentage of the third class bounding boxes is very low. After improving the labels of the datasets the method regenerates the image patches, and extends the dataset by augmenting more data. Distributed Training for YOLOv3 on Refined and Extended Dataset [0089] After relabelling the data and extending the datasets by more augmentation the YOLOv3 model is trained again. The data set was significantly extended. This time there are around 110,000 image patches in the training set and 14,000 image patches in the validation set. In the previous experiments, there were only around 37,000 and 800 image patches in the training and validation set respectively. The YOLOv3 model was trained on the single GPU adopted in the previous experiments. To decrease the overall training time with a larger dataset, the method applied distributed training strategy in this experiment. Two GPUs were used for model training. The DistributedDataParallel
module in Pytorch was used to implement the distributed training strategy. Even there were two GPUs for training, it still took around 27 minutes to complete one epoch. The model was trained for 120 epochs which took us 54 hours to finish. [0090] The table below shows the result comparisons between the new YOLOv3 model (YOLOv32) and the old YOLOv3 model. There is a remarkable improvement brought by YOLOv32. When the confidence score threshold is set to 0.5, the number of false positives has decreased to 3 from 14, and the number of true negatives increased to 59 from 46. However, the number of false negative increased to 7 and the number of true positive decreased to 15. Consequently, the specificity increased obviously, and the sensitivity dropped by 20%. The results are fantastic when the threshold set to 0.4. The sensitivity increased by 3%, and the specificity increased by 21%. The overall improvement of F1 score has reached 24%. These results show that this is an object detection model with high sensitivity and high specificity. [0091] Image-level evaluation results of YOLOv3 and YOLOv32 on the testing set. Previously, there were 24 infested images in the old set , the number dropped to 22 in the new set:

[0092] The table below shows the evaluation results of two YOLOv3 models on efficacy. There are no infected images in this folder, so we only have specificity as the metric. When the threshold is 0.4, new model has 48 false positives, while the old model has 136 false positives. The specificity of the new model is 27% higher than that of the old model. When the threshold is 0.5, the performance of both models improve.
The new model only has 15 false positives. By contrast, the old model still has 99 false positive. [0093] Image-level evaluation results of YOLOv3 and YOLOv32 on the testing set (efficacy healthy):

Prediction Refinement Approach Ablation Study [0094] In the above experiment, a distributed strategy was used to train YOLOv3 on the refined dataset. It turns out that the new YOLOv3 (YOLOv32) model achieves a better performance. Since there is a significant improvement in the above experiment, the method can applies the prediction refinement approach on YOLOv32. [0095] In the process of generating the training and testing sets for the ML classifier, note that different from the above experiment, the method feeds the original images into YOLOv3 to generate predictions directly. [0096] ML classifiers used in ablation study:

[0097] The method may use three different ML classifiers in this experiment, the details of these models are shown in the table above. In experiment four, confidence and class scores are used to generate the feature vectors as is illustrated in Figure 9. In this experiment, only class 0 score is used to form the feature vector. Models with ”1” index were trained using feature vectors generated by using confidence and class scores, e.g., SVM 1. Models with ”2” index were trained using feature vectors generated by using class 0 score. [0098] Image-level evaluation results comparison on testing set:
[0099] Image-level evaluation results on the testing set (Efficacytest healthy no label):

[0100] The two tables above summarise the comparison results of difference refinement approaches and YOLOv32 and Efficacytest healthy no label. The first table shows that ML classifiers have better performance than YOLOv32. Except for SVM 1, other ML classifiers achieve the same results. The number of false positives decreases to 2 from 4, and the true positives remain the same level as YOLOv3 (th=0.4). [0101] In the second table, SVM 1 and two random forest classifiers have worse performance than YOLOv32 (th=0.5), but they have better performance than YOLOv3 2 (th=0.4). The performance of other ML classifiers are better than YOLOv32. Decision tree classifiers only have 12 false positives which is the best record. From this experiment, the SVM trained with feature vectors generated by only using class 0 score seems to have a consistent better performance than the SVM trained with feature vectors generated by using both confidence and class scores. [0102] This experiment shows that the refinement approach consistently improves the image level prediction accuracy by YOLOv32. For YOLOv32, when we set the threshold to 0.4, it obtained better performance than YOLOv32 with 0.5 threshold. However, the evaluation results on Efficacytest healthy no label turn out to be opposite, the 0.5 threshold works better. This is the disadvantage of using the approach (top panel in Figure 9), the method keeps tuning the threshold to get a better performance. However, when using the prediction refinement approach, it always gets the best accuracy that a YOLOv3 model can get without tuning the threshold as is shown in the above tables. [0103] It is noted that although the examples above provide specific details on machine learning models and parameters, other models and other parameters can equally be used, such as other CNNs.
Example wavelengths [0104] The disclosed imaging system may be multi-band and consist of one or several bands of NIR (Near Infrared) LED lighting and regular colour or monochrome camera. Compared to other multi-band imaging systems with several cameras with specific bandpass filters, the disclosed system uses only two or three cameras to cover all the viewing angle of the fruit at once. The system detects the damage under different spectral light with different frames in milliseconds while plant products pass the imaging chamber. This is an easy approach for classification of damage symptoms on images. This approach can be easily implemented on the current optical scanning system in the packhouses, in which no change is required for existing camera and imaging system in order to use the disclosed pest detection model. [0105] The NIR LED diode are installed under a lighting arch above the imaging chamber, which consists diodes of spectral range of 730-850nm, the system may use one or multiple bands depending on the fruit type. This system can be potentially applied to vastly varieties of Rosaceae fruits, i.e., apple, cherry, blueberry, peach, nectarine and apricot. Experiments [0106] We used Hyperspectral images (HIS) to identify a specific wavelength/s for each fruit type where the damage became visible. We identified specific pattern/s of damage caused by specific insect on taken images at identified band for each specific fruit type. For example, we discovered that Q-fly damage on cherries became visible at 730nm of light and the pattern of damage was distinguishable from the other type of damage (e.g. mechanical damages). The following table provides examples that were tested:


[0107] This disclosure provides a methodology in which, instead of using spectral analysis, we used image recognition with spatial analysis under optimal wavelength/s to detect damage pattern/s. This methodology has increased the level of accuracy and sensitivity for detecting insect pests on fruits. [0108] Then the images were taken at identified wavelengths and were labelled based on identified pattern confirmed by entomologists to generate the image library for each specific insect damage on each fruit types. The image library would be used as training data for machine learning of pest detection model and automated sorting. Testing the feasibility of developing a detection model have yielded promising results. To train the object detection model for detecting Q-fly damage on cherries, we generated a dataset of 1.2k images. For this attempt, allocated 468 images were allocated for training purposes and 91 images for testing. The model achieved an accuracy and sensitivity rate of over 90% when it came to detecting Q-fly damage at the image level.. [0109] The results of the detection model for cherry data , the training set contains 468 images and the testing set has 91 images: Table 3.2: Image-level evaluation results of YOLOX_I on the testing dataset (91 unhealthy images).

[0110] In another instance, For the Medfly and Northern Territory fruit fly, we generated over 10,000 images on blueberries and cherries. (listed in the table below)

Video processing [0111] As noted above, the disclosed methods capture an image of plant product 101. To improve the effectiveness of the approach, it is possible to use frames of a video as those images. Processor 105 can then apply the trained machine learning model to each frame. This way, as the plant product rotates or rolls, a single camera can detect damage on all sides of the product. Through the utilization of the video recording system, a substantial dataset consisting of 39,477 images was created derived from 575 infested cherries, each containing Queensland fruit fly eggs. Out of these images, 4,104 have been labelled for the purpose of training the model. This extensive dataset can be leveraged to develop a pre-labelling model. This model automates the labelling process, eliminating the need for manual labelling by human annotators. By automating this task, it is possible to significantly reduce the time and labour associated with the manual labelling process, which can be both time-consuming and arduous. With the implementation of the pre-labelling model, it is possible to streamline the overall process and achieve greater efficiency in our data labelling efforts.
Experimental details [0112] Cherries used in experiments were either purchased from a local supermarket or were supplied by the Queensland Department of Agriculture and Fisheries (QDAF) and washed with warm water and soap to remove any insecticide residue. Bruised or cherries with surface pittings were discarded to eliminate the possibility of B.tryoni ovipositing into damaged sites. To inflict oviposition damage, four cherries were placed into a small insect rearing tent (60×60×60 cm) containing 10014–18 day old male and female B. tryoni (obtained from a laboratory culture maintained by QDAF). The rearing tent was kept in an insectary at 26 ± 1◦C, 59% RH, L12:D12. Female B. tryoni were allowed to oviposit into the cherries for two minutes, after which, cherries were removed, and the same method repeated for another four cherries. Oviposition- damaged and undamaged cherries were then stored for either 24h, 48h or for 72h in an incubator (25 ± 1◦C, 67% RH). [0113] Oviposition-damaged and undamaged cherries were photographed using a digital auto-montage imaging system consisting of a Nikon SMZ25 microscope and a Nikon DS-Fi2 camera. Additionally, two panels of 730nm light-emitting diode (LED) were used for illumination. Each LED containing 18 × 1W diode with 60◦ lens attached was placed on both sides, 80mm away from the cherries. The 730nm of the light spectrum was chosen based on research (such as hyperspectral image analysis), which showed that oviposition sites were more visible on cherries between a 700-750nm spectral range. In order to use DS-Fi2, a colour camera, for near infrared imaging, the infrared filter in front of the sensor was removed and only the red channel was saved as a monochrome image. The microscope and the camera settings used to take the images were magnification: 0.32-0.5×, exposure: 8ms, anolog gain: 1.2×. Images from four sides of an individual cherry were taken. [0114] This disclosure provides an automatic computer vision-based technique for cherry quality control and for other plant products. Specifically, the goal is to identify plant products, such as cherries, with damages caused by fruit fly oviposition. The damages are usually very small on the surface of the products in early stage after
oviposition and with a particular pattern that can be recognized by human eyes with the help of a microscope. Human inspection is usually time-consuming and inconsistent. The assumption here is that the computer vision techniques can find products with symptoms efficiently and accurately as it can extract the feature representations of damages from a large amount of data. [0115] An object detection model can produce bounding boxes to show the location and size of different objects. In the present case, there are two classes of objects, healthy spots and damage spots. These two objects are detected with bounding boxes by an object detection model. This allows quick locating of the spots that are likely to be the damages and evaluating the model performance comprehensively. [0116] As mentioned above, the YOLO set of object detection networks can be used. For example, YOLOv1 is a one-stage object detection model which means it can be trained end-to-end. Figure 5 shows the architecture of YOLOv1, it takes an image with 448 × 448 × 3 size as input, then the image goes through six convolutional layers and one fully connected layer. The output of the fully connected layer is then reshaped to 7 × 7 × 30. YOLOv1 divides the image into 7 × 7 grid, each grid cell is responsible for the detection of an object that falls in it. From the output, each grid cell is a 30- dimension vector (B × 5 + C). In this case, B refers to the number of bounding boxes in the cell which is 2, C refers to the number of classes of the task which is 20. For each bounding boxes, it has five numbers which are confidence score, coordinates of centre point, width and height. In YOLOv1, each cell can predicts two bounding boxes, however, only the bounding box with a higher confidence score will be used. [0117] Image normalisation is a useful step for the training of deep learning models. It ensures the normalised pixels have a similar distribution, which makes the training converge fast. The method calculates the mean and standard deviation (std) of pixel values of the whole dataset. Afterwards, it performs every time before the images is fed into model for training or testing, as follows: x − mean z = std
[0118] In one example, the method uses an algorithm to crop the original images into 448 × 448 patches. Since bounding boxes are available in the ground truth images, the algorithm guarantees each bounding box from a particular image will fall fully in a patch. Algorithm 1 in Figure 6 shows the detail of the cropping algorithm. Figure 7 shows the idea of image cropping. The method applies the image patch cropping algorithm on captured images and obtains the image patches. The image cropping algorithm results in 5,450 image patches. In addition, to get sufficient image data for training, the method also performs data augmentation. For an example, three operations were performed on each image patch. These operations included 180 degree rotation, image mirror, and image flip. This results in 16,350 new image patches. [0119] One way to evaluate the performance of the models includes specificity, sensitivity and F1 score as the metrics. Since the disclosed method employs object detection, given an image of a product, such as cherry, the model can predict multiple bounding boxes on this image. If there exist any infested bounding boxes, the predicted image-level label for this image is infested. Otherwise it is healthy. Similarly, the ground-truth image-level label of an image can be determined by the label file. If there are any infested bounding boxes in the label file, then the ground-truth image-level label is infested. Otherwise it is healthy. When the ground-truth image-level label is infested, if the predicted label is infested. This prediction is a True Positive (TP), otherwise it is a False Negative (FN). When the ground-truth image-level label is healthy, if the predicted label is healthy, this prediction is a True Negative (TN), otherwise it is a False Positive (FP). The specificity of the prediction is calculated as: TN specificit y = TN + F P where ”TN” refers to the number of true negatives in the predictions. Specificity refers to the model’s ability to correctly detect healthy cherries. High specificity means the model can correctly detect the most of the healthy cherries among all the healthy TP cherries. The sensitivity of the prediction is: sensitivity = TP + FN
[0120] Sensitivity refers to the model’s ability to correctly detect infected cherries. High sensitivity means the model can correctly detect the most of the infected cherries among all the infected cherries. The F1 score is the combination of both precision and recall. It comprehensively evaluates the model performance, and is calculated as:

[0121] It will be appreciated by persons skilled in the art that numerous variations and/or modifications may be made to the above-described embodiments, without departing from the broad general scope of the present disclosure. The present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive.