WO2020134532A1 - 深度模型训练方法及装置、电子设备及存储介质 - Google Patents

深度模型训练方法及装置、电子设备及存储介质 Download PDF

Info

Publication number
WO2020134532A1
WO2020134532A1 PCT/CN2019/114493 CN2019114493W WO2020134532A1 WO 2020134532 A1 WO2020134532 A1 WO 2020134532A1 CN 2019114493 W CN2019114493 W CN 2019114493W WO 2020134532 A1 WO2020134532 A1 WO 2020134532A1
Authority
WO
WIPO (PCT)
Prior art keywords
training
model
trained
labeling information
information
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2019/114493
Other languages
English (en)
French (fr)
Inventor
李嘉辉
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Sensetime Technology Development Co Ltd
Original Assignee
Beijing Sensetime Technology Development Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Sensetime Technology Development Co Ltd filed Critical Beijing Sensetime Technology Development Co Ltd
Priority to JP2021507067A priority Critical patent/JP7158563B2/ja
Priority to KR1020217004148A priority patent/KR20210028716A/ko
Priority to SG11202100043SA priority patent/SG11202100043SA/en
Publication of WO2020134532A1 publication Critical patent/WO2020134532A1/zh
Priority to US17/136,072 priority patent/US20210118140A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/10Segmentation; Edge detection
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/10Text processing
    • G06F40/166Editing, e.g. inserting or deleting
    • G06F40/169Annotation, e.g. comment data or footnotes
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • G06N3/0455Auto-encoder networks; Encoder-decoder networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/0895Weakly supervised learning, e.g. semi-supervised or self-supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20081Training; Learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20084Artificial neural networks [ANN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30004Biomedical image processing
    • G06T2207/30024Cell structures in vitro; Tissue sections in vitro
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30196Human being; Person
    • G06T2207/30201Face

Definitions

  • the present disclosure relates to the field of information technology but is not limited to the field of information technology, and particularly relates to a deep model training method and device, electronic equipment, and storage medium.
  • the training set usually includes training data and labeled data of the training data.
  • labeling data requires manual labeling manually.
  • all the training data is labeled manually, which has a large workload, low efficiency, and manual errors in the labeling process;
  • high-precision labeling is required, such as the labeling in the image field, it is necessary to achieve pixel-level Segmentation, pure manual labeling must achieve pixel-level segmentation, which is very difficult and the labeling accuracy is difficult to guarantee.
  • the training of deep learning models based on purely manually labeled training data may result in low training efficiency and the resulting model's accuracy because of the low accuracy of the training data, resulting in the model's classification or recognition ability being less accurate than expected.
  • the embodiments of the present disclosure are expected to provide a deep model training method and device, electronic equipment, and storage medium.
  • a first aspect of an embodiment of the present disclosure provides a deep learning model training method, including:
  • n+1th label information output by the model to be trained, and the model to be trained has been n rounds of training; n is an integer greater than or equal to 1;
  • the n+1th training sample is used to perform the n+1th round of training on the model to be trained.
  • the generating the n+1th training sample based on the training data and the n+1th labeling information includes:
  • n+1th training sample based on the training data and the n+1th labeling information and the nth training sample, wherein the nth training sample includes: the training data and the first labeling information 1
  • the training samples, and the labeled information obtained from the previous n-1 rounds of training and the training samples constitute the second training sample to the n-1th training sample to be trained model respectively.
  • the method includes:
  • N is the maximum number of training rounds of the model to be trained
  • the obtaining the n+1th label information output by the model to be trained includes:
  • n is less than N, obtain the n+1th labeling information output by the model to be trained.
  • the method includes:
  • the first labeling information is generated.
  • the acquisition of the training data and the initial annotation information of the training data includes:
  • the generating the first labeling information based on the initial labeling information further includes:
  • the drawing outlines consistent with the shape of the segmentation target in the circumscribed frame based on the circumscribed frame includes:
  • an inscribed ellipse of the circumscribed frame consistent with the cell shape is drawn in the circumscribed frame.
  • a second aspect of an embodiment of the present disclosure provides a deep learning model training device, including:
  • the labeling module is configured to obtain the n+1th labeling information output by the model to be trained.
  • the model to be trained has been trained for n rounds; n is an integer greater than or equal to 1;
  • a first generation module configured to generate an n+1th training sample based on the training data and the n+1th labeling information
  • the training module is configured to perform the n+1th round of training the model to be trained on the model to be trained with the n+1th training sample.
  • the first generation module is configured to generate an n+1th training sample based on the training data and the n+1th labeling information, and the first training sample; or, based on the training data and all Generating the n+1th training sample by the n+1th labeling information and the nth training sample, wherein the nth training sample includes: the first training sample composed of the training data and the first labeling information, and the first n The labeled information obtained from the -1 round of training and the training samples constitute the second to n-1th training samples, respectively.
  • the device includes:
  • a determination module configured to determine whether n is less than N, where N is the maximum number of training rounds of the model to be trained;
  • the labeling module is configured to obtain the n+1th labeling information output by the model to be trained if n is less than N.
  • the device includes:
  • An acquisition module configured to acquire the training data and the initial annotation information of the training data
  • the second generating module is configured to generate the first labeling information based on the initial labeling information.
  • the acquisition module is configured to acquire a training image including multiple segmentation targets and an external frame of the segmentation targets;
  • the second generating module is configured to draw a label outline in the circumscribed frame consistent with the shape of the segmentation target based on the circumscribed frame.
  • the first generation module is configured to generate a segmentation boundary of two segmentation targets with overlapping portions based on the circumscribed frame.
  • the second generation module is configured to draw an inscribed ellipse of the circumscribed frame that is consistent with the cell shape in the circumscribed frame based on the circumscribed frame.
  • a third aspect of an embodiment of the present disclosure provides a computer storage medium that stores computer-executable instructions; the computer-executable instructions; after the computer-executable instructions are executed, any of the foregoing technical solutions can be implemented Provided deep learning model training methods.
  • a fifth aspect of an embodiment of the present disclosure provides an electronic device, including:
  • a processor connected to the memory, is configured to implement the deep learning model training method provided by any one of the foregoing technical solutions by executing computer-executable instructions stored on the memory.
  • a fifth aspect of an embodiment of the present disclosure provides a computer program product, the program product including computer-executable instructions; after the computer-executable instructions are executed, the deep learning model training method provided by any one of the foregoing technical solutions can be implemented.
  • the technical solution provided by the embodiment of the present disclosure uses the deep learning model to mark the training data after the previous round of training is completed to obtain labeling information.
  • the labeling information is used as a training sample for the next round of training, and very few initial labels can be used (for example, the initial manual annotation or equipment annotation) training data is used for model training, and then the labeled data output by the self-identification of the model to be trained that gradually converges is used as the next round of training samples, because the model to be trained in the previous training process
  • the model parameters will be generated based on most of the correctly labeled data, and a small amount of incorrectly labeled or low-precision data will have little effect on the model parameters of the trained model, so iterative multiple times, the label information of the model to be trained will become more and more accurate.
  • the training results are getting better and better. Because the model uses its own labeling information to build training samples, it reduces the amount of data for initial labeling such as manual manual labeling, reduces the efficiency and artificial errors caused by initial labeling such as manual manual labeling, and has fast model training speed and training effect. Good characteristics, and the deep learning model trained in this way has the characteristics of high classification or recognition accuracy.
  • FIG. 1 is a schematic flowchart of a first deep learning model training method provided by an embodiment of the present disclosure
  • FIG. 2 is a schematic flowchart of a second deep learning model training method provided by an embodiment of the present disclosure
  • FIG. 3 is a schematic flowchart of a third deep learning model training method provided by an embodiment of the present disclosure
  • FIG. 4 is a schematic structural diagram of a deep learning model training device provided by an embodiment of the present disclosure.
  • FIG. 5 is a schematic diagram of a change of a training set provided by an embodiment of the present disclosure.
  • FIG. 6 is a schematic structural diagram of an electronic device according to an embodiment of the present disclosure.
  • this embodiment provides a deep learning model training method, including:
  • Step S110 Obtain the n+1th labeling information output by the model to be trained, and the model to be trained has been n rounds of training;
  • Step S120 Generate an n+1th training sample based on the training data and the n+1th annotation information
  • Step S130 Perform the n+1th round of training on the model to be trained with the n+1th training sample.
  • the deep learning model training method provided in this embodiment can be used in various electronic devices, for example, in various large data model training servers.
  • the model structure of the model to be trained is obtained.
  • the network structure of the neural network needs to be determined first.
  • the network structure may include: the number of layers of the network, the number of nodes included in each layer, and the connection relationship between the nodes between the layers, And the initial network parameters.
  • the network parameters include: node weights and/or thresholds.
  • the first training sample may include: training data and first labeled data of the training data; taking image segmentation as an example, the training data is an image; and the first labeled data may be an image segmentation target A mask image with a background; in the embodiment of the present disclosure, all the first annotation information and the second annotation information may include but are not limited to the annotation information of the image.
  • the image may include medical images and the like.
  • the medical image may be a planar (2D) medical image or a stereoscopic (3D) medical image composed of an image sequence formed by a plurality of 2D images.
  • Each of the first labeling information and the second labeling information may be a label for an organ and/or tissue in a medical image, or may be a label for different cell structures in a cell, such as a label for a cell nucleus.
  • the images are not limited to medical images, but can also be applied to images of traffic road conditions in the field of traffic roads.
  • the model parameters of the deep learning model (for example, the network parameters of the neural network) are changed; the model to be trained is used to process the image and output annotation information.
  • the annotation information and the initial The first label information is compared, and the current loss value of the deep learning model is calculated by the result of the comparison; if the current loss value is less than the loss threshold, the training round can be stopped.
  • step S110 in this embodiment the training data will first be processed using the model to be trained that has completed n rounds of training. At this time, the model to be trained will obtain an output, which is the n+1th labeled data, The n+1th labeled data corresponds to the training data to form a training sample.
  • the training data and the n+1th labeling information may be directly used as the n+1th training sample for the n+1th training sample of the model to be trained.
  • the training data and the n+1th labeled data, and the first training sample may form the n+1th round of training samples of the model to be trained.
  • the first training sample is a training sample for the first round of training of the training model
  • the Mth training sample is a training sample for the Mth round of training of the training module
  • M is a positive integer.
  • the first training sample here may be: the training data and the first labeling information of the training data are initially obtained, and the first labeling information here may be manually labeled information.
  • the training data and the n+1th label information, and the union of this training sample and the nth training sample used in the nth round of training constitutes the n+1th training sample.
  • the above three methods for generating the n+1th training sample are all methods for the device to automatically generate samples. Therefore, there is no need for the user to manually mark and other equipment to mark the training samples for the n+1th round of training, reducing manual manual labeling. Waiting for the time consumed by the initial annotation of the sample increases the training rate of the deep learning model and reduces the inaccuracies in the classification or recognition results of the model after training due to inaccurate or inaccurate manual annotation. The accuracy of the classification or recognition results after the deep learning model is trained.
  • Completing a round of training in this embodiment includes: the model to be trained completes at least one learning for each training sample in the training set.
  • step S130 the n+1th training sample is used to perform the n+1th round of training on the training model.
  • the first training sample may be the S images and the manual labeling results of the S images. If one of the S images is not accurate enough to label the image, but During the first training process of the training model, since the accuracy of the annotation structure of the remaining S-1 images reaches the expected threshold, the S-1 images and their corresponding annotation data are larger in the image of the model parameters of the training model.
  • the deep learning model includes but is not limited to a neural network; the model parameters include but are not limited to: weights and/or thresholds of network nodes in the neural network.
  • the neural network may be various types of neural networks, for example, U-net or V-net.
  • the neural network may include an encoding part that performs feature extraction on the training data and a decoding part that acquires semantic information based on the extracted features.
  • the encoding part can perform feature extraction on the area where the segmentation target is located in the image to obtain a mask image that distinguishes the segmentation target from the background.
  • the decoder can obtain some semantic information based on the mask image, for example, the target's Omics characteristics, etc.
  • the omics feature may include: morphological features such as area, volume, and shape of the target, and/or gray value features formed based on the gray value.
  • the characteristics of the gray value may include: statistical characteristics of the histogram and the like.
  • the model to be trained after the first round of training recognizes S images, it will initially label the image parameter of the image to be trained with an insufficient accuracy compared to the other S- 1 piece has a low loudness.
  • the model to be trained will use network parameters learned from other S-1 images for labeling, and the labeling accuracy of the image with insufficient initial labeling accuracy at this time is aligned with the labeling accuracy of other S-1 images, so this
  • the second annotation information corresponding to an image is more accurate than the original first annotation information.
  • the second training set composed includes: training data composed of S images and original first labeling information, and training data composed of S images and second labeling information that the model to be trained self-labels.
  • the model to be trained will be learned based on most correct or high-precision labeling information during the training process to gradually suppress the negative effects of training samples with insufficient or incorrect initial labeling accuracy, and thus adopt this
  • the automatic iteration of the deep learning model in this way can not only greatly reduce the manual annotation of training samples, but also gradually improve the training accuracy through its own iteration characteristics, so that the accuracy of the model to be trained after training reaches the expected effect.
  • the training data takes an image as an example.
  • the training data may also be a voice segment other than the image, text information other than the image, etc.
  • the training data has many forms It is not limited to any of the above.
  • the method includes:
  • Step S100 Determine whether n is less than N, where N is the maximum number of training rounds of the model to be trained;
  • the step S110 may include:
  • the model to be trained obtains the n+1th labeling information output by the model to be trained.
  • the n+1th training set before constructing the n+1th training set, it is first determined whether the current number of training rounds of the model to be trained reaches the predetermined maximum number of training rounds N, and the n+1st labeling information is generated if it is not reached to achieve Construct the n+1th training set, otherwise, it is determined that the model training is completed to stop the training of the deep learning model.
  • the value of N may be 4, 5, 6, 7 or 8 empirical values or statistical values.
  • the value of N may range from 3 to 10, and the value of N may be a user input value received by the training device from the human-computer interaction interface.
  • determining whether to stop the training of the model to be trained may further include:
  • test set Use the test set to test the model to be trained. If the test result indicates that the accuracy of the labeling result of the test data in the test set of the model to be trained reaches a specific value, then stop the training of the model to be trained, otherwise enter Said step S110 to enter the next round of training.
  • the test set may be an accurately labeled data set, so it can be used to measure the training result of each round of a model to be trained to determine whether to stop the training of the model to be trained.
  • the method includes:
  • Step S210 Obtain the training data and the initial annotation information of the training data
  • Step S220 Generate the first labeling information based on the initial labeling information.
  • the initial labeling information may be original labeling information of the training data, and the original labeling information may be information manually labeled manually, or may be information labeled by other devices. For example, information marked by other devices with certain marking capabilities.
  • the first labeling information is generated based on the initial labeling information.
  • the first label information here may directly include the initial label information and/or refined first label information generated according to the initial standard information.
  • the initial labeling information may be labeling information that roughly labels the location of the cell imaging
  • the first identification information may be an accurate indication of the location of the cell Labeling information.
  • the accuracy of labeling the segmentation object with the first labeling information may be higher than the accuracy of the initial labeling information.
  • the initial labeling information may be a circumscribed frame of cells drawn manually by a doctor.
  • the first labeling information may be: an inscribed ellipse generated by the training device based on a manually labeled outer frame. Compared with the circumscribed frame, the calculation of the inscribed ellipse reduces the number of pixels that do not belong to the cell imaging in the cell imaging, so the accuracy of the first labeling information is higher than that of the initial labeling information.
  • the step S210 may include: obtaining a training image including a plurality of segmentation targets and an external frame of the segmentation targets;
  • the step S220 may include: based on the circumscribed frame, drawing a marked outline in the circumscribed frame consistent with the shape of the segmentation target.
  • the annotated contour that is consistent with the segmentation target shape may be the aforementioned ellipse, or may be a circle, or, a triangle or other contralateral shape is equal to the segmentation target shape, and is not limited to an ellipse.
  • the marked outline is inscribed in the outer frame.
  • the external frame may be a rectangular frame.
  • the step S220 further includes:
  • the first labeling information further includes: a segmentation boundary between the two overlapping segmentation targets.
  • cell imaging A is superimposed on cell imaging B, then after cell imaging A is drawn out of the cell boundary and after cell B imaging is drawn out of the cell boundary, the two cell boundaries intersect to form part of the two Intersection between cell imaging.
  • the portion of the cell boundary of the cell imaging B located inside the cell imaging A may be erased, and the part of the cell imaging A located in the cell imaging B may be As the division boundary.
  • the step S220 may include: drawing the division boundary on the overlapping part of the two using the positional relationship of the two division targets.
  • the segmentation boundary when drawing the segmentation boundary, it can be achieved by modifying the boundary of one of the two segmentation targets with overlapping boundaries.
  • the pixel expansion can be used to thicken the boundary.
  • the cell boundary of the cell imaging A is expanded by a predetermined number of pixels in the direction of the overlapping portion toward the cell imaging B, for example, 1 or more pixels, and the cell of the overlapping portion is thickened to the boundary of the imaging A, thereby making the bolding
  • the boundary is recognized as a dividing boundary.
  • the drawing an outline corresponding to the shape of the segmentation target in the circumscribed frame based on the circumscribed frame includes: drawing the cell shape in the circumscribed frame based on the circumscribed frame The ellipse inside the outer frame is consistent.
  • the segmentation target is cell imaging
  • the marked outline includes an inscribed ellipse of a circumscribed frame of the cell shape.
  • the first labeling information includes at least one of the following:
  • the cell boundary of the cell imaging (corresponding to the inscribed ellipse);
  • the segmentation target is not a cell but other targets, for example, the segmentation target is a face in a collective phase, the outer frame of the face may still be a rectangular frame, but at this time the boundary of the face may be marked It is the border of an oval-shaped face, the border of a round face, etc. In this case, the shape is not limited to the inscribed ellipse.
  • the model to be trained uses its previous training results to output the labeling information of the training data during its own training process to construct the training set for the next round. Iterate multiple times to complete model training without manually labeling a large number of training samples. It has a fast training rate and can improve training accuracy through repeated iterations.
  • this embodiment provides a deep learning model training device, including:
  • the labeling module 110 is configured to obtain the n+1th labeling information output by the model to be trained.
  • the model to be trained has been trained for n rounds; n is an integer greater than or equal to 1;
  • the first generation module 120 is configured to generate an n+1th training sample based on the training data and the n+1th annotation information
  • the training module 130 is configured to perform the n+1th round of training on the model to be trained on the n+1th training sample.
  • the labeling module 110, the first generating module 120, and the training module 130 may be program modules, and the program modules, after being executed by the processor, can realize the generation of the n+1th labeling information and the nth The composition of the +1 training set and the training of the model to be trained.
  • the labeling module 110, the first generation module 120, and the training module 130 may be soft-hard combination models; the soft-hard combination modules may be various programmable arrays, for example, field programmable arrays Or complex programmable array.
  • the labeling module 110, the first generation module 120, and the training module 130 may be pure hardware modules, and the pure hardware modules may be application specific integrated circuits.
  • the first generation module 120 is configured to generate an n+1th training sample based on the training data and the n+1th labeling information, and the first training sample; or, based on the training Data and the n+1th labeling information and the nth training sample to generate an n+1th training sample, wherein the nth training sample includes: a first training sample composed of the training data and the first labeling information, The labeled information obtained from the previous n-1 rounds of training and the training samples constitute the second to n-1th training samples, respectively.
  • the device includes:
  • a determination module configured to determine whether n is less than N, where N is the maximum number of training rounds of the model to be trained;
  • the labeling module 110 is configured to, if n is less than N, the model to be trained acquire the n+1th labeling information output by the model to be trained.
  • the device includes:
  • An acquisition module configured to acquire the training data and the initial annotation information of the training data
  • the second generating module is configured to generate the first labeling information based on the initial labeling information.
  • the acquisition module is configured to acquire a training image including multiple segmentation targets and an external frame of the segmentation targets;
  • the first generating module 120 is configured to generate a segmentation boundary of two segmentation targets with overlapping portions based on the circumscribed frame.
  • the second generation module is configured to draw an inscribed ellipse of the circumscribed frame that is consistent with the cell shape in the circumscribed frame based on the circumscribed frame.
  • This example provides a self-learning weakly supervised learning method for deep learning models.
  • the supervision signal here is the training sample in the training set
  • the segmentation model is predicted on this graph, and the obtained predicted graph and the initial annotation graph are combined as a new supervision signal, and the segmentation model is repeatedly trained.
  • the original image is annotated to obtain a mask image to construct the first training set, and the first training set is used for the first round of training.
  • the deep learning model is used for image recognition to obtain the second annotation information.
  • the second training set is constructed based on the second annotation information.
  • the third labeling information is output, and the third training set is obtained based on the third labeling information. Stop training after repeated iteration training in this way.
  • the deep learning model training method does not perform any calculation on the output segmentation probability map, and directly takes it as a union with the annotation map, and then continues to train the model. This process is simple to implement.
  • an electronic device including:
  • Memory used to store information
  • a processor connected to the memory, is configured to execute the deep learning model training method provided by the foregoing one or more technical solutions by executing computer-executable instructions stored on the memory, for example, as shown in FIGS. 1 to 3 One or more of the methods shown.
  • the memory may be various types of memory, such as random access memory, read-only memory, flash memory, etc.
  • the memory can be used for information storage, for example, storing computer-executable instructions.
  • the computer executable instructions may be various program instructions, for example, target program instructions and/or source program instructions.
  • the processor may be various types of processors, for example, a central processor, a microprocessor, a digital signal processor, a programmable array, a digital signal processor, an application specific integrated circuit, or an image processor.
  • the processor may be connected to the memory through a bus.
  • the bus may be an integrated circuit bus or the like.
  • the terminal device may further include: a communication interface, and the communication interface may include: a network interface, for example, a local area network interface, a transceiver antenna, and the like.
  • the communication interface is also connected to the processor and can be used for information transmission and reception.
  • the electronic device further includes a camera, which can collect various images, such as medical images.
  • the terminal device further includes a human-machine interaction interface.
  • the human-machine interaction interface may include various input and output devices, such as a keyboard, a touch screen, and so on.
  • An embodiment of the present disclosure provides a computer storage medium that stores computer executable code; after the computer executable code is executed, the deep learning model training method provided by one or more of the foregoing technical solutions can be implemented For example, one or more of the methods shown in FIGS. 1-3.
  • the storage medium includes: mobile storage devices, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks or optical disks and other media that can store program codes.
  • the storage medium may be a non-transitory storage medium.
  • An embodiment of the present disclosure provides a computer program product, the program product including computer-executable instructions; after the computer-executable instructions are executed, the deep learning model training method provided by any of the foregoing implementations can be implemented, for example, as shown in FIGS. One or more of the methods shown in FIG. 3.
  • the disclosed device and method may be implemented in other ways.
  • the device embodiments described above are only schematic.
  • the division of the units is only a division of logical functions.
  • the displayed or discussed components are coupled to each other, or directly coupled, or the communication connection may be through some interfaces, and the indirect coupling or communication connection of the device or unit may be electrical, mechanical, or other forms of.
  • the above-mentioned units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units; Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
  • the functional units in the embodiments of the present disclosure may all be integrated into one processing module, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above integration
  • the unit can be implemented in the form of hardware, or in the form of hardware plus software functional units.
  • a computer program product includes computer-executable instructions; after the computer-executable instructions are executed, the deep model training method in the foregoing embodiment can be implemented.
  • the foregoing program may be stored in a computer-readable storage medium, and when the program is executed, Including the steps of the above method embodiments; and the foregoing storage media include: mobile storage devices, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks or optical disks, etc.
  • ROM read-only memory
  • RAM random access memory
  • magnetic disks or optical disks etc.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • General Health & Medical Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Computational Linguistics (AREA)
  • General Engineering & Computer Science (AREA)
  • Biomedical Technology (AREA)
  • Evolutionary Computation (AREA)
  • Molecular Biology (AREA)
  • Computing Systems (AREA)
  • Data Mining & Analysis (AREA)
  • Biophysics (AREA)
  • Mathematical Physics (AREA)
  • Software Systems (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Image Analysis (AREA)

Abstract

一种深度模型训练方法及装置、电子设备及存储介质,所述深度模型训练方法包括:获取待训练模型输出的第n+1标注信息,所述待训练模型已经过n轮训练(S110);基于所述训练数据及所述第n+1轮标注信息生成第n+1训练样本(S120);将所述第n+1训练样本对所述待训练模型进行第n+1轮训练(S130)。

Description

深度模型训练方法及装置、电子设备及存储介质
相关申请的交叉引用
本公开基于申请号为201811646430.5、申请日为2018年12月29日的中国专利申请提出,并要求该中国专利申请的优先权,该中国专利申请的全部内容在此引入本公开作为参考。
技术领域
本公开涉及信息技术领域但不限于信息技术领域,尤其涉及一种深度模型训练方法及装置、电子设备及存储介质。
背景技术
深度学习模型可以通过训练集的训练之后,具有一定的分类或识别能力。所述训练集通常包括:训练数据及训练数据的标注数据。但是一般情况下,标注数据都需要人工进行手动标注。一方面纯手动标注所有的训练数据,工作量大、效率低,且标注过程中存在人工错误;另一方面,若需要实现高精度的标注,例如以图像领域的标注为例,需要实现像素级分割,纯人工标注要达到像素级分割,难度非常大且标注精度也难以保证。
故基于纯人工标注的训练数据进行深度学习模型的训练,会存在训练效率低、训练得到的模型因为训练数据自身精度低导致模型的分类或识别能力精度达不到预期。
发明内容
有鉴于此,本公开实施例期望提供一种深度模型训练方法及装置、电子设备及存储介质。
本公开的技术方案是这样实现的:
本公开实施例第一方面提供一种深度学习模型训练方法,包括:
获取待训练模型输出的第n+1标注信息,所述待训练模型已经过n轮训练;n为大于或等于1的整数;
基于所述训练数据及所述第n+1标注信息生成第n+1训练样本;
将所述第n+1训练样本对所述待训练模型进行第n+1轮训练。
基于上述方案,所述基于所述训练数据及所述第n+1标注信息生成第n+1训练样本,包括:
基于所述训练数据及所述第n+1标注信息、及第1训练样本生成第n+1训练样本;
或者,
基于所述训练数据及所述第n+1标注信息、及第n训练样本生成第n+1训练样本,其中,所述第n训练样本包括:所述训练数据及第一标注信息构成的第1训练样本、及前n-1轮训练得到的标注信息与所述训练样本分别构成的第2训练样本至第n-1训练样本待训练模型。
基于上述方案,所所述方法包括:
确定n是否小于N,其中,N为所述待训练模型的最大训练轮数;
所述获取待训练模型输出的第n+1标注信息,包括:
若n小于N,获取所述待训练模型输出的第n+1标注信息。
基于上述方案,所述方法包括:
获取所述训练数据及所述训练数据的初始标注信息;
基于所述初始标注信息,生成所述第一标注信息。
基于上述方案,所所述获取所述训练数据及所述训练数据的初始标注信息,包括:
获取包含有多个分割目标的训练图像及所述分割目标的外接框;
所述基于所述初始标注信息,生成所述第一标注信息,包括:
基于所述外接框,在所述外接框内绘制与所述分割目标形状一致的标注轮廓。
基于上述方案,所述基于所述初始标注信息,生成所述第一标注信息,还 包括:
基于所述外接框,生成具有重叠部分的两个所述分割目标的分割边界。
基于上述方案,所述基于所述外接框,在所述外接框内绘制与所述分割目标形状一致的标注轮廓,包括:
基于所述外接框,在所述外接框内绘制与细胞形状一致的所述外接框的内接椭圆。
本公开实施例第二方面提供一种深度学习模型训练装置,包括:
标注模块,配置为获取待训练模型输出的第n+1标注信息,所述待训练模型已经过n轮训练;n为大于或等于1的整数;
第一生成模块,配置为基于所述训练数据及所述第n+1标注信息生成第n+1训练样本;
训练模块,配置为将所述第n+1训练样本对所述待训练模型进行第n+1轮训练待训练模型。
基于上述方案,所述第一生成模块,配置为基于所述训练数据及所述第n+1标注信息、及第1训练样本生成第n+1训练样本;或者,基于所述训练数据及所述第n+1标注信息、及第n训练样本生成第n+1训练样本,其中,所述第n训练样本包括:所述训练数据及第一标注信息构成的第1训练样本、及前n-1轮训练得到的标注信息与所述训练样本分别构成的第2训练样本至第n-1训练样本。
基于上述方案,所述装置包括:
确定模块,配置为确定n是否小于N,其中,N为所述待训练模型的最大训练轮数;
所述标注模块,配置为若n小于N,获取所述待训练模型输出的第n+1标注信息。
基于上述方案,所述装置包括:
获取模块,配置为获取所述训练数据及所述训练数据的初始标注信息;
第二生成模块,配置为基于所述初始标注信息,生成所述第一标注信息。
基于上述方案,所述获取模块,配置为获取包含有多个分割目标的训练图像及所述分割目标的外接框;
所述第二生成模块,配置为基于所述外接框,在所述外接框内绘制与所述分割目标形状一致的标注轮廓。
基于上述方案,所述第一生成模块,配置为基于所述外接框,生成具有重叠部分的两个所述分割目标的分割边界。
基于上述方案,所述第二生成模块,配置为基于所述外接框,在所述外接框内绘制与细胞形状一致的所述外接框的内接椭圆。
本公开实施例第三方面提供一种计算机存储介质,所述计算机存储介质存储有计算机可执行指令;所述计算机可执行指令;所述计算机可执行指令被执行后,能够实现前述任意一个技术方案提供的深度学习模型训练方法。
本公开实施例第五方面提供一种电子设备,包括:
存储器;
处理器,与所述存储器连接,用于通过执行存储在所述存储器上的计算机可执行指令实现前述任意一个技术方案提供的深度学习模型训练方法。
本公开实施例第五方面提供一种计算机程序产品,所述程序产品包括计算机可执行指令;所述计算机可执行指令被执行后,能够实现前述任意一个技术方案提供的深度学习模型训练方法。
本公开实施例提供的技术方案,会利用深度学习模型前一轮训练完成之后对训练数据进行标注获得标注信息,该标注信息用作下一轮训练的训练样本,可以利用非常少的初始标注(例如,初始的人工标注或者设备标注)的训练数据进行模型训练,然后面利用逐步收敛的待训练模型自身识别输出的标注数据作为下一轮训练样本,由于待训练模型在前一轮训练过程中模型参数会依据大部分标注正确的数据生成,而少量标注不正确或者标注精度低的数据对待训练模型的模型参数影响小,如此反复迭代多次,待训练模型的标注信息会越来越精确,训练结果也越来越好。由于模型利用自身的标注信息构建训练样本,如此,减少了人工手动标注等初始标注的数据量,减少了人工手动标注等初始标 注所导致的效率低及人工错误,具有模型训练速度快及训练效果好的特点,且采用这种方式训练的深度学习模型,具有分类或识别精确度高的特点。
附图说明
图1为本公开实施例提供的第一种深度学习模型训练方法的流程示意图;
图2为本公开实施例提供的第二种深度学习模型训练方法的流程示意图;
图3为本公开实施例提供的第三种深度学习模型训练方法的流程示意图;
图4为本公开实施例提供的一种深度学习模型训练装置的结构示意图;
图5为本公开实施例提供的一种训练集的变化示意图;
图6为本公开实施例提供的一种电子设备的结构示意图。
具体实施方式
以下结合说明书附图及具体实施例对本公开的技术方案做进一步的详细阐述。
如图1所示,本实施例提供一种深度学习模型训练方法,包括:
步骤S110:获取待训练模型输出的第n+1标注信息,所述待训练模型已经过n轮训练;
步骤S120:基于所述训练数据及所述第n+1标注信息生成第n+1训练样本;
步骤S130:将所述第n+1训练样本对所述待训练模型进行第n+1轮训练。
本实施例提供的深度学习模型训练方法可以用于各种电子设备中,例如,各种大数据模型训练的服务器中。
在进行第1轮训练时,获取待训练模型的模型结构。以待训练模型为神经网络为例进行说明,首先需要确定神经网络的网络结构,该网络结构可包括:网络的层数、每层包括的节点数,层与层之间的节点的连接关系,以及初始的网络参数。该网络参数包括:节点的权重和/或阈值。
获取第1训练样本,所述第一训练样本可包括:训练数据和训练数据的第一标注数据;以图像分割为例,所述训练数据为图像;所述第一标注数据可为 图像分割目标和背景的掩码图像;在本公开实施例中所有的第一标注信息和第二标注信息,可包括但不限于对图像的标注信息。该图像可包括医疗图像等。该医疗图像可为平面(2D)医疗图像或者由多个2D图像形成的图像序列构成的立体(3D)医疗图像。各所述第一标注信息和所述第二标注信息,可为对医疗图像中器官和/会组织的标注,也可以是对细胞内不同细胞结构的标注,如,细胞核的标注。在一些实施例中,所述图像不限于是医疗图像,还可应用于交通道路领域的交通道路状况的图像。
利用第1训练样本训练对待训练模型进行第一轮训练。神经网络等深度学习模型被训练之后,深度学习模型的模型参数(例如,神经网络的网络参数)发生改变;利用改变了模型参数的待训练模型对图像进行处理输出标注信息,该标注信息与初始的第一标注信息进行比对,通过比对的结果计算深度学习模型的当前的损失值;若当前的损失值小于损失阈值可以停止该轮训练。
在本实施例中的步骤S110中,首先会利用已经完成n轮训练的待训练模型对训练数据进行处理,此时待训练模型会获得输出,该输出即为所述第n+1标注数据,该第n+1标注数据与训练数据对应起来,就形成了训练样本。
在一些实施例中,可以将训练数据和第n+1标注信息直接作为第n+1训练样本,用于待训练模型的第n+1轮训练样本。
在还有一些实施例中,可以将训练数据和第n+1标注数据,及第1训练样本共同组成待训练模型的第n+1轮训练样本。
所述第1训练样本为对待训练模型进行第1轮训练的训练样本;第M训练样本为对待训练模块进行第M轮训练的训练样本,M为正整数。
此处的第1训练样本可为:初始获得训练数据及训练数据的第一标注信息,此处的第一标注信息可为人工手动标注的信息。
在还有一些实施例中,训练数据和第n+1标注信息,并将这个训练样本和第n轮训练时采用的第n训练样的并集构成第n+1训练样本。
总之,上述三种生成第n+1训练样本的方式都是设备自动生成样本的方式,如此,无需用户手动标注等其他设备来标注获得第n+1轮训练的训练样本,减少了人工手动标等初始标注注样本所消耗的时间,提升了深度学习模型的训练 速率,且减少深度学习模型因为手动标注的不准确或不精确导致的模型训练后的分类或识别结果的不够精确的现象,提升了深度学习模型训练后的分类或识别结果的精确度。
在本实施例中完成一轮训练包括:待训练模型对训练集中的每一个训练样本都完成了至少一次学习。
在步骤S130中利用第n+1训练样本对待训练模型进行第n+1轮训练。
在本实施例中,若初始标注中有少量错误,由于模型训练过程中会关注训练样本的共同特点,则这些错误的影响对模型训练就越来越小,从而模型的精确度也越来越高。
例如,以所述训练数据为S张图像为例,则第1训练样本可为S张图像及这S张图像的人工标注结果,若S张图像中有一张图像标注图像精确度不够,但是待训练模型在第一轮训练过程中,由于剩余S-1张图像的标注结构精确度达到预期阈值,则这S-1张图像及其对应的标注数据对待训练模型的模型参数影像更大。在本实施例中,所述深度学习模型包括但不限于神经网络;所述模型参数包括但不限于:神经网络中网络节点的权值和/或阈值。所述神经网络可为各种类型的神经网络,例如,U-net或V-net。所述神经网络可包括:对训练数据进行特征提取的编码部分和基于提取的特征获取语义信息的解码部分。
例如,编码部分可以对图像中分割目标所在区域等进行特征提取,得到区分分割目标和背景的掩码图像,解码器基于掩码图像可以得到一些语义信息,例如,通过像素统计等方式获得目标的组学特征等。
该组学特征可包括:目标的面积、体积、形状等形态特征,和/或,基于灰度值形成的灰度值特征等。
所述灰度值特征可包括:直方图的统计特征等。
总之,在本实施例中,经过第一轮训练后的待训练模型在识别S张图像时,会初始标注精度不够的那一张图像待训练模型的模型参数的影响度相比于另外S-1张的响度小的。待训练模型将利用从其他S-1张图像上学习获得网络参数来进行标注,而此时初始标注精度不够的图像的标注精度是向其他S-1张图像的 标注精度靠齐的,故这一张图像所对应的第2标注信息是会比原始的第1标注信息的精度提升的。如此,构成的第2训练集包括:S张图像和原始的第一标注信息构成的训练数据、及S张图像和待训练模型自行标注的第二标注信息构成的训练数据。故在本实施例中,可以利用待训练模型在训练过程中会基于大多数正确或高精度的标注信息进行学习,逐步抑制初始标注精度不够或不正确的训练样本的负面影响,从而采用这种方式进行深度学习模型的自动迭代,不仅能够实现训练样本的人工标注大大的减少,而且还会通过自身迭代的特性逐步提升训练精度,使得训练后的待训练模型的精确度达到预期效果。
在上述举例中所述训练数据以图像为例,在一些实施例中,所述训练数据还可以图像以外的语音片段、所述图像以外的文本信息等;总之,所述训练数据的形式有多种,不局限于上述任意一种。
在一些实施例中,如图2所示,所述方法包括:
步骤S100:确定n是否小于N,其中,N为所述待训练模型的最大训练轮数;
所述步骤S110可包括:
若n小于N,待训练模型获取待训练模型输出的第n+1标注信息。
在本实施例中在构建第n+1训练集之前,首先会确定目前待训练模型的训练轮数是否达到预定的最大训练轮数N,若未大达到才生成第n+1标注信息,以构建第n+1训练集,否则,则确定模型训练完成停止所述深度学习模型的训练。
在一些实施例中,所述N的取值可为4、5、6、7或8等经验值或者统计值。
在一些实施例中,所述N的取值范围可为3到10之间,所述N的取值可以是训练设备从人机交互接口接收的用户输入值。
在还有一些实施例中,确定是否停止待训练模型的训练还可包括:
利用测试集进行所述待训练模型的测试,若测试结果表明所述待训练模型的对测试集中测试数据的标注结果的精确度达到特定值,则停止所述待训练模 型的训练,否则进入到所述步骤S110以进入下一轮训练。此时,所述测试集可为精确标注的数据集,故可以用于衡量一个待训练模型的每一轮的训练结果,以判定是否停止待训练模型的训练。
在一些实施例中,如图3所示,所述方法包括:
步骤S210:获取所述训练数据及所述训练数据的初始标注信息;
步骤S220:基于所述初始标注信息,生成所述第一标注信息。
在本实施例中,所述初始标注信息可为所述训练数据的原始标注信息,该原始标注信息可为人工手动标注的信息,也可以是其他设备标注的信息。例如,具有一定标注能力的其他设备标注的信息。
本实施例中,获取到训练数据及初始标注信息之后,会基于初始标注信息生成第一标注信息。此处的第一标注信息可直接包括所述初始标注信息和/或根据所述初始标准信息生成的精细化的第一标注信息。
例如,若训练数据为图像,图像包含有细胞成像,所述初始标注信息可为大致标注所述细胞成像所在位置的标注信息,而所述第一标识信息可为精确指示所述细胞所在位置的标注信息,总之,在本实施例中,所述第一标注信息对分割对象的标注精确度可高于所述初始标注信息的精确度。
如此,即便由人工进行所述初始标注信息的标注,也降低了人工标注的难度,简化了人工标注。
例如,以细胞成像为例,细胞由于其椭圆球体的形态,一般在二维平面图像内细胞的外轮廓都呈现为椭圆形。所述初始标注信息可为医生手动绘制的细胞的外接框。所述第一标注信息可为:训练设备基于手动标注的外接框生成的内接椭圆。在计算内接椭圆相对于外接框,减少细胞成像中不属于细胞成像的像素个数,故第一标注信息的精确度是高于所述初始标注信息的精确度的。
故进一步地,所述步骤S210可包括:获取包含有多个分割目标的训练图像及所述分割目标的外接框;
所述步骤S220可包括:基于所述外接框,在所述外接框内绘制与所述分割目标形状一致的标注轮廓。
在一些实施例中,所述与分割目标形状一致的标注轮廓可为前述椭圆形,还可以为圆形,或者,三角形或者其他对边形等于分割目标形状一致的形状,不局限于椭圆形。
在一些实施例中,所述标注轮廓为内接于所述外接框的。所述外接框可为矩形框。
在一些实施例中,所述步骤S220还包括:
基于所述外接框,生成具有重叠部分的两个所述分割目标的分割边界。
在一些图像中,两个分割目标之间会有重叠,在本实施例中所述第一标注信息还包括:两个重叠分割目标之间的分割边界。
例如,两个细胞成像,细胞成像A叠在细胞成像B上,则细胞成像A被绘制出细胞边界之后和细胞B成像被绘制出细胞边界之后,两个细胞边界交叉形成一部分框出了两个细胞成像之间的交集。在本实施例中,可以根据细胞成像A和细胞成像B之间的位置关系,擦除细胞成像B的细胞边界位于细胞成像A内部的部分,并以细胞成像A的位于细胞成像B中的部分作为所述分割边界。
总之,在本实施例中,所述步骤S220可包括:利用两个分割目标的位置关系,在两者的重叠部分绘制分割边界。
在一些实施例中,在绘制分割边界时,可以通过修正两个具有重叠边界的分割目标其中一个的边界来实现。为了突出边界,可以通过像素膨胀的方式,可以加粗边界。例如,通过细胞成像A的细胞边界在所述重叠部分向细胞成像B方向上扩展预定个像素,例如,1个或多个像素,加粗重叠部分的细胞成像A的边界,从而使得该加粗边界被识别为分割边界。
在一些实施例中,所述基于所述外接框,在所述外接框内绘制与所述分割目标形状一致的标注轮廓,包括:基于所述外接框,在所述外接框内绘制与细胞形状一致的所述外接框的内接椭圆。
在本本实施例中分割目标为细胞成像,所述标注轮廓包括所述细胞形状这一张的外接框的内接椭圆。
在本实施例中,所述第一标注信息包括以下至少之一:
所述细胞成像的细胞边界(对应于所述内接椭圆);
重叠细胞成像之间的分割边界。
若在一些实施例中,所述分割目标不是细胞而是其他目标,例如,分割目标为集体相中的人脸,人脸的外接框依然可以是矩形框,但是此时人脸的标注边界可能是鹅蛋形脸的边界,圆形脸的边界等,此时,所述形状不局限于所述内接椭圆。
当然以上仅是举例,总之在本实施例中,所述待训练模型在自身的训练过程中利用自身前一轮的训练结果输出训练数据的标注信息,以构建下一轮的训练集,通过反复迭代多次完成模型训练,无需手动标注大量的训练样本,具有训练速率快及通过反复迭代可以提升训练精确度。
如图5所示,本实施例提供一种深度学习模型训练装置,包括:
标注模块110,配置为获取待训练模型输出的第n+1标注信息,所述待训练模型已经过n轮训练;n为大于或等于1的整数;
第一生成模块120,配置为基于所述训练数据及所述第n+1标注信息生成第n+1训练样本;
训练模块130,配置为将所述第n+1训练样本对所述待训练模型进行第n+1轮训练。
在一些实施例中,所述标注模块110,第一生成模块120及训练模块130可为程序模块,所述程序模块被处理器执行后,能够实现前述第n+1标注信息的生成、第n+1训练集的构成及待训练模型的训练。
在还有一些实施例中,所述标注模块110,第一生成模块120及训练模块130可为软硬结合模型;所述软硬结合模块可为各种可编程阵列,例如,现场可编程阵列或复杂可编程阵列。
在另外一些实施例中,所述标注模块110,第一生成模块120及训练模块130可纯硬件模块,所述纯硬件模块可为专用集成电路。
在一些实施例中,所述第一生成模块120,配置为基于所述训练数据及所述第n+1标注信息、及第1训练样本生成第n+1训练样本;或者,基于所述训 练数据及所述第n+1标注信息、及第n训练样本生成第n+1训练样本,其中,所述第n训练样本包括:所述训练数据及第一标注信息构成的第1训练样本、及前n-1轮训练得到的标注信息与所述训练样本分别构成的第2训练样本至第n-1训练样本。
在一些实施例中,所述装置包括:
确定模块,配置为确定n是否小于N,其中,N为所述待训练模型的最大训练轮数;
所述标注模块110,配置为若n小于N,待训练模型获取所述待训练模型输出的第n+1标注信息。
在一些实施例中,所述装置包括:
获取模块,配置为获取所述训练数据及所述训练数据的初始标注信息;
第二生成模块,配置为基于所述初始标注信息,生成所述第一标注信息。
在一些实施例中,所述获取模块,配置为获取包含有多个分割目标的训练图像及所述分割目标的外接框;
所述基于所述初始标注信息,生成所述第一标注信息,包括:
基于所述外接框,在所述外接框内绘制与所述分割目标形状一致的标注轮廓。
在一些实施例中,所述第一生成模块120,配置为基于所述外接框,生成具有重叠部分的两个所述分割目标的分割边界。
在一些实施例中,所述第二生成模块,配置为基于所述外接框,在所述外接框内绘制与细胞形状一致的所述外接框的内接椭圆。
以下结合上述实施例提供一个具体示例:
示例1:
本示例提供一种深度学习模型的自学习式弱监督学学习方法。
以图5中每一个物体的包围矩形框作为输入,进行自我学习,能够输出该物体,以及其他没有标注的物体的像素分割结果。
以细胞分割为例子,一开始有图中部分细胞的包围矩形标注。观察发现细 胞大部分是椭圆,于是在矩形中画个最大内接椭圆,不同椭圆之间画好分割线,椭圆边缘也画上分割线;作为初始监督信号。此处的监督信号即为训练集中的训练样本;
训练一个分割模型。
此分割模型在此图上预测,得到的预测图和初始标注图作并集,作为新的监督信号,再重复训练该分割模型。
通过观测发现图中的分割结果变得越来越好。
如图5所示,对原始图像进行标注得到一个掩膜图像构建第一训练集,利用第一训练集进行第一轮训练,训练完之后,利用深度学习模型进行图像识别得到第2标注信息,基于第二标注信息构建第2训练集。利用第二训练集完成第二轮训练之后输出第3标注信息,基于第三标注信息得到第3训练集。如此反复迭代训练多轮之后停止训练。
在相关技术中,总是复杂的考虑第一次分割结果的概率图,做峰值、平缓区域等等的分析,然后做区域生长等,对于阅读者而言,复现工作量大,实现困难。本示例提供的深度学习模型训练方法,不对输出分割概率图做任何计算,直接拿来和标注图做并集,再继续训练模型,这个过程实现简单。
如图6示,本公开实施例提供了一种电子设备,包括:
存储器,用于存储信息;
处理器,与所述存储器连接,用于通过执行存储在所述存储器上的计算机可执行指令,能够实现前述一个或多个技术方案提供的深度学习模型训练方法,例如,如图1至图3所示的方法中的一个或多个。
该存储器可为各种类型的存储器,可为随机存储器、只读存储器、闪存等。所述存储器可用于信息存储,例如,存储计算机可执行指令等。所述计算机可执行指令可为各种程序指令,例如,目标程序指令和/或源程序指令等。
所述处理器可为各种类型的处理器,例如,中央处理器、微处理器、数字信号处理器、可编程阵列、数字信号处理器、专用集成电路或图像处理器 等。
所述处理器可以通过总线与所述存储器连接。所述总线可为集成电路总线等。
在一些实施例中,所述终端设备还可包括:通信接口,该通信接口可包括:网络接口、例如,局域网接口、收发天线等。所述通信接口同样与所述处理器连接,能够用于信息收发。
在一些实施例中,所述电子设备还包括摄像头,该摄像头可采集各种图像,例如,医疗影像等。
在一些实施例中,所述终端设备还包括人机交互接口,例如,所述人机交互接口可包括各种输入输出设备,例如,键盘、触摸屏等。
本公开实施例提供了一种计算机存储介质,所述计算机存储介质存储有计算机可执行代码;所述计算机可执行代码被执行后,能够实现前述一个或多个技术方案提供的深度学习模型训练方法,例如,如图1至图3所示的方法中的一个或多个。
所述存储介质包括:移动存储设备、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、磁碟或者光盘等各种可以存储程序代码的介质。所述存储介质可为非瞬间存储介质。
本公开实施例提供一种计算机程序产品,所述程序产品包括计算机可执行指令;所述计算机可执行指令被执行后,能够实现前述任意实施提供的深度学习模型训练方法,例如,如图1至图3所示的方法中的一个或多个。
在本公开所提供的几个实施例中,应该理解到,所揭露的设备和方法,可以通过其它的方式实现。以上所描述的设备实施例仅仅是示意性的,例如,所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,如:多个单元或组件可以结合,或可以集成到另一个系统,或一些特征可以忽略,或不执行。另外,所显示或讨论的各组成部分相互之间的耦合、或直接耦合、或通信连接可以是通过一些接口,设备或单元的间接耦合或通信连接,可以是电性的、机械的或其它形式的。
上述作为分离部件说明的单元可以是、或也可以不是物理上分开的,作为单元显示的部件可以是、或也可以不是物理单元,即可以位于一个地方,也可以分布到多个网络单元上;可以根据实际的需要选择其中的部分或全部单元来实现本实施例方案的目的。
另外,在本公开各实施例中的各功能单元可以全部集成在一个处理模块中,也可以是各单元分别单独作为一个单元,也可以两个或两个以上单元集成在一个单元中;上述集成的单元既可以采用硬件的形式实现,也可以采用硬件加软件功能单元的形式实现。
本公开一种实施例中公开一种计算机程序产品,程序产品包括计算机可执行指令;该计算机可执行指令被执行后,能够实现上述实施例中的深度模型训练方法。
本领域普通技术人员可以理解:实现上述方法实施例的全部或部分步骤可以通过程序指令相关的硬件来完成,前述的程序可以存储于一计算机可读取存储介质中,该程序在执行时,执行包括上述方法实施例的步骤;而前述的存储介质包括:移动存储设备、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述,仅为本公开的具体实施方式,但本公开的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本公开揭露的技术范围内,可轻易想到变化或替换,都应涵盖在本公开的保护范围之内。因此,本公开的保护范围应以所述权利要求的保护范围为准。

Claims (17)

  1. 一种深度学习模型训练方法,其中,包括:
    获取待训练模型输出的第n+1标注信息,所述待训练模型已经过n轮训练;n为大于或等于1的整数;
    基于所述训练数据及所述第n+1标注信息生成第n+1训练样本;
    将所述第n+1训练样本对所述待训练模型进行第n+1轮训练。
  2. 根据权利要求1所述的方法,其中,所述基于所述训练数据及所述第n+1标注信息生成第n+1训练样本,包括:
    基于所述训练数据及所述第n+1标注信息、及第1训练样本生成第n+1训练样本;
    或者,
    基于所述训练数据及所述第n+1标注信息、及第n训练样本生成第n+1训练样本,其中,所述第n训练样本包括:所述训练数据及第一标注信息构成的第1训练样本、及前n-1轮训练得到的标注信息与所述训练样本分别构成的第2训练样本至第n-1训练样本待训练模型。
  3. 根据权利要求1或2所述的方法,其中,所述方法包括:
    确定n是否小于N,其中,N为所述待训练模型的最大训练轮数;
    所述获取待训练模型输出的第n+1标注信息,包括:
    若n小于N,获取所述待训练模型输出的第n+1标注信息。
  4. 根据权利要求2所述的方法,其中,所述方法包括:
    获取所述训练数据及所述训练数据的初始标注信息;
    基于所述初始标注信息,生成所述第一标注信息。
  5. 根据权利要求4所述的方法,其中,
    所述获取所述训练数据及所述训练数据的初始标注信息,包括:
    获取包含有多个分割目标的训练图像及所述分割目标的外接框;
    所述基于所述初始标注信息,生成所述第一标注信息,包括:
    基于所述外接框,在所述外接框内绘制与所述分割目标形状一致的标注轮廓。
  6. 根据权利要求5所述的方法,其中,所述基于所述初始标注信息,生成所述第一标注信息,还包括:
    基于所述外接框,生成具有重叠部分的两个所述分割目标的分割边界。
  7. 根据权利要求5所述的方法,其中,
    所述基于所述外接框,在所述外接框内绘制与所述分割目标形状一致的标注轮廓,包括:
    基于所述外接框,在所述外接框内绘制与细胞形状一致的所述外接框的内接椭圆。
  8. 一种深度学习模型训练装置,其中,包括:
    标注模块,配置为获取待训练模型输出的第n+1标注信息,所述待训练模型已经过n轮训练;n为大于或等于1的整数;
    第一生成模块,配置为基于所述训练数据及所述第n+1标注信息生成第n+1训练样本;
    训练模块,配置为将所述第n+1训练样本对所述待训练模型进行第n+1轮训练待训练模型。
  9. 根据权利要求8所述的装置,其中,
    所述第一生成模块,配置为基于所述训练数据及所述第n+1标注信息、及第1训练样本生成第n+1训练样本;或者,基于所述训练数据及所述第n+1标注信息、及第n训练样本生成第n+1训练样本,其中,所述第n训练样本包括:所述训练数据及第一标注信息构成的第1训练样本、及前n-1轮训练得到的标注信息与所述训练样本分别构成的第2训练样本至第n-1 训练样本。
  10. 根据权利要求8或9所述的装置,其中,所述装置包括:
    确定模块,配置为确定n是否小于N,其中,N为所述待训练模型的最大训练轮数;
    所述标注模块,配置为若n小于N,获取所述待训练模型输出的第n+1标注信息。
  11. 根据权利要求9所述的装置,其中,所述装置包括:
    获取模块,配置为获取所述训练数据及所述训练数据的初始标注信息;
    第二生成模块,配置为基于所述初始标注信息,生成所述第一标注信息。
  12. 根据权利要求11所述的装置,其中,
    所述获取模块,配置为获取包含有多个分割目标的训练图像及所述分割目标的外接框;
    所述第二生成模块,配置为于基于所述外接框,在所述外接框内绘制与所述分割目标形状一致的标注轮廓。
  13. 根据权利要求12所述的装置,其中,所述第一生成模块,配置为基于所述外接框,生成具有重叠部分的两个所述分割目标的分割边界。
  14. 根据权利要求12所述的装置,其中,
    所述第二生成模块,配置为基于所述外接框,在所述外接框内绘制与细胞形状一致的所述外接框的内接椭圆。
  15. 一种计算机存储介质,所述计算机存储介质存储有计算机可执行指令;所述计算机可执行指令;所述计算机可执行指令被执行后,能够实现权利要求1至7任一项所述的方法。
  16. 一种电子设备,其中,包括:
    存储器;
    处理器,与所述存储器连接,用于通过执行存储在所述存储器上的计算机可执行指令实现前述权利要求1至7任一项所述的方法。
  17. 一种计算机程序产品,所述程序产品包括计算机可执行指令;所述计算机可执行指令被执行后,能够实现前述权利要求1至7任一项所述的方法。
PCT/CN2019/114493 2018-12-29 2019-10-30 深度模型训练方法及装置、电子设备及存储介质 Ceased WO2020134532A1 (zh)

Priority Applications (4)

Application Number Priority Date Filing Date Title
JP2021507067A JP7158563B2 (ja) 2018-12-29 2019-10-30 深層モデルの訓練方法及びその装置、電子機器並びに記憶媒体
KR1020217004148A KR20210028716A (ko) 2018-12-29 2019-10-30 딥러닝 모델의 트레이닝 방법, 장치, 전자 기기 및 저장 매체
SG11202100043SA SG11202100043SA (en) 2018-12-29 2019-10-30 Deep model training method and apparatus, electronic device, and storage medium
US17/136,072 US20210118140A1 (en) 2018-12-29 2020-12-29 Deep model training method and apparatus, electronic device, and storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201811646430.5A CN109740752B (zh) 2018-12-29 2018-12-29 深度模型训练方法及装置、电子设备及存储介质
CN201811646430.5 2018-12-29

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US17/136,072 Continuation US20210118140A1 (en) 2018-12-29 2020-12-29 Deep model training method and apparatus, electronic device, and storage medium

Publications (1)

Publication Number Publication Date
WO2020134532A1 true WO2020134532A1 (zh) 2020-07-02

Family

ID=66362804

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2019/114493 Ceased WO2020134532A1 (zh) 2018-12-29 2019-10-30 深度模型训练方法及装置、电子设备及存储介质

Country Status (7)

Country Link
US (1) US20210118140A1 (zh)
JP (1) JP7158563B2 (zh)
KR (1) KR20210028716A (zh)
CN (1) CN109740752B (zh)
SG (1) SG11202100043SA (zh)
TW (1) TW202026958A (zh)
WO (1) WO2020134532A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111881966A (zh) * 2020-07-20 2020-11-03 北京市商汤科技开发有限公司 神经网络训练方法、装置、设备和存储介质

Families Citing this family (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109740752B (zh) * 2018-12-29 2022-01-04 北京市商汤科技开发有限公司 深度模型训练方法及装置、电子设备及存储介质
CN110399927B (zh) * 2019-07-26 2022-02-01 玖壹叁陆零医学科技南京有限公司 识别模型训练方法、目标识别方法及装置
CN110909688B (zh) * 2019-11-26 2020-07-28 南京甄视智能科技有限公司 人脸检测小模型优化训练方法、人脸检测方法及计算机系统
TWI771004B (zh) 2021-05-14 2022-07-11 財團法人工業技術研究院 物件姿態估測系統及其執行方法與圖案化使用者介面
CN113487575B (zh) * 2021-07-13 2024-01-16 中国信息通信研究院 用于训练医学影像检测模型的方法及装置、设备、可读存储介质
CN113947771B (zh) * 2021-10-15 2023-06-27 北京百度网讯科技有限公司 图像识别方法、装置、设备、存储介质以及程序产品
CN114842377B (zh) * 2022-04-22 2025-04-29 上海西井科技股份有限公司 目标检测模型训练方法、系统、设备及存储介质

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106250874A (zh) * 2016-08-16 2016-12-21 东方网力科技股份有限公司 一种服饰及随身物品的识别方法和装置
CN107169556A (zh) * 2017-05-15 2017-09-15 电子科技大学 基于深度学习的干细胞自动计数方法
US20180114123A1 (en) * 2016-10-24 2018-04-26 Samsung Sds Co., Ltd. Rule generation method and apparatus using deep learning
CN108764372A (zh) * 2018-06-08 2018-11-06 Oppo广东移动通信有限公司 数据集的构建方法和装置、移动终端、可读存储介质
CN109740752A (zh) * 2018-12-29 2019-05-10 北京市商汤科技开发有限公司 深度模型训练方法及装置、电子设备及存储介质

Family Cites Families (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN102074034B (zh) * 2011-01-06 2013-11-06 西安电子科技大学 多模型人体运动跟踪方法
CN102184541B (zh) * 2011-05-04 2012-09-05 西安电子科技大学 多目标优化人体运动跟踪方法
CN102622766A (zh) * 2012-03-01 2012-08-01 西安电子科技大学 多目标优化的多镜头人体运动跟踪方法
JP2015114172A (ja) * 2013-12-10 2015-06-22 オリンパスソフトウェアテクノロジー株式会社 画像処理装置、顕微鏡システム、画像処理方法、及び画像処理プログラム
US20180268292A1 (en) * 2017-03-17 2018-09-20 Nec Laboratories America, Inc. Learning efficient object detection models with knowledge distillation
AU2018269941A1 (en) * 2017-05-14 2019-12-05 Digital Reasoning Systems, Inc. Systems and methods for rapidly building, managing, and sharing machine learning models
US20190102674A1 (en) * 2017-09-29 2019-04-04 Here Global B.V. Method, apparatus, and system for selecting training observations for machine learning models
US10997727B2 (en) * 2017-11-07 2021-05-04 Align Technology, Inc. Deep learning for tooth detection and evaluation
CN109066861A (zh) * 2018-08-20 2018-12-21 四川超影科技有限公司 基于机器视觉的智能巡检机器人自动充电控制方法

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106250874A (zh) * 2016-08-16 2016-12-21 东方网力科技股份有限公司 一种服饰及随身物品的识别方法和装置
US20180114123A1 (en) * 2016-10-24 2018-04-26 Samsung Sds Co., Ltd. Rule generation method and apparatus using deep learning
CN107169556A (zh) * 2017-05-15 2017-09-15 电子科技大学 基于深度学习的干细胞自动计数方法
CN108764372A (zh) * 2018-06-08 2018-11-06 Oppo广东移动通信有限公司 数据集的构建方法和装置、移动终端、可读存储介质
CN109740752A (zh) * 2018-12-29 2019-05-10 北京市商汤科技开发有限公司 深度模型训练方法及装置、电子设备及存储介质

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111881966A (zh) * 2020-07-20 2020-11-03 北京市商汤科技开发有限公司 神经网络训练方法、装置、设备和存储介质

Also Published As

Publication number Publication date
CN109740752B (zh) 2022-01-04
JP2021533505A (ja) 2021-12-02
CN109740752A (zh) 2019-05-10
US20210118140A1 (en) 2021-04-22
KR20210028716A (ko) 2021-03-12
TW202026958A (zh) 2020-07-16
SG11202100043SA (en) 2021-02-25
JP7158563B2 (ja) 2022-10-21

Similar Documents

Publication Publication Date Title
TWI747120B (zh) 深度模型訓練方法及裝置、電子設備及儲存介質
CN109740752B (zh) 深度模型训练方法及装置、电子设备及存储介质
US11842487B2 (en) Detection model training method and apparatus, computer device and storage medium
US12430900B2 (en) Object recognition using spatial and timing information of object images at diferent times
EP3961500A1 (en) Medical image detection method based on deep learning, and related device
WO2020125495A1 (zh) 一种全景分割方法、装置及设备
WO2018108129A1 (zh) 用于识别物体类别的方法及装置、电子设备
KR20210082234A (ko) 이미지 처리 방법 및 장치, 전자 기기 및 기억 매체
CN110348294A (zh) Pdf文档中图表的定位方法、装置及计算机设备
CN112767329A (zh) 图像处理方法及装置、电子设备
WO2022227218A1 (zh) 药名识别方法、装置、计算机设备和存储介质
CN112634369A (zh) 空间与或图模型生成方法、装置、电子设备和存储介质
CN109376756B (zh) 基于深度学习的上腹部转移淋巴结节自动识别系统、计算机设备、存储介质
CN111445440A (zh) 一种医学图像分析方法、设备和存储介质
CN111598025A (zh) 图像识别模型的训练方法和装置
CN114820679A (zh) 图像标注方法、装置、电子设备和存储介质
CN111932557B (zh) 基于集成学习和概率图模型的图像语义分割方法及装置
JP2025526192A (ja) 画像処理方法およびその装置、機器、記憶媒体、並びにコンピュータプログラム
US12079950B2 (en) Image processing method and apparatus, smart microscope, readable storage medium and device
CN112750124B (zh) 模型生成、图像分割方法、装置、电子设备及存储介质
CN112597328B (zh) 标注方法、装置、设备及介质
CN115527092A (zh) 模型训练方法、装置及设备
HK40006395B (zh) 深度模型训练方法及装置、电子设备及存储介质
CN116363628A (zh) 标志检测方法、装置、非易失性存储介质及计算机设备
HK40006395A (zh) 深度模型训练方法及装置、电子设备及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19906033

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2021507067

Country of ref document: JP

Kind code of ref document: A

ENP Entry into the national phase

Ref document number: 20217004148

Country of ref document: KR

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 19906033

Country of ref document: EP

Kind code of ref document: A1