WO2025185014A1 - 图像标注数据生成、模型训练方法、装置、设备及介质 - Google Patents
图像标注数据生成、模型训练方法、装置、设备及介质Info
- Publication number
- WO2025185014A1 WO2025185014A1 PCT/CN2024/100320 CN2024100320W WO2025185014A1 WO 2025185014 A1 WO2025185014 A1 WO 2025185014A1 CN 2024100320 W CN2024100320 W CN 2024100320W WO 2025185014 A1 WO2025185014 A1 WO 2025185014A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- visual
- target recognition
- sample image
- target
- adjective
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/40—Document-oriented image-based pattern recognition
- G06V30/41—Analysis of document content
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
- G06V30/19—Recognition using electronic means
- G06V30/19007—Matching; Proximity measures
- G06V30/19093—Proximity measures, i.e. similarity or distance measures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
- G06V30/19—Recognition using electronic means
- G06V30/191—Design or setup of recognition systems or techniques; Extraction of features in feature space; Clustering techniques; Blind source separation
- G06V30/19147—Obtaining sets of training patterns; Bootstrap methods, e.g. bagging or boosting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
- G06V30/19—Recognition using electronic means
- G06V30/191—Design or setup of recognition systems or techniques; Extraction of features in feature space; Clustering techniques; Blind source separation
- G06V30/1918—Fusion techniques, i.e. combining data from various sources, e.g. sensor fusion
Definitions
- This application relates to the fields of image processing, artificial intelligence, computer vision, large language models and smart cities, for example, to a method, device, equipment and medium for generating image annotation data, training a model.
- the present application provides a method, apparatus, device and medium for generating image annotation data and model training.
- a method for generating image annotation data comprising:
- the target description information of the target recognition object is used as the annotation data of the target recognition object; wherein the annotation data is used to train a visual semantic understanding model, and the visual semantic understanding model is used to output a description text of the visual content based on the input visual content.
- a method for training a visual semantic understanding model comprising:
- Acquire sample data wherein the sample data includes sample images and annotation data.
- the annotation data is generated by the image annotation data generation method as described in any embodiment of the present application;
- the sample data is used to train a visual semantic understanding model.
- a device for generating image annotation data comprising:
- a sample data acquisition module is configured to acquire a sample image, a target object in the sample image, and visual description information of the target object;
- a visual description verification module configured to verify the words in the visual description information based on the sample image and the target recognition object, and obtain a verification result of the visual description information
- a target description information adjustment module configured to adjust the words in the visual description information according to the verification result of the visual description information to obtain the target description information of the target recognition object;
- the annotation data generation module is configured to use the target description information of the target recognition object as the annotation data of the target recognition object; wherein the annotation data is used to train the visual semantic understanding model, and the visual semantic understanding model is used to output the description text of the visual content based on the input visual content.
- a training device for a visual semantic understanding model comprising:
- sample data acquisition module configured to acquire sample data, wherein the sample data includes a sample image and annotation data, and the annotation data is generated by the image annotation data generation method as described in any embodiment of the present application;
- the model training module is configured to use the sample data to train a visual semantic understanding model.
- an electronic device including:
- the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image annotation data generation method described in any embodiment of the present application, or the visual semantic understanding model training method described in any embodiment of the present application.
- a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the image annotation data generation method described in any embodiment of the present application, or the visual semantic understanding model training method described in any embodiment of the present application.
- a computer program product comprising a computer program, wherein when the computer program is executed by a processor, the image annotation data according to any embodiment of the present application is implemented. According to the generation method, or the training method of the visual semantic understanding model described in any embodiment of the present application.
- FIG1 is a flowchart of a method for generating image annotation data according to an embodiment of the present application
- FIG2 is a flowchart of another method for generating image annotation data according to an embodiment of the present application.
- FIG3 is a flowchart of another method for generating image annotation data according to an embodiment of the present application.
- FIG4 is a flowchart of a method for training a visual semantic understanding model according to an embodiment of the present application
- FIG5 is a scene diagram of a method for generating image annotation data according to an embodiment of the present application.
- FIG6 is a schematic diagram of the structure of an apparatus for generating image annotation data according to an embodiment of the present application.
- FIG7 is a schematic diagram of the structure of a training device for a visual semantic understanding model according to an embodiment of the present application.
- FIG8 is a block diagram of an electronic device provided according to an embodiment of the present application.
- FIG. 1 is a flowchart of a method for generating image annotation data according to an embodiment of the present application.
- This embodiment can be applied to the case of generating descriptive content of an image.
- the method of this embodiment can be executed by an image annotation data generating device, which can be implemented in software and/or hardware, and can be configured in an electronic device with a certain data computing capability.
- the electronic device can be a client device or a server device.
- the client device can include: a personal computer, a laptop computer, a smart phone, a tablet computer, an Internet of Things device or a portable wearable device, etc.
- the Internet of Things device can be a smart speaker, a smart TV, a smart air conditioner or a smart car device, etc.
- the portable wearable device can be a smart watch, a smart bracelet or a head-mounted device, etc.
- S101 Acquire a sample image, a target object in the sample image, and visual description information of the target object.
- the sample image includes at least one target recognition object. If there are multiple target recognition objects, different target recognition objects may correspond to different objects, and/or different target recognition objects may correspond to different types of objects.
- the sample image includes a target recognition object of a person, a target recognition object of a cat, and
- a sample image may include at least one target object of a cat.
- Object detection can be performed on the sample image to obtain a target detection frame for the object in the sample image, and each target detection frame can be identified as a target object.
- the sample image can be segmented to obtain segmented regions in the sample image, and each region of the object can be identified as a target object, while the background region does not include the object.
- the target object can be classified to obtain a category of the target object.
- Visual description information is used to describe the image content of the target object and represent the semantic information of the image corresponding to the target object.
- Visual description information can refer to a statement describing the visually visible content.
- the visual description information can include at least one of the following: target object content, subject content, and relationship content.
- the subject content can refer to key information about the sample image.
- the subject content can include scenes and/or events.
- the target object content can refer to the content of the target object itself.
- the target object can include at least one of the following: size, state, position, and color.
- Relationship content can refer to the relationship between the target object and other objects in the image.
- other objects can include objects within the target object and/or objects other than the target object in the sample image.
- Relationships can include at least one of positional relationships, dynamic and static relationships, and state relationships.
- the visual description information can include other content, which is not limited to this.
- the sample image is an image of the surrounding environment captured by a vehicle while driving;
- the target object can include pedestrians, vehicles, road signs, indicator lights, and guardrails.
- the corresponding visual description information can include: "There is a pedestrian directly ahead.”
- Sample images and target recognition objects can be input into a pre-trained generative model to obtain visual description information.
- the generative model's processing process is as follows: the sample image and target recognition object are encoded separately to obtain vector representations corresponding to the sample image and target recognition object, respectively; prompt word text is obtained and encoded to obtain a vector representation of the prompt word text; the two vector representations are fused to obtain a fusion result; the fusion result is decoded to obtain visual description information.
- sample images, target recognition objects, and the category of the target recognition object can also be input into the generative model to obtain visual description information.
- the generative model can perform semantic understanding of visual content and generate description information.
- S102 Verify the words in the visual description information according to the sample image and the target recognition object to obtain a verification result of the visual description information.
- the visual description information is segmented to obtain at least one word. Each word is verified to determine if it is correct.
- Image information can be extracted from a sample image and a target object for verification of the words in the visual description information. Verification can include detecting whether a word is correct or exists based on the target object and the sample image. The verification result of the visual description information can determine if each word in the visual description information is correct, retaining correct words and removing incorrect words.
- S103 Adjust the words in the visual description information according to the verification result of the visual description information to obtain target description information of the target recognition object.
- the visual description information may refer to a sentence. Accordingly, the words in the visual description information may be adjusted by retaining correct words and modifying incorrect words to obtain a new sentence, which is then determined as the target description information for the target recognition object. Furthermore, synonyms of at least one word in the new sentence may be obtained, and the at least one word and the synonyms may be reorganized to form more new sentences. These multiple new sentences may then be screened to obtain more accurate and standardized sentences, which are then determined as the target description information.
- the target description information is added to the sample data as annotation data. This can be understood as the true value of the description information of the target recognition object.
- the sample images and their target description information are used to train the visual semantic understanding model.
- the trained visual semantic understanding model can semantically understand the input image and output a description text of the image as the semantic understanding content of the image.
- the sample image may include at least one target recognition object, and each target recognition object generates target description information.
- the sample image and the target description information of at least one target recognition object in the sample image can be used as training data to train the visual semantic understanding model.
- manual image annotation can provide optimal labeling, it requires significant manpower and resources.
- Large models require enormous amounts of data, and some large enterprises have spent thousands of people working together for months to complete all required annotations, resulting in expensive labeling costs.
- many enterprises face the dilemma of data and large amounts of data annotation, such as in the current era of large models.
- machine learning and deep learning methods can be used for automated annotation in related fields (such as detection and segmentation), but this approach carries significant uncertainty and inaccuracy.
- the visual description information of the target recognition object is automatically extracted, and the words in the visual description information are verified to obtain the verification result.
- the words in the visual description information are adjusted to obtain the target description information as the annotation data of the target recognition object.
- the image semantic understanding model is trained to obtain a model for semantic understanding of the input visual content, thereby improving the generation efficiency of image annotation data, reducing the annotation cost, and taking into account the annotation accuracy, and then quickly generating correct annotation data to train the model and improve the accuracy of the description text output by the visual semantic understanding model.
- FIG2 is a flowchart of another method for generating image annotation data according to an embodiment of the present application, which is described based on the above technical solution and can be combined with the above multiple optional implementation methods.
- the method of verifying the words in the visual description information according to the sample image and the target recognition object to obtain the verification result of the visual description information includes: dividing the visual description information into at least one visual noun and at least one visual adjective; verifying the at least one visual noun according to the sample image and the target recognition object to obtain a noun verification result corresponding to the at least one visual noun; verifying the at least one visual adjective according to the sample image and the target recognition object to obtain an adjective verification result corresponding to the at least one visual adjective; and determining the noun verification result and the adjective verification result as the verification result of the visual description information.
- S201 Acquire a sample image, a target object in the sample image, and visual description information of the target object.
- S202 Divide the visual description information into at least one visual noun and at least one visual adjective.
- Visual nouns refer to visually visible nouns related to the target object.
- Visual adjectives refer to visually visible adjectives related to visual nouns.
- Visual nouns are used to describe the category of the target object.
- Visual adjectives are used to limit the target object.
- Visual adjectives can be divided into material, color, or location, etc.
- the visual description information is a white ragdoll cat
- the visual noun can include cat
- the visual adjectives can include white and ragdoll.
- the visual description information is a black dog lying on the carpet.
- Visual nouns include: dog and carpet.
- Visual adjectives can include black and dog on the carpet.
- the visual description information can be divided into at least one noun-adjective pair.
- the visual description information is "a red apple”
- the noun-adjective pair extracted is ⁇ red>, ⁇ apple> ⁇ .
- the visual description information is "a red apple”
- the noun-adjective pair extracted is ⁇ red>, ⁇ apple> ⁇ .
- the visual description information is "a black dog lying on the carpet”
- the noun-adjective pairs extracted are ⁇ black>, ⁇ dog> ⁇ , ⁇ lying>, ⁇ dog> ⁇ , ⁇ up>, ⁇ dog> ⁇ , and ⁇ down>, ⁇ carpet> ⁇ .
- This visual description information corresponds to the target recognition object.
- a sample image can contain multiple target recognition objects.
- the visual description information of object 1 is "white ragdoll cat”
- the visual description information of object 2 is "black dog lying on the carpet”.
- the noun-adjective pairs that can be formed in a sample image include: ⁇ white>, ⁇ cat>, ⁇ object 1> ⁇ , ⁇ ragdoll>, ⁇ cat>, ⁇ object 1> ⁇ , ⁇ black>, ⁇ dog>, ⁇ object 2> ⁇ , ⁇ lying down>, ⁇ dog>, ⁇ object 2> ⁇ , ⁇ up>, ⁇ dog>, ⁇ object 2> ⁇ and ⁇ down>, ⁇ carpet>, ⁇ object 2> ⁇ .
- S203 Verify the at least one visual noun according to the sample image and the target recognition object to obtain a noun verification result corresponding to the at least one visual noun.
- the category of the target object can be obtained based on the sample image and the target object, and the visual noun can be verified. For example, it can be detected whether the category of the target object is similar to the visual noun.
- the noun verification result may refer to a verification result of whether a visual noun is correct.
- One visual noun corresponds to one noun verification result.
- the at least one visual noun is verified based on the sample image and the target recognition object to obtain a noun verification result corresponding to the at least one visual noun, including: obtaining category description information of the target recognition object; for the at least one visual noun, detecting the consistency between the category description information of the target recognition object and the at least one visual noun to obtain a noun verification result corresponding to the at least one visual noun.
- the category description information may be the target object and the category of the target object detected when detecting the object in the sample image, and the category of the target object determined as the category description information. For example, a detection frame of the target object is detected in the sample image, along with the category of the detection frame, and the image region corresponding to the detection frame is determined as the target object; the category of the detection frame is determined as the category description information of the target object.
- the category description information and the target object can be stored in association to facilitate subsequent retrieval of the category description information of the target object.
- the category description information is compared with the visual noun. If they are consistent, the noun verification result of the visual noun is determined to be correct; if they are inconsistent, the noun verification result of the visual noun is determined to be incorrect.
- the similarity between the category description information and the visual noun can be calculated and the similarity can be determined as the noun verification result. Alternatively, when the similarity is greater than or equal to a corresponding similarity threshold, the noun verification result of the visual noun is determined to be correct; when the similarity is less than the corresponding similarity threshold, the noun verification result of the visual noun is determined to be incorrect.
- the category description information of the target recognition object By comparing the category description information of the target recognition object with the visual noun for consistency, a comparison result is obtained, and based on the comparison result, the noun verification result of the visual noun is determined.
- the visual noun can be verified according to the category obtained by the sample image when detecting the object, avoiding the verification of the visual noun through additional complex and redundant steps, simplifying the verification operation of the visual noun, and improving the verification efficiency of the visual noun.
- the accuracy of the category of the target recognition object is high, and the category of the target recognition object is used for verification to improve the verification accuracy of the visual noun.
- S204 Verify the at least one visual adjective according to the sample image and the target recognition object to obtain an adjective verification result corresponding to the at least one visual adjective.
- Visual adjectives typically qualify or describe a visual noun. Verification of visual adjectives typically involves checking whether the characteristics of the visual noun they describe are consistent with the visual adjective. If they are consistent, the visual adjective is considered correct; if not, it is considered incorrect. Each visual adjective corresponds to one adjective verification result.
- the at least one visual adjective is verified based on the sample image and the target recognition object to obtain an adjective verification result corresponding to the at least one visual adjective.
- the method includes: generating interactive dialogue content for the at least one visual adjective based on the at least one visual adjective and the visual noun corresponding to the at least one visual adjective; using the sample image and the target recognition object as context, and inputting the interactive dialogue content into a large language model to obtain a result of determining the existence of the at least one visual adjective; and generating an adjective verification result for the at least one visual adjective based on the result of determining the existence of the at least one visual adjective.
- the interactive dialogue content is used to detect whether the description content of the visual noun contains content corresponding to the corresponding visual adjective.
- the interactive dialogue content is usually a correctness judgment statement.
- the interactive dialogue content includes a visual adjective, the visual noun corresponding to the visual adjective, and the judgment content of whether the visual noun corresponding to the visual adjective is associated with the visual adjective.
- the judgment content includes: a judgment word or a predicate;
- the association between a visual noun and a visual adjective can mean that the visual noun has the property corresponding to the visual adjective, for example, the visual noun is cat, the visual adjective is white, and the color of the cat in the sample image is white, that is, the visual noun has the color property represented by the visual adjective, and it is determined that the visual noun is associated with the visual adjective.
- a visual noun and a visual adjective are not associated. This can mean that the visual noun does not possess the property corresponding to the visual adjective.
- the property represented by the visual adjective is different from the property of the visual noun.
- the color of the cat is black, meaning that the visual noun does not possess the color property represented by the visual adjective, thus determining that the visual noun and the visual adjective are not associated.
- the property represented by the visual noun is unrelated to the property represented by the visual adjective.
- the visual noun is "cat" and the visual adjective is "cat's eyes are blue.”
- the cat is shown from behind, so it is impossible to determine the eye color of the visual noun, thus determining that the visual noun and the visual adjective are not associated.
- the interactive dialogue content is used as input to the large language model.
- the large language model relies on its semantic understanding of the target recognition object in the sample image to determine whether the interactive dialogue content is correct.
- the large language model outputs the result of whether the interactive dialogue content is correct.
- the visual adjective presence determination result can refer to whether the visual noun corresponding to the visual adjective possesses the property of the visual adjective.
- Inputting the sample image and target recognition object as context into the large language model may mean that the large language model responds to the interactive dialogue content based on the content of the sample image and target recognition object, and outputs the existence judgment result of the visual adjective. It may be that the sample image is input into the large language model as context.
- the sample image and target recognition object may be input into the large language model as context, and the interactive dialogue content may be input into the large language model.
- the large language model judges the correctness of the interactive dialogue content based on the semantic understanding content of the sample image and target recognition object, and obtains the existence judgment result of the visual adjective.
- the sample image and target recognition object may be input into the large language model at the same time as the interactive dialogue content, or multiple rounds of dialogue may be conducted, first inputting the sample image and target recognition object into the large language model for dialogue, and then inputting the interactive dialogue content into the large language model, that is, in multiple rounds of dialogue, the previous dialogue includes the sample image and target recognition object, and the subsequent dialogue includes the interactive dialogue content.
- the large language model may be the aforementioned generative model, which is used to input the sample image and target recognition object.
- Image and target recognition objects are input, and visual description information is output.
- the next round of dialogue is carried out, the interactive dialogue content is input, and the judgment result of the interactive dialogue content is output, that is, the judgment result of the existence of visual adjectives.
- the visual description information is: white ragdoll cat.
- the visual adjective is white
- the corresponding visual noun is cat.
- the interactive dialogue content is: Is the cat white?
- the interactive dialogue content is: Do white cats exist? Taking the sample image and the target recognition object as the context, if the large language model outputs the result of "yes” for the aforementioned interactive dialogue content, then the existence judgment result of the visual adjective is "yes", and the corresponding adjective verification result of the visual adjective is correct; if the large language model outputs the result of "does not exist" for the aforementioned interactive dialogue content, then the existence judgment result of the visual adjective is "does not exist", and the corresponding adjective verification result of the visual adjective is incorrect.
- a visual noun can correspond to multiple visual adjectives, and one visual adjective corresponds to one visual noun.
- the existence judgment result is obtained for each visual adjective, that is, each visual adjective generates corresponding interactive dialogue content, and the existence judgment result is obtained through the large language model. If the verification result of a visual adjective is wrong, only the visual adjective will be eliminated or corrected. Since there may be other visual adjectives corresponding to the same visual noun as the visual adjective, the visual noun corresponding to the visual adjective cannot be directly eliminated or corrected, and the corresponding visual noun needs to be verified using other verification methods.
- interactive dialogue content is generated for judging whether visual adjectives exist, and the interactive dialogue content is input into a large language model with sample images and target recognition objects as context to detect whether the visual noun has the properties of the visual adjective, thereby judging whether the visual adjective exists, and then determining the verification result of the visual adjective.
- the judgment can be made based on the visual content of the sample image itself, avoiding the verification of visual adjectives through additional complex and redundant steps, simplifying the verification operation of visual adjectives, and improving the verification efficiency of visual adjectives.
- the visual nouns are associated with the visual adjectives, and the association judgment of the visual nouns and visual adjectives is made, making full use of the characteristic that visual adjectives cannot exist alone, verifying both the existence of the visual adjective and the association of the visual noun with the visual adjective, thereby improving the verification accuracy of the visual noun.
- S205 Determine the noun verification result and the adjective verification result as the verification result of the visual description information.
- the verification results of all nouns and all adjectives are determined as the verification results of the visual description information.
- S206 Adjust the words in the visual description information according to the verification result of the visual description information to obtain target description information of the target recognition object.
- Correct or delete the words corresponding to the incorrect verification results Obtain similar words for the words corresponding to the correct verification results. Reorganize the correct words, similar words, and corrected words to obtain new visual description information. Filter the multiple description information to obtain the target description information of the target recognition object.
- the wording in the visual description information is adjusted according to the verification result of the visual description information to obtain the target description information of the target recognition object, including: screening the at least one visual noun according to the verification result of each noun to obtain at least one target noun; screening the at least one visual adjective according to the verification result of each adjective to obtain at least one target adjective; generating at least one descriptive phrase according to the at least one target noun and the at least one target adjective; and generating the target description information of the target recognition object according to the at least one descriptive phrase.
- a description phrase includes a noun and an adjective.
- the target adjective can be combined with the defined target noun to obtain a description phrase.
- a visual adjective and the visual noun corresponding to the visual adjective are formed into a phrase, and the phrase is filtered according to the target noun and the target adjective to determine the description phrase.
- the target noun and the target adjective in the description phrase correspond to each other, that is, the target adjective defines and describes the target noun in the same description phrase.
- Sentences are constructed based on the nouns and adjectives in the description phrase to form sentences, which are determined as target description information of the target recognition object.
- a target recognition object can generate at least one target description information. All target description information can be used as the annotation data of the target recognition object, or the target description information can be filtered, for example, by scoring and filtering based on grammar, standardization and expression richness, to obtain a target description information as the annotation data of the target recognition object.
- corresponding synonyms can be added to at least one target noun and at least one target adjective as target nouns or target adjectives, and based on the added target nouns and target adjectives, they can be arranged and combined to form descriptive phrases, which can enrich the content of the descriptive phrases, thereby enriching the content of the target description information and increasing the representativeness of the target description content.
- the target nouns and target adjectives are obtained, and the descriptive phrases are generated in combination.
- the target description information of the target recognition object is generated.
- the words of different parts of speech can be screened separately to achieve accurate verification, and the screened and retained words are combined accordingly, that is, the correlation between nouns and adjectives is screened to achieve correlation verification, and the target description information is generated according to the verified descriptive phrases to improve the accuracy of the target description information.
- the visual description information by segmenting the visual description information, visual nouns and visual adjectives are obtained, and the visual nouns and visual adjectives are verified respectively to obtain the verification results of the visual description information.
- the visual description information can be disassembled and verified based on words as units, and the description information can be verified at a finer granularity.
- words with different parts of speech are verified separately to achieve targeted description word verification and improve the verification accuracy of the description information.
- FIG3 is a flowchart of another method for generating image annotation data according to an embodiment of the present application, which is described based on the above technical solution and can be combined with the above multiple optional implementation methods.
- the sample image and target recognition object are obtained by: obtaining a real-time updated object category and adding the obtained real-time updated object category to a dynamic vocabulary; obtaining a sample image; identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image, and obtaining the target recognition object in the sample image.
- S301 Acquire a sample image, a target object in the sample image, and visual description information of the target object.
- S302 Verify the words in the visual description information according to the sample image and the target recognition object to obtain a verification result of the visual description information.
- S303 Adjust the words in the visual description information according to the verification result of the visual description information to obtain target description information of the target recognition object.
- the sample image and the target recognition object are obtained by: acquiring the object category updated in real time, and adding the object category updated in real time to a dynamic vocabulary.
- the object category can be the category of the object that can be identified in the image.
- the object category that is updated in real time refers to the category of the object that can be identified by the image and collected from the network hot spots or real-time events.
- the dynamic vocabulary is used to store the object category. Exemplarily, real-time images or hot spot images can be acquired periodically, and objects can be identified therefrom, and the object category of the object can be added to the dynamic vocabulary.
- the steps of obtaining the sample image and the target identification object when generating the annotation data can be independent of each other.
- the sample image and the target identification object can be updated periodically, and a large number of sample images and target identification objects can be stored.
- the sample image and at least one object identified in the sample image can be updated periodically.
- a target object is used as a piece of sample data; or a sample image and a target object are used as a piece of sample data. If multiple target objects are recognized in the sample image, multiple pieces of sample data are generated accordingly.
- at least one piece of sample data is obtained from a large amount of pre-stored sample data, for example, the latest multiple pieces of sample data can be obtained.
- static vocabulary is not open to images, meaning that in the future, new business needs will not allow for the extraction of more diverse frame information. Therefore, tagging is used to extract unique words from each image and add them to the dynamic vocabulary.
- the combination of static and dynamic vocabulary greatly enriches the content of object categories.
- obtaining the object category updated in real time includes: obtaining a visual set updated in real time; performing target detection on the visual data in the visual set to obtain at least one first object category; and/or performing semantic understanding on the visual data in the visual set to obtain at least one second object category; and determining the first object category and/or the second object category as the object category updated in real time.
- a visual set may refer to a collection of images or videos.
- the visual data may include a captured image or a frame of video.
- the first object category may be an object category obtained by object detection in the latest visual data. Object detection may be performed on the visual data to obtain at least one first object and a first object category for the at least one first object.
- a Recognize anything plus model can be proposed to perform target detection on visual data through the Inject Semantic Concepts Into Image Tagging Open-Set Recognition method to obtain a first object category.
- the extracted first object category can ensure the word diversity of the dynamic vocabulary and the different extracted visual data, thereby avoiding the homogenization of the first object category caused by extracting the first object category from the same or similar visual data, and increasing the richness and representativeness of the first object category.
- the second object category may be an object category obtained by semantically understanding the latest visual data.
- the semantic understanding of the visual data is performed to obtain at least one second object and at least one second object category of the second object.
- the first object category and the second object category are obtained in different ways, but both are obtained from Object categories extracted from visual data.
- a large language model can be used to describe visual data. For example, if the large language model is fed the question "What objects might be included in the following scene?", it will respond based on the visual data, extracting the category terms from the response to obtain a second object category. Describing the second object category based on the visual data ensures the integrity and diversity of the dynamic vocabulary.
- Only the first object category or only the second object category may be acquired and added to the dynamic vocabulary.
- a first object category and/or a second object category is obtained, which is added to a dynamic vocabulary to enrich the words in the dynamic vocabulary.
- objects in sample images are identified, thereby increasing the diversity and representativeness of the identified objects, thereby making the labeled data real-time, reliable and diverse.
- the sample image may be the same as or different from the image used to identify the object category updated in real time.
- the sample image may be collected, obtained from a public channel, or obtained from an undisclosed channel with authorization.
- Each word in the dynamic vocabulary that is updated in real time is used to identify whether there is an object corresponding to each word in the sample image. If there is an object, the object is determined as the target recognition object.
- the corresponding word can be determined as the target category of the target recognition object, or the target category of at least one target recognition object can be identified while identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image.
- the sample image and all target recognition objects identified in the sample image are combined to generate a piece of sample data.
- the target recognition object can be a target detection frame.
- the sample image, at least one target recognition object identified in the sample image, and the target category of the at least one identified target recognition object can be combined to generate sample data.
- the target identification object can also be pre-processed to reduce erroneous and redundant data.
- the method further includes: detecting the intersection-and-union ratio between multiple target recognition objects in the sample image; detecting the multiple target recognition objects according to the intersection-and-union ratio between the multiple target recognition objects to obtain multiple similar target recognition objects; fusing the multiple similar target recognition objects to obtain a fusion result; and fusing the multiple target recognition objects in the sample image according to the fusion result. to update.
- multiple target recognition objects there are multiple target recognition objects in the sample image, and these multiple target recognition objects may contain redundant data. If there is only one target recognition object, this step can be omitted.
- Multiple target recognition objects can be deduplicated. Deduplication can be performed using the intersection-of-union (IoU) ratio between multiple target recognition objects. For each of the multiple target recognition objects, the IoU ratio is calculated between the target detection frames or image regions of each pair of target recognition objects. When the IoU ratio is greater than or equal to a preset IoU threshold, the two target recognition objects are determined to be similar, resulting in two similar target recognition objects. A similarity comparison is performed between each pair of target recognition objects to obtain multiple similar target recognition objects.
- IoU intersection-of-union
- semantic understanding can be performed on the detection frames of the target recognition objects to obtain target description information for the detection frames and divide it into visual noun-visual adjective pairs.
- IoU ratio of two target recognition objects is greater than or equal to a preset IoU threshold, and the corresponding visual noun-visual adjective pairs in the target description information are the same, the two target recognition objects are determined to be similar. If only the IoU ratio is greater than or equal to the preset IoU threshold, or if only the word pairs are the same, the two target recognition objects cannot be determined to be similar.
- a fusion method for a group of similar multiple target recognition objects may be: selecting one of the target recognition objects as the fusion result, eliminating the remaining target recognition objects, or calculating the intersection or union of the multiple target recognition objects to determine the fusion result.
- each target recognition object has a target category, and the number of identical target categories is counted as the number of occurrences of the identical target category.
- the target category with the largest number of occurrences in the group of similar multiple target recognition objects is counted and determined as the target category of the fusion result.
- Updating the target recognition object in the sample image according to the fusion result may be to use the fusion result as the target recognition object finally identified in the sample image to update the target recognition object in the sample image.
- redundant target recognition objects in the sample image are eliminated to update the target recognition object in the sample image.
- intersection-over-union method By using the intersection-over-union method to eliminate redundant data from multiple target recognition objects identified in the sample image, the redundant data of the target recognition objects is reduced, the multiple annotations of the same object are reduced, and the amount of annotation data is reduced, thereby improving the annotation efficiency and accuracy.
- identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image it also includes: detecting the number of pixels of at least one target recognition object in the sample image; and updating the target recognition object in the sample image according to the number of pixels of the at least one target recognition object.
- the number of pixels is used to characterize the amount of information about the target object.
- the number of pixels of a target object is too small, for example, if there is only one pixel, it is difficult to distinguish the target object from other targets, and it is also difficult to obtain effective information about the target object.
- the recognition significance of the target object is small, and even if it can be labeled, the labeled data cannot be provided. Effective information to train the model.
- Target objects with a pixel count less than the threshold are eliminated, while those with a pixel count greater than or equal to the threshold are retained.
- Different pixel count thresholds can be set for different categories.
- Updating the target recognition objects in the sample image according to the number of pixels of at least one target recognition object may be eliminating redundant target recognition objects in the sample image according to the number of pixels of at least one target recognition object. For example, target recognition objects with too few pixels are eliminated to achieve the updating of the target recognition objects in the sample image.
- the target recognition objects By screening the target recognition objects according to the number of pixels of the target recognition objects, the target recognition objects with more effective information and more key information can be retained, so as to increase the representativeness of the target recognition objects, increase the content richness and representativeness of the annotation data, and make the annotation data contain more effective information.
- the target identification objects are screened by the dimensions of redundancy and effectiveness. In addition, other dimensions can also be used to screen the target identification objects, which is not limited.
- the relative area ratio of the target recognition object and the sample image can be adjusted before generating sample data.
- identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image it also includes: detecting the size ratio of at least one target recognition object in the sample image to the sample image; when the size ratio of the target recognition object is less than a preset threshold, expanding the boundary of the target recognition object to increase the size ratio of the expanded target recognition object to the sample image.
- the size ratio of the target object is used to indicate its importance to the sample image. To extract richer and more complete data from the target object, the size ratio of the target object is typically kept within a certain range.
- the size ratio refers to the ratio of the area of the target object's detection box or image region to the sample image area. If the size ratio of the target object is less than the preset threshold, it indicates that the size ratio is too small and should generally be increased.
- the target object's detection box is rectangular. You can move the edges of the rectangle outward by a distance x.
- the expanded target object includes the original target object. Typically, the center of the target object remains unchanged. The target object can be enlarged proportionally.
- the target recognition unit corresponding to each word in the dynamic vocabulary is identified in the sample image.
- it also includes: detecting the size ratio of at least one target identification object in the sample image to the sample image; when the size ratio of the target identification object is less than a preset threshold, cropping the sample image and updating the sample image to increase the size ratio of the target identification object to the cropped sample image.
- the image can be cropped to avoid the location of at least one target object, while preserving the shape, angle, and position of the sample image.
- the target recognition object is recognized according to each word in the dynamic vocabulary to obtain the target recognition object, which can increase the diversity of identifiable objects, thereby increasing the diversity of annotation data and enriching the content of the annotation task.
- the object category can be updated in real time, and the recognized object can be updated accordingly, thereby updating the annotation data in real time, which can make the annotated data open and dynamically changeable, thereby improving the accuracy and real-time performance of the annotation data.
- Figure 4 is a flow chart of a method for training a visual semantic understanding model according to an embodiment of the present application.
- This embodiment can be applicable to the case where a visual semantic understanding model is trained based on the aforementioned generated annotation data in combination with sample images, and the visual semantic understanding model is used to input an image and output a description of the image content.
- the method of this embodiment can be executed by a training device for a visual semantic understanding model, which can be implemented in software and/or hardware and configured in an electronic device with a certain data computing capability, which can be a client device or a server device, and the client device can include: a personal computer, a laptop computer, a smart phone, a tablet computer, an Internet of Things device or a portable wearable device, etc.
- the labeled data is used as the true description content of the sample image, and the sample image is input into the visual semantic understanding model to train the visual semantic understanding model.
- the visual semantic understanding model can be a large language model.
- the accuracy of the descriptive text output by the visual semantic understanding model can be improved.
- the automatic generation of annotated data can reduce the model training cost and improve the model training efficiency.
- FIG. 5 is a scene diagram of a method for generating image annotation data according to an embodiment of the present application.
- This embodiment of the present application is mainly divided into two phases: Phase 1 and Phase 2.
- Phase 1 includes five small steps, such as generating sample data, which includes an image, a detection box in the image, and a category label.
- Phase 2 includes four small steps, such as generating target description information for the detection box, that is, implementing the annotation of the detection box.
- RAM++ is used to perform object detection on sample images, and category labels are added to the detected objects.
- Open-World Localization version 2 (OWLv2) is used to implement open vocabulary object detection (OVD).
- OWL Open-World Localization version 2
- the generated detection boxes are filtered and suppressed.
- This generates input for the visual semantic understanding model, which is used to generate captions.
- the input requires the object's detection box (BOX) and label (LABEL) (which can be a category label).
- BOX object's detection box
- LABEL label
- the image, the object in the image (i.e., the BOX and LABEL of the target object), and the label are input into the large language model to obtain the initial version of the target caption.
- the large language model can be a multimodal large language model (MLLM), exemplarily a Shikra model.
- an MLLM model typically includes a visual encoding layer, a fusion layer, a text encoding layer, and a decoding layer.
- the initial version of the target caption is decomposed into semantic terms to obtain corresponding nouns and adjectives, namely visual nouns and visual adjectives.
- the word meanings are parsed, and adjectives are divided into categories such as material, color, behavior, and location. Each adjective is double-checked to ensure the accuracy of the filtered adjectives.
- the nouns are matched with tagging information (i.e., category labels) to ensure the accuracy of the noun pairs.
- the generated parsed words are semantically concatenated and diversified to obtain the target caption.
- MLLM is used for verification to improve the ability to generate captions.
- a dynamic vocabulary for images through tagging Execute step 1-1 in Figure 5 to add tags (tagging).
- RAM++ crawls millions of visual data (images or images in videos) on the Internet.
- Automated labeling of visual data enables coarse-grained labeling.
- Image-level labels are added to visual data, resulting in a large number of image-text pairs.
- corresponding word information is extracted.
- This system can identify over 6,000 categories, as well as the categories and background within an image, and then add the extracted word information to a dynamic vocabulary.
- Tagging ensures the diversity of the image vocabulary, ensuring that multiple extracted images are distinct.
- a dynamic vocabulary can include: people, advertisements, cars, animals, scenes, food, furniture, and appliances. There are also many more diverse words, and there is no limit to this.
- Imaginative vocabulary obtained through a large language model We also extracted background categories from RAM++, such as park and kitchen, and used the large language model to generate the question "What objects might be included in the following scenes?" This ensures the integrity of the vocabulary and the diversity of the generated image vocabulary.
- OWLv2 is used to generate detection frames corresponding to the preset vocabulary. For example, by inputting the original image, i.e., the sample image, and combining it with the preset vocabulary, detection frames corresponding to the words in the preset vocabulary in the sample image can be generated. For example, in the sample image, object 1 and object 2 are identified, where the categories of object 1 and object 2 can be different or the same.
- the generated detection frames are input into the Filter module after executing steps 1-4.
- the Filter module performs frame suppression and outputs multiple sets of images, detection frames, and category labels based on steps 1-5.
- static categories are divided into at least one major category, such as at least one of food, tools, people, traffic scenes, and tools.
- the main suppression content is food and people. The operation is as follows:
- An image will generate at least one adjective-noun pair and the corresponding box position, and the adjective, noun and box will be combined to generate a list [ ⁇ adj.1>, ⁇ n.1>,box1 ⁇ , ⁇ adj.2>, ⁇ n.2>,box2 ⁇ ...].
- Box_new is the fused detection box.
- the center coordinates of the fused detection box are the minimum of the center coordinates of the two fused detection boxes.
- the width and height of the fused detection box are the maximum of the width and height of the two fused detection boxes.
- ii Merge the corresponding label content.
- the enhanced label refers to the label content with the richest semantics. You can perform semantic segmentation on the label to obtain noun-adjective pairs and retain the largest number of labels.
- the sample data includes a sample image, at least one detection box, and at least one class label for the detection box. Each detection box represents a target object.
- step 2-1 Execute step 2-1 and input the filtered box, label, and image into the MLLM to obtain the initial target description information caption.
- the caption has many problems, including at least one of the following: incorrect expression, incorrect color description, and incorrect object description.
- the above problems can be solved by combining semantic segmentation with semantic reorganization.
- the sample image is cropped into a new sample image of appropriate size based on the position and ratio of the detection frame. This ensures that the sample image used in the annotation process is the most informative, improving the accuracy of the annotation.
- the original sample image is used directly without cropping.
- the annotation filtering module makes a preliminary judgment on the generated visual description information to detect whether there is any inconsistency with the actual sample image content. If there is a difference, it will enter the processing flow of the annotation correction module.
- the annotation filtering module performs semantic division on the visual description to obtain nouns and adjectives.
- relations can also be treated as adjectives or separated from adjectives. Among them, nouns can refer to detection
- adjectives can be descriptive words, and relationships can refer to the relationship between the object and surrounding objects. Semantic segmentation combined with semantic reorganization can clean nouns and adjectives, and reorganize correct nouns and correct adjectives to obtain accurate visual description information, thus achieving clean visual description information.
- the annotation correction module is responsible for correcting nouns and adjectives accordingly. This module ensures that the final annotation results more accurately reflect the actual image content.
- the annotation correction module inputs the corrected terms into the MLLM, performs semantic reorganization of the description information, and obtains multiple corrected visual descriptions.
- the corrected visual descriptions are output to the annotation filtering module.
- the annotation filtering module filters the multiple corrected descriptions to obtain the final target description.
- the filtered target description is output as the final annotation result, completing the entire image annotation process.
- steps 2-4 directly use the visual description information in the previous stage as the target description information, and output it as the final annotation result to complete the entire image annotation process.
- the second phase from steps 2-1 to 2-4, is complete, generating target descriptions for the sample data.
- This systematic process improves the automation and accuracy of annotation.
- the resulting annotated data can be used for both MLLM and OWLv2 training.
- the technical solution of this application through automatic tagging technology, part-of-speech checking and box filtering, and making full use of the calibration process of the generative pre-trained transformer model (GPT), can effectively improve the accuracy of the tagging process and effectively improve the credibility of the automated tagging process.
- GPT generative pre-trained transformer model
- Figure 6 is a structural diagram of an apparatus for generating image annotation data in an embodiment of the present application.
- This embodiment of the present application is applicable to generating image descriptions.
- the apparatus is implemented using software and/or hardware and is configured in an electronic device with certain data processing capabilities.
- an image annotation data generating device 600 includes: a sample data acquisition module 601, a visual description verification module 602, an object description information adjustment module 603 and an annotation data generating module 604.
- the sample data acquisition module 601 is configured to acquire a sample image, a target object in the sample image, and visual description information of the target object;
- a visual description verification module 602 configured to verify the words in the visual description information based on the sample image and the target recognition object, and obtain a verification result of the visual description information
- the target description information adjustment module 603 is configured to adjust the words in the visual description information according to the verification result of the visual description information to obtain the target description information of the target recognition object;
- the annotation data generation module 604 is configured to use the target description information of the target recognition object as the annotation data of the target recognition object; the annotation data is used to train the visual semantic understanding model, and the visual semantic understanding model is used to output the description text of the visual content based on the input visual content.
- the visual description information of the target recognition object is automatically extracted, and the words in the visual description information are verified to obtain the verification result.
- the words in the visual description information are adjusted to obtain the target description information, which is used as the annotation data of the target recognition object and added to the sample data to train the image semantic understanding model to obtain a model for semantic understanding of the input visual content, thereby improving the generation efficiency of image annotation data, reducing the annotation cost, and taking into account the annotation accuracy, thereby quickly generating correct annotation data to train the model and improve the accuracy of the description text of the visual semantic understanding model.
- the visual description verification module 602 includes: a description segmentation unit, configured to divide the visual description information into at least one visual noun and at least one visual adjective; a noun verification unit, configured to verify the at least one visual noun based on the sample image and the target recognition object, and obtain a noun verification result corresponding to the at least one visual noun; an adjective verification unit, configured to verify the at least one visual adjective based on the sample image and the target recognition object, and obtain an adjective verification result corresponding to the at least one visual adjective; and a verification result determination unit, configured to determine the noun verification result and the adjective verification result as the verification result of the visual description information.
- the noun verification unit includes: a category acquisition subunit, configured to obtain category description information of the target recognition object based on the sample image and the target recognition object; a category verification subunit, configured to detect the consistency between the category description information of the target recognition object and the visual noun for at least one visual noun, and obtain a noun verification result corresponding to the visual noun.
- the adjective verification unit includes: an interaction generation subunit, configured to generate interactive dialogue content for the at least one visual adjective based on the visual adjective and the visual noun corresponding to the visual adjective; an adjective semantic understanding subunit, configured to use the sample image and the target recognition object as context, and input the interactive dialogue content into a large language model to obtain an existence judgment result of the visual adjective; and an adjective verification subunit, configured to generate an adjective verification result of the visual adjective based on the existence judgment result of the visual adjective.
- the target description information adjustment module 603 includes: a noun screening unit, configured to screen the at least one visual noun based on the verification result of each noun to obtain at least one target noun; an adjective screening unit, configured to screen the at least one visual adjective based on the verification result of each adjective to obtain at least one target adjective; a description phrase generation unit, configured to generate at least one description phrase based on the at least one target noun and the at least one target adjective; and a description information generation unit, configured to generate target description information of the target recognition object based on the at least one description phrase.
- the image annotation data generating device further includes: a real-time category acquisition module, configured to acquire real-time updated object categories and add the acquired real-time updated object categories to a dynamic vocabulary; an image acquisition module, configured to acquire a sample image; and an object update acquisition module, configured to identify, in the sample image, a target recognition object corresponding to each word in the dynamic vocabulary, and obtain the target recognition object in the sample image.
- a real-time category acquisition module configured to acquire real-time updated object categories and add the acquired real-time updated object categories to a dynamic vocabulary
- an image acquisition module configured to acquire a sample image
- an object update acquisition module configured to identify, in the sample image, a target recognition object corresponding to each word in the dynamic vocabulary, and obtain the target recognition object in the sample image.
- the real-time category acquisition module includes: a visual set acquisition unit, configured to acquire a visual set updated in real time; a target detection unit, configured to perform target detection on the visual data in the visual set to obtain at least one first object category; and/or a semantic understanding unit, configured to perform semantic understanding on the visual data in the visual set to obtain at least one second object category; and a category update unit, configured to determine the first object category and/or the second object category as the real-time updated object category.
- the image annotation data generating device further includes: an intersection-over-union (IoU) detection module, configured to detect the IoU between multiple target recognition objects in the sample image after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image; a similar object detection module, configured to detect the multiple target recognition objects based on the IoU between the multiple target recognition objects to obtain similar multiple target recognition objects; an object fusion module, configured to fuse the similar multiple target recognition objects to obtain a fusion result; and a redundancy screening module, configured to update the multiple target recognition objects in the sample image based on the fusion result.
- IoU intersection-over-union
- the image annotation data generating device further includes: an intersection-over-union detection module, configured to detect the number of pixels of at least one target recognition object in the sample image after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image; and a pixel number screening module, configured to update the target recognition object in the sample image based on the number of pixels of the at least one target recognition object.
- the image annotation data generation device further includes: a size ratio detection module configured to detect the size ratio of at least one target recognition object in the sample image to the sample image after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image; an object boundary expansion module configured to detect the size ratio of at least one target recognition object in the sample image to the sample image after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image; When a threshold is preset, the boundary of the target recognition object is expanded to increase the size ratio of the expanded target recognition object to the sample image.
- the image annotation data generating device further includes: a size ratio detection module, configured to detect the size ratio of at least one target recognition object in the sample image and the sample image after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image; an image cropping module, configured to crop the sample image when the size ratio of the target recognition object is less than a preset threshold, and update the sample image to increase the size ratio of the target recognition object and the cropped sample image.
- a size ratio detection module configured to detect the size ratio of at least one target recognition object in the sample image and the sample image after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image
- an image cropping module configured to crop the sample image when the size ratio of the target recognition object is less than a preset threshold, and update the sample image to increase the size ratio of the target recognition object and the cropped sample image.
- the above-mentioned image annotation data generation device can execute the image annotation data generation method provided by any embodiment of the present application, and has the corresponding functional modules and effects of executing the image annotation data generation method.
- FIG7 is a structural diagram of a training device for a visual semantic understanding model in an embodiment of the present application.
- the embodiment of the present application is applicable to training a visual semantic understanding model.
- the device is implemented using software and/or hardware and is configured in an electronic device with certain data processing capabilities.
- a visual semantic understanding model training device 700 includes: a sample data acquisition module 701 and a model training module 702 .
- a sample data acquisition module 701 is configured to acquire sample data, where the sample data includes a sample image and annotation data, where the annotation data is generated by the image annotation data generation method described in any embodiment of the present application;
- the model training module 702 is configured to use the sample data to train a visual semantic understanding model.
- the accuracy of the descriptive text output by the visual semantic understanding model can be improved.
- the automatic generation of annotated data can reduce the model training cost and improve the model training efficiency.
- the above-mentioned training device for the visual semantic understanding model can execute the training method for the visual semantic understanding model provided in any embodiment of the present application, and has the corresponding functional modules and effects for executing the training method for the visual semantic understanding model.
- the present application also provides an electronic device, a readable storage medium and a computer program product.
- FIG8 shows a schematic area diagram of an example electronic device 800 that can be used to implement an embodiment of the present application.
- the electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.
- the electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices.
- the components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and/or required herein.
- electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803.
- RAM 803 can also store various programs and data required for the operation of electronic device 800.
- Computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804.
- An input/output (I/O) interface 805 is also connected to bus 804.
- the I/O interface 805 includes an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc.
- the communication unit 809 allows the electronic device 800 to exchange information/data with other devices via a computer network such as the Internet and/or various telecommunication networks.
- the computing unit 801 can be a variety of general and/or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), a variety of dedicated artificial intelligence (AI) computing chips, a variety of computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc.
- the computing unit 801 performs the multiple methods and processes described above, such as an image annotation data generation method or a visual semantic understanding model training method.
- the image annotation data generation method or the visual semantic understanding model training method can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 808.
- part or all of the computer program can be loaded and/or installed on the electronic device 800 via the ROM 802 and/or the communication unit 809.
- the computing unit 801 may be configured to perform the method for generating image annotation data or the method for training a visual semantic understanding model by any other appropriate means (e.g., by means of firmware).
- Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and/or combinations thereof.
- FPGAs field programmable gate arrays
- ASICs application specific integrated circuits
- ASSPs application specific standard parts
- SOCs system on chips
- CPLDs complex programmable logic devices
- computer hardware firmware, software, and/or combinations thereof.
- the program code for implementing the method of the present application can be written in any combination of one or more programming languages.
- Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions/instructions specified in the flow chart and/or area diagram are implemented.
- the program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
- a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.
- a machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
- a machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.
- Machine-readable storage media include electrical connections based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, Erasable Programmable Read-Only Memory (EPROM), flash memory, optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
- the storage medium may be a non-transitory storage medium.
- the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer.
- a display device e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor
- a keyboard and pointing device e.g., a mouse or trackball
- Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
- the systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components.
- the components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: Local Area Networks (LANs), Wide Area Networks (WANs), blockchain networks, and the Internet.
- a computer system may include a client and a server.
- the client and server are generally remote from each other and typically interact through a communication network.
- the client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other.
- the server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system. It solves the management difficulties and poor business scalability of traditional physical hosts and virtual private servers (VPS) services.
- the server may also be a server in a distributed system or a server integrated with blockchain.
- AI Artificial intelligence
- AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing.
- AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning/deep learning, big data processing, and knowledge graphs.
- Cloud computing refers to a technology system that provides network access to elastically scalable shared pools of physical or virtual resources. These resources can include servers, instruction sets, networks, software, applications, and storage devices, and can be deployed and managed on-demand in a self-service manner. Cloud computing technology provides efficient and powerful data processing capabilities for the application of technologies such as artificial intelligence and blockchain, as well as for model training.
Landscapes
- Engineering & Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Multimedia (AREA)
- Theoretical Computer Science (AREA)
- Artificial Intelligence (AREA)
- Image Analysis (AREA)
- Image Processing (AREA)
Abstract
一种图像标注数据生成、模型训练方法、装置、设备及介质。图像标注数据生成方法包括:获取样本图像、所述样本图像中目标识别对象和所述目标识别对象的视觉描述信息(S101);根据所述样本图像和所述目标识别对象,对所述视觉描述信息中词语进行校验,得到所述视觉描述信息的校验结果(S102);根据所述视觉描述信息的校验结果,对所述视觉描述信息中词语进行调整,得到所述目标识别对象的目标描述信息(S103);将所述目标识别对象的目标描述信息,作为所述目标识别对象的标注数据(S104)。
Description
本申请要求在2024年03月08日提交中国专利局、申请号为202410268967.1的中国专利申请的优先权,该申请的全部内容通过引用结合在本申请中。
本申请涉及图像处理领域,涉及人工智能、计算机视觉、大语言模型和智慧城市领域,例如涉及一种图像标注数据生成、模型训练方法、装置、设备及介质。
随着图像数量的剧增,人们迫切地需要实现图像内容的高效标注,以实现大规模图像的有效检索与管理。
随着人工智能技术的不断进步,大型模型的问世引起了行业的普遍关注。在人工智能新浪潮的生态架构中,算力基础设施层的确定机遇以及大型模型的应用革新为行业带来了全新的发展机遇。
发明内容
本申请提供了一种图像标注数据生成、模型训练方法、装置、设备及介质。
根据本申请的一方面,提供了一种图像标注数据生成方法,包括:
获取样本图像、所述样本图像中目标识别对象和所述目标识别对象的视觉描述信息;
根据所述样本图像和所述目标识别对象,对所述视觉描述信息中词语进行校验,得到所述视觉描述信息的校验结果;
根据所述视觉描述信息的校验结果,对所述视觉描述信息中词语进行调整,得到所述目标识别对象的目标描述信息;
将所述目标识别对象的目标描述信息,作为所述目标识别对象的标注数据;其中,所述标注数据用于训练视觉语义理解模型,所述视觉语义理解模型用于根据输入的视觉内容,输出所述视觉内容的描述文本。
根据本申请的一方面,提供了一种视觉语义理解模型的训练方法,包括:
获取样本数据,其中,所述样本数据包括样本图像和标注数据,所述标
注数据通过如本申请任一实施例所述的图像标注数据生成方法生成;
采用所述样本数据,训练视觉语义理解模型。
根据本申请的一方面,提供了一种图像标注数据生成装置,包括:
样本数据获取模块,设置为获取样本图像、所述样本图像中目标识别对象和所述目标识别对象的视觉描述信息;
视觉描述校验模块,设置为根据所述样本图像和所述目标识别对象,对所述视觉描述信息中词语进行校验,得到所述视觉描述信息的校验结果;
目标描述信息调整模块,设置为根据所述视觉描述信息的校验结果,对所述视觉描述信息中词语进行调整,得到所述目标识别对象的目标描述信息;
标注数据生成模块,设置为将所述目标识别对象的目标描述信息,作为所述目标识别对象的标注数据;其中,所述标注数据用于训练视觉语义理解模型,所述视觉语义理解模型用于根据输入的视觉内容,输出所述视觉内容的描述文本。
根据本申请的一方面,提供了一种视觉语义理解模型的训练装置,包括:
样本数据获取模块,设置为获取样本数据,其中,所述样本数据包括样本图像和标注数据,所述标注数据通过如本申请任一实施例所述的图像标注数据生成方法生成;
模型训练模块,设置为采用所述样本数据,训练视觉语义理解模型。
根据本申请的另一方面,提供了一种电子设备,包括:
至少一个处理器;以及
与所述至少一个处理器通信连接的存储器;其中,
所述存储器存储有可被所述至少一个处理器执行的指令,所述指令被所述至少一个处理器执行,以使所述至少一个处理器能够执行本申请任一实施例所述的图像标注数据生成方法,或本申请任一实施例所述的视觉语义理解模型的训练方法。
根据本申请的另一方面,提供了一种存储有计算机指令的非瞬时计算机可读存储介质,其中,所述计算机指令用于使所述计算机执行本申请任一实施例所述的图像标注数据生成方法,或本申请任一实施例所述的视觉语义理解模型的训练方法。
根据本申请的另一方面,提供了一种计算机程序产品,包括计算机程序,所述计算机程序在被处理器执行时实现本申请任一实施例所述的图像标注数
据生成方法,或本申请任一实施例所述的视觉语义理解模型的训练方法。
图1是根据本申请实施例提供的一种图像标注数据生成方法的流程图;
图2是根据本申请实施例提供的另一种图像标注数据生成方法的流程图;
图3是根据本申请实施例提供的另一种图像标注数据生成方法的流程图;
图4是根据本申请实施例提供的一种视觉语义理解模型的训练方法的流程图;
图5是根据本申请实施例提供的一种图像标注数据生成方法的场景图;
图6是根据本申请实施例提供的图像标注数据生成装置的结构示意图;
图7是根据本申请实施例提供的视觉语义理解模型的训练装置的结构示意图;
图8是根据本申请实施例提供的一种电子设备的框图。
以下结合附图对本申请的示范性实施例做出说明,其中包括本申请实施例的多种细节以助于理解,应当将它们认为仅仅是示范性的。因此,可以对这里描述的实施例做出多种改变和修改。同样,为了清楚和简明,以下的描述中省略了对公知功能和结构的描述。
图1是根据本申请实施例提供的一种图像标注数据生成方法的流程图,本实施例可以适用于生成图像的描述内容的情况。本实施例方法可以由图像标注数据生成装置来执行,该装置可采用软件和/或硬件的方式实现,并可配置于具有一定数据运算能力的电子设备中,该电子设备可以是客户端设备或服务器设备,客户端设备可以包括:个人计算机、笔记本电脑、智能手机、平板电脑、物联网设备或便携式可穿戴设备等,物联网设备可为智能音箱、智能电视、智能空调或智能车载设备等。便携式可穿戴设备可为智能手表、智能手环或头戴设备等。
S101、获取样本图像、所述样本图像中目标识别对象和所述目标识别对象的视觉描述信息。
样本图像包括至少一个目标识别对象。若目标识别对象的数量为多个,不同目标识别对象可以对应不同对象,和/或不同目标识别对象对应不同类型的对象。示例性的,样本图像包括人的目标识别对象、猫的目标识别对象和
桌子的目标识别对象等。又如,样本图像包括至少一个猫的目标识别对象。可以对样本图像进行目标检测,得到样本图像中的对象的目标检测框,将一个目标检测框确定为一个目标识别对象。或者还可以对样本图像进行分割,得到样本图像中分割区域,将一个对象的区域确定为一个目标识别对象,背景区域不包括对象。此外,还可以在得到目标识别对象的同时,对目标识别对象进行分类,得到目标识别对象的类别。
视觉描述信息用于描述目标识别对象的图像内容,表征目标识别对象对应的图像的语义信息。视觉描述信息可以是指对视觉可见内容进行描述的语句。视觉描述信息可以包括下述至少一项:目标对象内容、主题内容和关系内容等。其中,主题内容可以是指样本图像的关键信息。示例性的,主题内容可以包括场景和/或事件等。目标对象内容可以是指目标对象本身的内容,例如,目标对象可以包括下述至少一项:尺寸、状态、位置和颜色等。关系内容可以是指目标识别对象的目标对象与图像中其他事物之前的关系,例如,其他事物可以包括目标识别对象的事物和/或样本图像中目标识别对象以外的事物等。关系可以包括位置关系、动静关系和状态关系等中的至少一项。此外,视觉描述信息还可以包括其他内容,对此不限定。示例性的,样本图像为车辆在行驶过程中所采集的周围环境的图像;目标识别对象可以是行人、车辆、标志线、指示灯和护栏等,例如目标识别对象为行人,相应的,视觉描述信息可以是:正前方存在行人。
可以将样本图像和目标识别对象输入到预先训练的生成模型中,得到视觉描述信息。示例性的,生成模型的处理过程为:对样本图像和目标识别对象分别进行编码,得到样本图像和目标识别对象分别对应的向量表示;获取提示词文本,并进行编码得到提示词文本的向量表示,将两个向量表示进行融合,得到融合结果,对融合结果进行解码,得到视觉描述信息。此外,还可以将样本图像、目标识别对象和目标识别对象的类别,输入到生成模型中,得到视觉描述信息。生成模型可以是对视觉内容进行语义理解,并生成描述信息。
S102、根据所述样本图像和所述目标识别对象,对所述视觉描述信息中词语进行校验,得到所述视觉描述信息的校验结果。
对视觉描述信息进行分词,得到至少一个词语。对每个词语进行校验,检测每个词语是否正确。可以从样本图像和目标识别对象提取图像信息,对视觉描述信息中词语进行校验。校验可以是根据目标识别对象和样本图像,检测词语是否正确或者是否存在等。视觉描述信息的校验结果可以是指视觉描述信息中每个词语是否正确,将正确的词语保留,将错误的词语剔除。
S103、根据所述视觉描述信息的校验结果,对所述视觉描述信息中词语进行调整,得到所述目标识别对象的目标描述信息。
视觉描述信息可以是指语句,相应的,对视觉描述信息中词语进行调整,可以是保留正确的词语,修改错误的词语,得到新的语句,确定为目标识别对象的目标描述信息。此外,还可以获取新的语句中至少一个词语的近义词,并对至少一个词语以及近义词进行重组,形成更多的新的语句,对多个新的语句进行筛选,得到更加准确更加规范的语句,确定为目标描述信息。
S104、将所述目标识别对象的目标描述信息,作为所述目标识别对象的标注数据;所述标注数据用于训练视觉语义理解模型,所述视觉语义理解模型用于根据输入的视觉内容,输出所述视觉内容的描述文本。
将目标描述信息作为标注数据,添加到样本数据中,可以理解为,目标描述信息是目标识别对象的描述信息的真值。采用样本图像以及样本图像的目标描述信息,训练视觉语义理解模型,这样训练得到的视觉语义理解模型可以对输入的图像进行语义理解,输出该图像的描述文本,作为该图像的语义理解内容。
实际上,样本图像中可以包括至少一个目标识别对象,每个目标识别对象生成目标描述信息,可以将样本图像以及该样本图像中至少一个目标识别对象的目标描述信息,作为训练数据,训练视觉语义理解模型。
在一个实施例中,采用人工标注图像,虽然可以为图像带来最佳的标签,但需要花费大量的人力和资源,而大模型所需要的数据量庞大,一些大型企业花费了上千人力共同进行多个月才标注完全部需求,带来了昂贵的标注费用,一般在面对新业务需求时,很多企业面临着数据以及大量数据标注的困境,例如在目前的大模型时代。或者还可以采用机器学习和深度学习的方法来进行相关领域(检测和分割等)自动标注,但这种方式带有很强烈地不确定性和不准确度。
根据本申请的技术方案,通过自动提取目标识别对象的视觉描述信息,并对视觉描述信息中词语进行校验,得到校验结果,同时根据校验结果,对视觉描述信息中词语进行调整,得到目标描述信息,作为目标识别对象的标注数据,并据此训练图像语义理解模型,以得到对输入的视觉内容进行语义理解的模型,提高图像标注数据的生成效率,降低标注成本,同时兼顾标注准确性,进而快速生成正确标注数据,以训练模型,提高视觉语义理解模型输出的描述文本的准确性。
图2是根据本申请实施例提供的另一种图像标注数据生成方法的流程图,基于上述技术方案进行说明,并可以与上述多个可选实施方式进行结合。所述根据所述样本图像和所述目标识别对象,对所述视觉描述信息中词语进行校验,得到所述视觉描述信息的校验结果,包括:将视觉描述信息划分为至少一个视觉名词和至少一个视觉形容词;根据所述样本图像和所述目标识别对象,对所述至少一个视觉名词进行校验,得到所述至少一个视觉名词对应的名词校验结果;根据所述样本图像和所述目标识别对象,对所述至少一个视觉形容词进行校验,得到所述至少一个视觉形容词对应的形容词校验结果;将所述名词校验结果和所述形容词校验结果,确定为所述视觉描述信息的校验结果。
S201、获取样本图像、所述样本图像中目标识别对象和所述目标识别对象的视觉描述信息。
S202、将视觉描述信息划分为至少一个视觉名词和至少一个视觉形容词。
视觉名词可以是指视觉可见的名词,与目标识别对象有关。视觉形容词可以是指视觉可见的形容词,与视觉名词有关。视觉名词用于描述目标识别对象的类别。视觉形容词用于对目标识别对象进行限定。视觉形容词可以划分为材质、颜色或位置等。示例性的,视觉描述信息为白色布偶猫,视觉名词可以包括猫,视觉形容词可以包括白色和布偶。又如,视觉描述信息为黑狗趴在地毯上。视觉名词包括:狗和地毯。视觉形容词可以包括黑色和狗在地毯上。
示例性的,视觉描述信息可以划分为至少一个名词形容词对,例如,视觉描述信息为a red apple,提取出名词形容词对为{<red>,<apple>}。又如,黑狗趴在地毯上,提取出名词形容词对为{<黑>,<狗>}、{<趴>,<狗>}、{<上>,<狗>}和{<下>,<地毯>}。该视觉描述信息与目标识别对象对应。此外,一个样本图像可以存在多个目标识别对象,例如,对象1的视觉描述信息为:白色布偶猫,对象2的视觉描述信息为黑狗趴在地毯上。相应的,一个样本图像中可以形成的名词形容词对包括:{<白>,<猫>,<对象1>}、{<布偶>,<猫>,<对象1>}、{<黑>,<狗>,<对象2>}、{<趴>,<狗>,<对象2>}、{<上>,<狗>,<对象2>}和{<下>,<地毯>,<对象2>}。
S203、根据所述样本图像和所述目标识别对象,对所述至少一个视觉名词进行校验,得到所述至少一个视觉名词对应的名词校验结果。
可以根据样本图像和目标识别对象,获取目标识别对象的类别,对视觉名词进行校验。例如可以是检测目标识别对象的类别是否与视觉名词相似。
名词校验结果可以是指视觉名词是否正确的校验结果。一个视觉名词对应存在一个名词校验结果。
可选的,所述根据所述样本图像和所述目标识别对象,对所述至少一个视觉名词进行校验,得到所述至少一个视觉名词对应的名词校验结果,包括:获取所述目标识别对象的类别描述信息;针对所述至少一个视觉名词,检测所述目标识别对象的类别描述信息与所述至少一个视觉名词的一致性,得到所述至少一个视觉名词对应的名词校验结果。
类别描述信息可以是在样本图像中检测对象时,检测得到目标识别对象,以及目标识别对象的类别,将目标识别对象的类别,确定为类别描述信息。例如,在样本图像中检测目标识别对象的检测框,以及该检测框的类别,将该检测框对应的图像区域确定为目标识别对象;将该检测框的类别,确定为目标识别对象的类别描述信息。可以将类别描述信息与目标识别对象进行对应存储,以便后续获取目标识别对象的类别描述信息。
将类别描述信息与视觉名词进行比较,若一致,确定该视觉名词的名词校验结果为正确;若不一致,确定该视觉名词的名词校验结果为错误。此外,还可以计算类别描述信息与视觉名词之间的相似度,将相似度确定为名词校验结果,或者在相似度大于或等于相应相似阈值时,确定该视觉名词的名词校验结果为正确;在相似度小于相应相似阈值时,确定该视觉名词的名词校验结果为错误。
通过将目标识别对象的类别描述信息与视觉名词进行一致性比较,得到比较结果,并根据比较结果,确定视觉名词的名词校验结果,可以根据样本图像在检测对象时获取的类别对视觉名词进行校验,避免通过额外的复杂冗余步骤对视觉名词进行校验,简化视觉名词的校验操作,提高视觉名词的校验效率,同时,目标识别对象的类别的准确性较高,采用目标识别对象的类别进行校验,提高视觉名词的校验准确性。
S204、根据所述样本图像和所述目标识别对象,对所述至少一个视觉形容词进行校验,得到所述至少一个视觉形容词对应的形容词校验结果。
视觉形容词通常是针对一个视觉名词的限定内容和描述内容。对视觉形容词进行校验,通常是检测该视觉形容词所描述的视觉名词具有的特质与该视觉形容词是否一致,若一致,确定该视觉形容词正确;若不一致,确定该视觉形容词错误。一个视觉形容词对应存在一个形容词校验结果。
可选的,所述根据所述样本图像和所述目标识别对象,对所述至少一个视觉形容词进行校验,得到所述至少一个视觉形容词对应的形容词校验结果,
包括:针对所述至少一个视觉形容词,根据所述至少一个视觉形容词,以及所述至少一个视觉形容词对应的视觉名词,生成交互对话内容;将所述样本图像和所述目标识别对象作为上下文,以及将所述交互对话内容输入到大语言模型中,得到所述至少一个视觉形容词的存在判断结果;根据所述至少一个视觉形容词的存在判断结果,生成所述至少一个视觉形容词的形容词校验结果。
交互对话内容用于检测视觉名词的描述内容是否存在与对应的视觉形容词对应的内容。交互对话内容通常是正确性判断语句。交互对话内容包括视觉形容词、该视觉形容词对应的视觉名词以及该视觉形容词对应的视觉名词是否与视觉形容词关联的判断内容。示例性的,判断内容包括:判断词或谓语;视觉名词与视觉形容词关联可以是指视觉名词具有视觉形容词对应的性质,例如,视觉名词为猫,视觉形容词为白色,样本图像中猫的颜色是白色,即该视觉名词具有视觉形容词所表示的颜色性质,确定该视觉名词与该视觉形容词关联。视觉名词与视觉形容词不关联,可以是指视觉名词不具有视觉形容词对应的性质,一方面,视觉形容词所表示的性质与视觉名词具有的性质不同,如前例,样本图像中猫的颜色为黑色,即该视觉名词不具有视觉形容词所表示的颜色性质,确定该视觉名词与该视觉形容词不关联;另一方面,样本图像中,视觉名词具有的性质与视觉形容词所表示的性质无关,如前例,视觉名词为猫,视觉形容词为猫眼睛为蓝色,样本图像中猫是背影,无法确定视觉名词有关眼睛的颜色,确定该视觉名词与该视觉形容词不关联。交互对话内容用于输入到大语言模型中,依赖大语言模型对样本图像中目标识别对象的语义理解内容,判断交互对话内容是否正确,大语言模型输出的是交互对话内容是否正确的结果。视觉形容词的存在判断结果可以是指视觉形容词对应的视觉名词是否存在该视觉形容词的性质。
将样本图像和目标识别对象作为上下文输入到大语言模型中,可以是指大语言模型是基于样本图像和目标识别对象的内容,对交互对话内容作出回复,并输出视觉形容词的存在判断结果。可以是将样本图像作为上下文输入到大语言模型中。可以将样本图像和目标识别对象作为上下文输入到大语言模型中,以及将交互对话内容输入到大语言模型中,大语言模型根据样本图像和目标识别对象的语义理解内容,对交互对话内容进行正确性判断,得到视觉形容词的存在判断结果。样本图像和目标识别对象可以和交互对话内容同时输入到大语言模型中,也可以进行多轮对话,先将样本图像和目标识别对象输入到大语言模型中进行对话,再将交互对话内容输入到大语言模型中,即多轮对话中,在先的对话包括样本图像和目标识别对象,在后的对话包括交互对话内容。其中,大语言模型可以是前述的生成模型,用于输入样本图
像和目标识别对象,输出视觉描述信息,在输入样本图像和目标识别对象之后,即进行下轮对话,输入交互对话内容,输出交互对话内容的判断结果,即视觉形容词的存在判断结果。
示例性的,视觉描述信息为:白色布偶猫。视觉形容词为白色,对应的视觉名词为猫。交互对话内容为:猫是白色吗?又如,交互对话内容为:是否存在白色的猫?将样本图像和目标识别对象作为上下文,若大语言模型针对前述交互对话内容输出的结果为:是的,此时,视觉形容词的存在判断结果为存在,相应的该视觉形容词的形容词校验结果为正确;若大语言模型针对前述交互对话内容输出的结果为:不存在或不是,此时,视觉形容词的存在判断结果为不存在,相应的该视觉形容词的形容词校验结果为错误。
一个视觉名词可以对应有多个视觉形容词,一个视觉形容词对应一个视觉名词。针对多个视觉形容词,对每个视觉形容词进行存在判断结果的获取,即每个视觉形容词生成相应的交互对话内容,经过大语言模型得到存在判断结果。若视觉形容词的校验结果为错误,仅会对该视觉形容词进行剔除或修正,由于可能存在其他视觉形容词与该视觉形容词对应同一视觉名词,因而对该视觉形容词对应的视觉名词不能直接进行剔除或修正,需要对该对应的视觉名词以其他校验方式进行校验。
通过将视觉形容词与对应的视觉名词结合,生成用于判断是否存在视觉形容词的交互对话内容,并以样本图像和目标识别对象作为上下文,将交互对话内容输入到大语言模型中,用于检测视觉名词是否存在该视觉形容词的性质,从而判断视觉形容词是否存在,进而确定视觉形容词的校验结果,可以根据样本图像本身的视觉内容进行判断,避免通过额外的复杂冗余步骤对视觉形容词进行校验,简化视觉形容词的校验操作,提高视觉形容词的校验效率,同时,将视觉名词与视觉形容词进行关联,并对视觉名词和视觉形容词进行关联判断,充分利用视觉形容词不能单独存在的特性,既校验视觉形容词是否存在,又校验视觉名词是否与视觉形容词关联,提高视觉名词的校验准确性。
S205、将所述名词校验结果和所述形容词校验结果,确定为所述视觉描述信息的校验结果。
将全部名词校验结果和全部形容词校验结果,确定为视觉描述信息的校验结果。
S206、根据所述视觉描述信息的校验结果,对所述视觉描述信息中词语进行调整,得到所述目标识别对象的目标描述信息。
将错误的校验结果对应的词语进行修正,或者删除。针对正确的校验结果对应的词语,获取相似词语。对正确的词语、相似的词语以及修正的词语进行重组,得到新的视觉描述信息。对多个描述信息进行筛选,得到目标识别对象的目标描述信息。
可选的,所述根据所述视觉描述信息的校验结果,对所述视觉描述信息中词语进行调整,得到所述目标识别对象的目标描述信息,包括:根据每个名词校验结果,在所述至少一个视觉名词中筛选,得到至少一个目标名词;根据每个形容词校验结果,在所述至少一个视觉形容词中筛选,得到至少一个目标形容词;根据所述至少一个目标名词和所述至少一个目标形容词,生成至少一个描述词组;根据所述至少一个描述词组,生成目标识别对象的目标描述信息。
在至少一个视觉名词中筛选名词校验结果为正确的词语,确定为目标名词。在至少一个视觉形容词中筛选形容词校验结果为正确的词语,确定为目标形容词。一个描述词组包括一个名词和一个形容词。可以针对每个目标形容词,将目标形容词与所限定的目标名词进行组合,得到一个描述词组。或者,在前述例子中,在对视觉形容词进行校验时,将一个视觉形容词与该视觉形容词对应的视觉名词形成词组,按照目标名词和目标形容词对词组进行筛选,确定描述词组。描述词组中目标名词和目标形容词对应,即目标形容词是对同一描述词组中的目标名词进行限定和描述。根据描述词组中的名词和形容词进行造句,形成语句,确定为目标识别对象的目标描述信息。
一个目标识别对象可以生成至少一个目标描述信息,可以将全部目标描述信息作为目标识别对象的标注数据,也可以对目标描述信息进行筛选,例如基于语法、规范和表达丰富性等方面进行评分筛选,得到一个目标描述信息作为目标识别对象的标注数据。
此外,还可以针对至少一个目标名词和至少一个目标形容词,添加相应近义词,作为目标名词或目标形容词,并基于添加后的目标名词和目标形容词,进行排列组合,形成描述词组,可以丰富描述词组的内容,从而丰富目标描述信息的内容,增加目标描述内容的代表性。
通过根据词语的校验结果,对视觉名词和视觉形容词分别进行筛选,得到目标名词和目标形容词,并组合生成描述词组,根据描述词组中词语,生成目标识别对象的目标描述信息,可以基于不同词性的词语分别进行筛选,实现精准校验,并对筛选保留的词语进行对应组合,即对名词和形容词的关联性进行筛选,实现关联校验,根据校验后的描述词组生成目标描述信息,提高目标描述信息的准确性。
S207、将所述目标识别对象的目标描述信息,作为所述目标识别对象的标注数据;所述标注数据用于训练视觉语义理解模型,所述视觉语义理解模型用于根据输入的视觉内容,输出所述视觉内容的描述文本。
根据本申请的技术方案,通过对视觉描述信息进行分词,得到视觉名词和视觉形容词,并分别对视觉名词和视觉形容词进行校验,得到视觉描述信息的校验结果,可以对视觉描述信息进行拆解,以词语为单元进行校验,以更细粒度对描述信息进行校验,同时针对不同词性的词语分别进行校验,实现针对性的描述词语校验,提高描述信息的校验准确性。
图3是根据本申请实施例提供的另一种图像标注数据生成方法的流程图,基于上述技术方案进行说明,并可以与上述多个可选实施方式进行结合。所述样本图像和目标识别对象,通过以下方式得到:获取实时更新的对象类别,并将获取的实时更新的对象类别添加到动态词表中;获取样本图像;在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象,得到所述样本图像中目标识别对象。
S301、获取样本图像、所述样本图像中目标识别对象和所述目标识别对象的视觉描述信息。
S302、根据所述样本图像和所述目标识别对象,对所述视觉描述信息中词语进行校验,得到所述视觉描述信息的校验结果。
S303、根据所述视觉描述信息的校验结果,对所述视觉描述信息中词语进行调整,得到所述目标识别对象的目标描述信息。
S304、将所述目标识别对象的目标描述信息,作为所述目标识别对象的标注数据;所述标注数据用于训练视觉语义理解模型,所述视觉语义理解模型用于根据输入的视觉内容,输出所述视觉内容的描述文本。
S305、所述样本图像和所述目标识别对象通过以下方式得到:获取实时更新的对象类别,并将实时更新的对象类别添加到动态词表中。
对象类别可以是图像中可识别到的对象的类别。实时更新的对象类别是指,可以从网络热点或实时事件中采集的可图像识别的对象的类别。动态词表用于存储对象类别。示例性的,可以周期性获取实时图像或热点图像等,并从中识别出对象,将对象的对象类别添加到动态词表中。更新得到样本图像和目标识别对象,与生成标注数据时需要获取样本图像和目标识别对象的步骤可以相互独立。可以周期性更新得到样本图像和目标识别对象,并存储大量的样本图像和目标识别对象。将样本图像和样本图像中识别的至少一个
目标识别对象,作为一条样本数据;或者将样本图像和一个目标识别对象,作为一条样本数据,若样本图像中识别出多个目标识别对象,相应生成多条样本数据。在需要生成标注数据时,从预先存储的大量样本数据中获取至少一条样本数据,例如,可以获取最新的多条样本数据。
也可以获取热点词语或者实时词语,作为对象类别,但这样获取的词语存在无法可视化的情况,这样的词语难以作为对象类别,并且最新的可识别对象的类别的影响较小。由此,可以选择实时图像或热点图像,并对这些最新的图像进行检测,得到新的对象类别,添加到动态词表,实现精准对象类别的更新。
此外,还可以设置静态词表,例如,可以通过提取OBJECT365数据集(通用物体检测数据集)、上下文通用对象(Common Objects in Context,COCO)数据集、OPENIMAGES数据集(开放大型图像标注数据集)中的类别,并结合业务需求,收集类别添加到静态词表,但静态词表对图片不具有开放性,即在后期新的业务需求中,没有办法提取到更多元的框的信息,于是,通过标记(Tagging)能力去提取每一张图片的特有的词语,添加到动态词表中。采用静态词表和动态词表结合,极大丰富了对象类别的内容。
可选的,所述获取实时更新的对象类别,包括:获取实时更新的视觉集;对所述视觉集中视觉数据进行目标检测,得到至少一个第一对象类别;和/或对所述视觉集中视觉数据进行语义理解,得到至少一个第二对象类别;将所述第一对象类别和/或所述第二对象类别,确定为实时更新的对象类别。
视觉集可以是指图像或视频的集合。视觉数据可以包括采集图像或一帧视频图像。第一对象类别可以是最新的视觉数据中目标检测得到的对象类别。对视觉数据进行目标检测,可以得到至少一个第一对象,以及至少一个第一对象的第一对象类别。
在一个例子中,可以采用提出识别一切的模型(Recognize anything plus model,RAM++),通过注入语义概念到图像标记的开集识别方法(Inject Semantic Concepts Into Image Tagging Open-Set Recognition),对视觉数据进行目标检测,得到第一对象类别,这样提取的第一对象类别可以保证动态词表的词语多样性,并且还可以保证提取的视觉数据不同,避免从相同或相似的视觉数据中提取第一对象类别而导致第一对象类别的同质化,可以增加第一对象类别的丰富性和代表性。
第二对象类别可以是对最新的视觉数据进行语义理解得到的对象类别。对视觉数据进行语义理解,得到至少一个第二对象,以及至少一个第二对象的第二对象类别。第一对象类别和第二对象类别的获取方式不同,但均是从
视觉数据中提取得到的对象类别。
在一个例子中,可以采用大语言模型,对视觉数据进行描述。例如,向大语言模型输入“以下的场景可能会包含哪些物体?”,大语言模型基于视觉数据对输入内容进行回复,输出回复内容,从回复内容中提取类别的词语,得到第二对象类别。通过对视觉数据进行描述得到第二对象类别,保证了动态词表的完整性,也保证动态词表中词语的多样性。
可以仅获取第一对象类别,或者仅获取第二对象类别,添加到动态词表中。
通过获取实时更新的视觉数据,并对视觉数据进行目标检测和/或语义理解,从而得到第一对象类别和/或第二对象类别,添加到动态词表中,丰富动态词表中的词语,并基于更加完整丰富和更具代表性的词语,识别样本图像中的对象,增加识别出的对象的多样性和代表性,从而使得标注数据实时可靠多样。
S306、获取样本图像。
样本图像可以与前述用于识别出实时更新的对象类别的图像相同,也可以不同。样本图像可以采集,也可以从公开渠道获取,或者经过授权从未公开的渠道获取。
S307、在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象,得到所述样本图像中目标识别对象。
采用实时更新的动态词表中的每个词语,在样本图像中识别是否存在每个词语对应的对象,在存在的情况下,将存在的对象,确定为目标识别对象。
可以将对应的词语,确定为该目标识别对象的目标类别,或者在样本图像中识别动态词表中每个词语对应的目标识别对象的同时,识别出至少一个目标识别对象的目标类别。将样本图像和该样本图像中识别出的全部目标识别对象进行组合,生成一条样本数据。其中,目标识别对象可以是目标检测框。此外,还可以将样本图像、样本图像中识别出的至少一个目标识别对象,以及识别出的至少一个目标识别对象的目标类别,组合生成样本数据。
对目标识别对象还可以进行预处理,减少错误冗余数据。
可选的,在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,还包括:检测所述样本图像中多个目标识别对象之间的交并比;根据所述多个目标识别对象之间的交并比,对所述多个目标识别对象进行检测,得到相似的多个目标识别对象;对所述相似的多个目标识别对象进行融合,得到融合结果;根据所述融合结果对所述样本图像中多个目标识别对象
进行更新。
实际上,样本图像中存在多个目标识别对象,多个目标识别对象可能存在冗余数据。若仅存在一个目标识别对象,可以不执行该步骤。可以对多个目标识别对象进行去重。可以通过多个目标识别对象之间的交并比去重。多个目标识别对象中,分别计算每两个目标识别对象的目标检测框或者图像区域之间的交并比。在交并比大于或等于预设交并比阈值时,确定这两个目标识别对象相似,得到相似的两个目标识别对象。每两个目标识别对象进行相似度比较,可以得到相似的多个目标识别对象。此外,还可以对目标识别对象的检测框进行语义理解,得到检测框的目标描述信息,并划分为视觉名词视觉形容词对。在两个目标识别对象的交并比大于或等于预设交并比阈值,且对应的目标描述信息中视觉名词视觉形容词对相同时,确定这两个目标识别对象相似;若仅交并比大于或等于预设交并比阈值,或仅词对相同,均不能确定这两个目标识别对象相似。
示例性的,一组相似的多个目标识别对象的融合方式可以是:选择其中一个目标识别对象作为融合结果,剔除剩余的目标识别对象,或者计算该多个目标识别对象的交集或并集,确定为融合结果。同时,在这组相似的多个目标识别对象中,每个目标识别对象存在目标类别,统计相同的目标类别的数量,作为该相同的目标类别的出现次数。统计该组相似的多个目标识别对象中,出现次数最多的目标类别,确定为融合结果的目标类别。根据融合结果对样本图像中目标识别对象进行更新可以是,将融合结果作为样本图像中最终识别到的目标识别对象,实现对样本图像中目标识别对象进行更新。或者根据融合结果,剔除样本图像中冗余的目标识别对象,实现对样本图像中目标识别对象进行更新。
通过对样本图像中识别的多个目标识别对象中,采用交并比剔除冗余数据,减少目标识别对象的冗余数据,减少同一物体的多次标注,以及减少标注的数据量,从而提高标注效率和准确性。
可选的,在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,还包括:检测所述样本图像中至少一个目标识别对象的像素点数量;根据所述至少一个目标识别对象的像素点数量,对所述样本图像中目标识别对象进行更新。
像素点数量用于表征目标识别对象的信息量。实际上,目标识别对象的像素点数量过小,例如,只有一个像素点,很难区分该目标识别对象与其他目标识别对象的区别,同时也难以获取到该目标识别对象的有效信息,该目标识别对象的识别意义较小,即使最终可以进行标注,标注数据也无法提供
有效信息,以训练模型。
将每个目标识别对象的像素点数量与相应像素点数量阈值进行比较,剔除小于相应像素点数量阈值的目标识别对象,保留大于或等于相应像素点数量阈值的目标识别对象。可以设置不同类别对应不同像素点数量阈值。
根据至少一个目标识别对象的像素点数量对样本图像中目标识别对象进行更新可以是,根据至少一个目标识别对象的像素点数量,剔除样本图像中冗余的目标识别对象,示例性的,剔除像素点数量过少的目标识别对象,实现对样本图像中目标识别对象进行更新。
通过根据目标识别对象的像素点数量,对目标识别对象进行筛选,可以保留有效信息更多,关键信息更多的目标识别对象,以增加目标识别对象的代表性,增加标注数据的内容丰富性和代表性,使得标注数据包含更多的有效信息。
前述通过冗余和有效性的维度,对目标识别对象进行筛选,此外,还可以采用其他维度对目标识别对象进行筛选,对此不限定。
实际上,为了可以从目标识别对象中提取到更多有效内容,增加标注数据的准确性,在生成样本数据之前,可以对目标识别对象和样本图像的相对面积占比进行调整。
可选的,在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,还包括:检测所述样本图像中至少一个目标识别对象与所述样本图像的尺寸占比;在所述目标识别对象的尺寸占比小于预设阈值时,对所述目标识别对象的边界进行扩展,以调大扩展后的目标识别对象与所述样本图像的尺寸占比。
目标识别对象的尺寸占比用于表征目标识别对象对于样本图像的重要程度。实际上,为了可以从目标识别对象中提取到更加丰富和完整的数据,通常要求目标识别对象的尺寸占比不能过大或不能过小。尺寸占比,可以是指一个目标识别对象的目标检测框或图像区域的面积与样本图像的面积的占比。目标识别对象的尺寸占比小于预设阈值,表明尺寸占比过小,一般选择调大尺寸占比。
可以向外扩展目标识别对象的边界。通常目标识别对象的目标检测框为矩形,可以将矩形框的边向外移动x距离。扩展后的目标识别对象包括扩展前的目标识别对象,通常扩展前后的目标识别对象的中心无变化。目标识别对象可以等比例扩大。
可选的,在所述样本图像中识别所述动态词表中每个词语对应的目标识
别对象之后,还包括:检测所述样本图像中至少一个目标识别对象与所述样本图像的尺寸占比;在所述目标识别对象的尺寸占比小于预设阈值时,对所述样本图像进行裁剪,并更新所述样本图像,以调大所述目标识别对象与裁剪后的样本图像的尺寸占比。
对样本图像进行裁剪,缩小样本图像的尺寸,同时裁剪的区域与目标识别对象不存在相交区域。可以在样本图像中确定至少一个目标识别对象的位置,并避开这些位置,裁剪部分背景区域,同时裁剪前后样本图像的形状、角度和位置等无变化。
可以仅扩大目标识别对象的尺寸,或者仅裁剪样本图像,缩小样本图像的尺寸,又或者可以同时扩大目标识别对象的尺寸,以及裁剪样本图像,使得尺寸占比变大,可以增加标注过程中所使用的图像区域是最具信息量的,提高了标注的准确性。
根据本申请的技术方案,通过获取最新的对象类别,并添加到动态词表中,根据动态词表中每个词语对目标识别对象进行识别,得到目标识别对象,可以增加可识别的对象的多样性,从而增加标注数据的多样性,丰富标注任务的内容,并且可以实时更新对象类别,相应更新识别的对象,从而实时更新标注数据,可以使得标注的数据具有开放性,可以动态变化,提高标注数据的准确性和实时性。
图4是根据本申请实施例提供的一种视觉语义理解模型的训练方法的流程图。本实施例可以适用于根据前述生成的标注数据结合样本图像训练视觉语义理解模型,视觉语义理解模型用于输入图像,输出图像的描述内容的情况。本实施例方法可以由视觉语义理解模型的训练装置来执行,该装置可采用软件和/或硬件的方式实现,并配置于具有一定数据运算能力的电子设备中,该电子设备可以是客户端设备或服务器设备,客户端设备可以包括:个人计算机、笔记本电脑、智能手机、平板电脑、物联网设备或便携式可穿戴设备等。
S401、获取样本数据,所述样本数据包括样本图像和标注数据,所述标注数据通过如本申请任一个实施例所述的图像标注数据生成方法生成。
将标注数据作为样本图像的描述真值内容,和样本图像输入到视觉语义理解模型中训练视觉语义理解模型。
S402、采用所述样本数据,训练视觉语义理解模型。
视觉语义理解模型可以是大语言模型。
根据本申请的技术方案,通过经过校验的标注数据作为训练数据训练视觉语义理解模型,可以提高视觉语义理解模型输出的描述文本的准确性,同时自动化生成标注数据,可以减少模型训练成本,并提高模型训练效率。
图5是根据本申请实施例提供的一种图像标注数据生成方法的场景图。本申请实施例主要划分两个阶段:一阶段和二阶段,其中,一阶段包括5个小步骤,例如可以是生成样本数据,样本数据包括图像、图像中检测框和类别标签。二阶段包括4个小步骤,例如可以是生成检测框的目标描述信息,即实现对检测框的标注。
一阶段和二阶段概述:采用RAM++对样本图像进行目标检测,并为检测得到的对象添加类别的标签,以及采用开放世界本地化版本2(Open-World Localization version2,OWLv2)实现面向开放词汇下的目标检测(Open Vocabulary object Detection,OVD)。对生成的检测框通过过滤(Filter)模块进行框抑制;生成视觉语义理解模型的输入,该视觉语义理解模型用于生成描述信息(Caption),输入需要包含物体的检测框(BOX)以及标签(LABEL)(可以是指类别标签)。将图像、图像中物体,即目标识别对象的BOX以及LABEL输入至大语言模型中,得到初始版本的目标描述信息Caption。大语言模型可以是多模态大语言模型(MultiModal Large Language Models,MLLM),示例性的,MLLM模型为shikra模型。通常MLLM模型包括视觉编码层、融合层、文本编码层和解码层。对初始版本的目标描述信息Caption进行词义分解,得到对应的名词与形容词,即视觉名词和视觉形容词。对词义进行解析,对“形容词”划分为材质、颜色、行为和位置等;分别对于不同的“形容词”进行双重校验(double check),保证过滤后的“形容词”的准确性;对“名词”与tagging信息,即类别标签,进行匹配,确保“名词”对准确性。对生成后的解析词进行语义拼接与多样性生成,得到生成的目标描述信息Caption。通过MLLM进行核准,提高生成描述信息Caption的能力。
示例性的,1、通过人工设置至少一个静态词表:通过提取OBJECT365数据集、COCO数据集和OPENIMAGES数据集中的类别,并结合业务需求,收集了大于500的数量的类别作为静态词表,但静态词表对图片不具有开放性,即在后期新的业务需求中,没有办法提取到更多元的框的信息,于是,通过Tagging能力去提取每一张图片的特有的词表信息。
2、通过Tagging得到图片的动态词表:执行如图5中步骤1-1添加标签(tagging),RAM++通过爬取网络上成百万的视觉数据(图像或视频中图像),
对视觉数据进行自动化标注,实现粗粒度添加标签的操作,在视觉数据中添加图像级别的标签,得到大量数据图文对,并根据自动化标注的标签,提取对应的词信息,可识别大于6000的数量的类别,并可以识别出图片中的类别以及识别背景,将提取的词信息添加到动态词表中。通过tagging,可以保证图片词表的多样性,保证提取的多个图片不同。例如,动态词表可以包括:人物、广告、车、动物、场景、食物、家具和电器等。此外还有更多更丰富的词语,对此不限定。
3、通过大语言模型得到的想象词表:同时提取了RAM++拥有的背景类,例如park(公园)和kitchen(厨房)等,利用大语言模型去生成“以下的场景可能会包含哪些物体?”通过这样的方式保证了词表的完整性,也保证图片词表生成的多样性。
执行步骤1-2添加,将动态词表中词语添加到预设词表中,此外,还可以将前述得到的静态词表和想象词表中词语也添加到预设词表中。
4、通过开放域检测能力生成检测框:利用基于前述的词表的开放域词表模型,基于预设词表,执行步骤1-3推理,在样本图像中检测得到检测框。在本申请实施例中利用了OWLv2,生成预设词表对应的检测框。例如可以通过输入图片原图,即样本图像,并结合预设词表,生成样本图像中预设词表中词语对应的检测框。例如,在样本图像中识别对象1和对象2,其中,对象1和对象2的类别可以不同也可以相同。
5、生成的检测框执行步骤1-4输入到Filter模块,通过Filter模块进行框抑制,基于步骤1-5输出多组图像、检测框和类别标签。通过人工标注,将静态类别分为至少一个大类,例如:食物、工具、人、交通场景和工具等中的至少一项等类别,主要的抑制内容为食物和人等,操作为:
a)对caption提取相应的名词和形容词,例如,描述信息为a red apple,可以提取出{<red>,<apple>}。
b)一张图像将生成至少一个形容词名词对,以及对应的box位置,将形容词、名词和box联合生成列表[{<adj.1>,<n.1>,box1},{adj.2>,<n.2>,box2}…]。
c)循环列表,计算列表中box间的交并比(Intersection over Union,IOU),如果两个box的IOU大于或等于预设交并比阈值,且形容词(adj)和名词(n)相同,则融合两个box为一个box。此外,还可以是如果两个box的IOU大于或等于预设交并比阈值,则融合两个box为一个box。其他情况不融合。
示例性的,
i.Box_new=
[min(box1.x,box2.x),min(box1.y,box2.y),max(box1.w,box2.w),max(box1.h,box2.h)
其中,Box_new为融合后的检测框。融合后的检测框的中心坐标为融合的两个检测框的中心坐标的最小值;融合后的检测框的宽高为融合的两个检测框的宽高的最大值。
ii.合并对应的label内容,例如可以是提取增强label作为融合后的检测框对应的label,增强label是指语义最丰富的label内容,可以对label进行语义分割,得到名词形容词对,并保留数量最多的label。
此时,一阶段的任务,从步骤1-1到步骤1-5,执行完成,生成样本数据。其中,样本数据包括样本图像、至少一个检测框和至少一个检测框的类别标签,一个检测框表示一个目标识别对象。下面执行二阶段任务。
6、执行步骤2-1输入,将Filter后的box、label以及image输入至MLLM得到初始目标描述信息caption,此时的caption具有很多问题,出现的问题可以包括下述至少一项:表述错误、颜色描述错误和物体描述错误等问题,通过语义划分结合语义重组的方式解决前述问题。
7、在将Filter后的box、label以及image输入至自动标注MLLM之前,可以检测样本图像中的检测框对样本图像的尺寸占比,并根据此占比来决定后续处理流程。
首先,如果该尺寸占比低于预设阈值,根据检测框的位置和占比,将样本图像裁剪成新的适宜尺寸的样本图像。这确保了标注过程中所使用的样本图像是最具信息量的,提高了标注的准确性。
相反,如果该尺寸占比高于或等于预设阈值,直接使用原始的样本图像,无需进行裁剪操作。
8、执行将裁剪后的样本图像、相应的检测框和类别标签输入到MLLM中。该模型能够生成与样本图像中对应检测框相关的目标描述信息,包括物体主体、形容描述词以及物体主体同周围物体的关系等。为了提高标注结果的质量,引入了标注过滤模块和标注校正模块。
执行步骤2-2,向标注过滤模块和标注校正模块输出视觉描述信息。通过标注过滤模块对生成的视觉描述信息进行初步判断,检测是否存在与实际样本图像内容不一致的情况。如果有差异,将进入标注校正模块的处理流程。标注过滤模块对视觉描述进行语义划分,得到名词和形容词,此外,关系也可以当做形容词,或者从形容词中单独拆分出来。其中,名词可以是指检测
框的物体主体,形容词可以是形容描述词,关系可以是指物体主题与周围物体的关系。语义划分结合语义重组实现对名词和形容词进行清洗,以及正确名词和正确形容词的重组,得到准确的视觉描述信息,实现清洗视觉描述信息。
标注校正模块的任务是对名词和形容词进行相应的校正。标注校正模块确保最终的标注结果更加准确地反映了实际图像内容。执行步骤2-3,标注校正模块将校正后词语输入到MLLM中,进行描述信息的语义重组,得到多个校正后的视觉描述信息。执行步骤2-2,向标注过滤模块输出校正后的视觉描述信息,标注过滤模块可以对多个校正后的描述信息进行筛选,得到最终输出的目标描述信息。执行步骤2-4,输出筛选得到的目标描述信息,并作为最终标注结果,完成整个图像标注流程。
如果前一阶段的视觉描述信息与实际图像内容吻合,执行步骤2-4,直接将前一阶段的视觉描述信息作为目标描述信息,并作为最终标注结果输出,完成整个图像标注流程。
此时,二阶段任务,从步骤2-1到步骤2-4,执行完成,生成样本数据的目标描述信息。通过这一系统化的处理流程,提高了标注的自动化程度和标注结果的准确性。最后得到的标注数据可以用于训练MLLM,也可以用于训练OWLv2。
本申请技术方案,通过自动标注技术,通过词性检查和box过滤,并充分利用生成式预训练转换模型(Generative Pre-Trained Transformer,GPT)的校准过程,能够有效的提高标注过程的准确性,有效提高自动化标注过程的可信度。
根据本申请的实施例,图6是本申请实施例中的图像标注数据生成装置的结构图,本申请实施例适用于生成图像的描述内容的情况。该装置采用软件和/或硬件实现,并配置于具备一定数据运算能力的电子设备中。
如图6所示的一种图像标注数据生成装置600,包括:样本数据获取模块601、视觉描述校验模块602、目标描述信息调整模块603和标注数据生成模块604。其中,
样本数据获取模块601,设置为获取样本图像、所述样本图像中目标识别对象和所述目标识别对象的视觉描述信息;
视觉描述校验模块602,设置为根据所述样本图像和所述目标识别对象,对所述视觉描述信息中词语进行校验,得到所述视觉描述信息的校验结果;
目标描述信息调整模块603,设置为根据所述视觉描述信息的校验结果,对所述视觉描述信息中词语进行调整,得到所述目标识别对象的目标描述信息;
标注数据生成模块604,设置为将所述目标识别对象的目标描述信息,作为所述目标识别对象的标注数据;所述标注数据用于训练视觉语义理解模型,所述视觉语义理解模型用于根据输入的视觉内容,输出所述视觉内容的描述文本。
根据本申请的技术方案,通过自动提取目标识别对象的视觉描述信息,并对视觉描述信息中词语进行校验,得到校验结果,同时根据校验结果,对视觉描述信息中词语进行调整,得到目标描述信息,作为目标识别对象的标注数据,并添加到样本数据,训练图像语义理解模型,以得到对输入的视觉内容进行语义理解的模型,提高图像标注数据的生成效率,降低标注成本,同时兼顾标注准确性,进而快速生成正确标注数据,以训练模型,提高视觉语义理解模型的描述文本的准确性。
在一个或多个实施例中,所述视觉描述校验模块602,包括:描述分词单元,设置为将视觉描述信息划分为至少一个视觉名词和至少一个视觉形容词;名词校验单元,设置为根据所述样本图像和所述目标识别对象,对所述至少一个视觉名词进行校验,得到所述至少一个视觉名词对应的名词校验结果;形容词校验单元,设置为根据所述样本图像和所述目标识别对象,对所述至少一个视觉形容词进行校验,得到所述至少一个视觉形容词对应的形容词校验结果;校验结果确定单元,设置为将所述名词校验结果和所述形容词校验结果,确定为所述视觉描述信息的校验结果。
在一个或多个实施例中,所述名词校验单元,包括:类别获取子单元,设置为根据所述样本图像和所述目标识别对象,获取所述目标识别对象的类别描述信息;类别校验子单元,设置为针对所述至少一个视觉名词,检测所述目标识别对象的类别描述信息与所述视觉名词的一致性,得到所述视觉名词对应的名词校验结果。
在一个或多个实施例中,所述形容词校验单元,包括:交互生成子单元,设置为针对所述至少一个视觉形容词,根据所述视觉形容词,以及所述视觉形容词对应的视觉名词,生成交互对话内容;形容词语义理解子单元,设置为将所述样本图像和所述目标识别对象作为上下文,以及将所述交互对话内容输入到大语言模型中,得到所述视觉形容词的存在判断结果;形容词校验子单元,设置为根据所述视觉形容词的存在判断结果,生成所述视觉形容词的形容词校验结果。
在一个或多个实施例中,所述目标描述信息调整模块603,包括:名词筛选单元,设置为根据每个名词校验结果,在所述至少一个视觉名词中筛选,得到至少一个目标名词;形容词筛选单元,设置为根据每个形容词校验结果,在所述至少一个视觉形容词中筛选,得到至少一个目标形容词;描述词组生成单元,设置为根据所述至少一个目标名词和所述至少一个目标形容词,生成至少一个描述词组;描述信息生成单元,设置为根据所述至少一个描述词组,生成目标识别对象的目标描述信息。
在一个或多个实施例中,图像标注数据生成装置,还包括:实时类别获取模块,设置为获取实时更新的对象类别,并将获取的实时更新的对象类别添加到动态词表中;图像获取模块,设置为获取样本图像;对象更新获取模块,设置为在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象,得到所述样本图像中目标识别对象。
在一个或多个实施例中,所述实时类别获取模块,包括:视觉集获取单元,设置为获取实时更新的视觉集;目标检测单元,设置为对所述视觉集中视觉数据进行目标检测,得到至少一个第一对象类别;和/或语义理解单元,设置为对所述视觉集中视觉数据进行语义理解,得到至少一个第二对象类别;类别更新单元,设置为将所述第一对象类别和/或所述第二对象类别,确定为实时更新的对象类别。
在一个或多个实施例中,图像标注数据生成装置,还包括:交并比检测模块,设置为在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,检测所述样本图像中多个目标识别对象之间的交并比;相似对象检测模块,设置为根据所述多个目标识别对象之间的交并比,对所述多个目标识别对象进行检测,得到相似的多个目标识别对象;对象融合模块,设置为对所述相似的多个目标识别对象进行融合,得到融合结果;冗余筛选模块,设置为根据所述融合结果对所述样本图像中多个目标识别对象进行更新。
在一个或多个实施例中,图像标注数据生成装置,还包括:交并比检测模块,设置为在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,检测所述样本图像中至少一个目标识别对象的像素点数量;像素点数量筛选模块,设置为根据所述至少一个目标识别对象的像素点数量,对所述样本图像中目标识别对象进行更新。
在一个或多个实施例中,图像标注数据生成装置,还包括:尺寸占比检测模块,设置为在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,检测所述样本图像中至少一个目标识别对象与所述样本图像的尺寸占比;对象边界扩展模块,设置为在所述目标识别对象的尺寸占比小
于预设阈值时,对所述目标识别对象的边界进行扩展,以调大扩展后的目标识别对象与所述样本图像的尺寸占比。
在一个或多个实施例中,图像标注数据生成装置,还包括:尺寸占比检测模块,设置为在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,检测所述样本图像中至少一个目标识别对象与所述样本图像的尺寸占比;图像裁剪模块,设置为在所述目标识别对象的尺寸占比小于预设阈值时,对所述样本图像进行裁剪,并更新所述样本图像,以调大所述目标识别对象与裁剪后的样本图像的尺寸占比。
上述图像标注数据生成装置可执行本申请任意实施例所提供的图像标注数据生成方法,具备执行图像标注数据生成方法相应的功能模块和效果。
根据本申请的实施例,图7是本申请实施例中的视觉语义理解模型的训练装置的结构图,本申请实施例适用于训练视觉语义理解模型的情况。该装置采用软件和/或硬件实现,并配置于具备一定数据运算能力的电子设备中。
如图7所示的一种视觉语义理解模型的训练装置700,包括:样本数据获取模块701和模型训练模块702。
样本数据获取模块701,设置为获取样本数据,所述样本数据包括样本图像和标注数据,所述标注数据通过如本申请任一实施例所述的图像标注数据生成方法生成;
模型训练模块702,设置为采用所述样本数据,训练视觉语义理解模型。
根据本申请的技术方案,通过经过校验的标注数据作为训练数据训练视觉语义理解模型,可以提高视觉语义理解模型输出的描述文本的准确性,同时自动化生成标注数据,可以减少模型训练成本,并提高模型训练效率。
上述视觉语义理解模型的训练装置可执行本申请任意实施例所提供的视觉语义理解模型的训练方法,具备执行视觉语义理解模型的训练方法相应的功能模块和效果。
本申请的技术方案中,所涉及的用户个人信息的收集、存储、使用、加工、传输、提供和公开等处理,均符合相关法律法规的规定,且不违背公序良俗。
根据本申请的实施例,本申请还提供了一种电子设备、一种可读存储介质和一种计算机程序产品。
图8示出了可以用来实施本申请的实施例的示例电子设备800的示意性区域图。电子设备旨在表示多种形式的数字计算机,诸如,膝上型计算机、台式计算机、工作台、个人数字助理、服务器、刀片式服务器、大型计算机、和其它适合的计算机。电子设备还可以表示多种形式的移动装置,诸如,个人数字处理、蜂窝电话、智能电话、可穿戴设备和其它类似的计算装置。本文所示的部件、它们的连接和关系、以及它们的功能仅仅作为示例,并且不意在限制本文中描述的和/或者要求的本申请的实现。
如图8所示,电子设备800包括计算单元801,计算单元801可以根据存储在只读存储器(Read-Only Memory,ROM)802中的计算机程序或者从存储单元808加载到随机访问存储器(Random Access Memory,RAM)803中的计算机程序,来执行多种适当的动作和处理。在RAM 803中,还可存储电子设备800操作所需的多种程序和数据。计算单元801、ROM 802以及RAM 803通过总线804彼此相连。输入/输出(Input/Output,I/O)接口805也连接至总线804。
电子设备800中的多个部件连接至I/O接口805,包括:输入单元806,例如键盘、鼠标等;输出单元807,例如多种类型的显示器、扬声器等;存储单元808,例如磁盘、光盘等;以及通信单元809,例如网卡、调制解调器、无线通信收发机等。通信单元809允许电子设备800通过诸如因特网的计算机网络和/或多种电信网络与其他设备交换信息/数据。
计算单元801可以是多种具有处理和计算能力的通用和/或专用处理组件。计算单元801的一些示例包括但不限于中央处理单元(Central Processing Unit,CPU)、图形处理单元(Graphics Processing Unit,GPU)、多种专用的人工智能(Artificial Intelligence,AI)计算芯片、多种运行机器学习模型算法的计算单元、数字信号处理器(Digital Signal Processor,DSP)、以及任何适当的处理器、控制器、微控制器等。计算单元801执行上文所描述的多个方法和处理,例如图像标注数据生成方法或视觉语义理解模型的训练方法。例如,在一些实施例中,图像标注数据生成方法或视觉语义理解模型的训练方法可被实现为计算机软件程序,其被有形地包含于机器可读介质,例如存储单元808。在一些实施例中,计算机程序的部分或者全部可以经由ROM 802和/或通信单元809而被载入和/或安装到电子设备800上。当计算机程序加载到RAM 803并由计算单元801执行时,可以执行上文描述的图像标注数据生成方法或视觉语义理解模型的训练方法的一个或多个步骤。备选地,在其他实施例中,计算单元801可以通过其他任何适当的方式(例如,借助于固件)而被配置为执行图像标注数据生成方法或视觉语义理解模型的训练方法。
本文中以上描述的系统和技术的多种实施方式可以在数字电子电路系统、集成电路系统、现场可编程门阵列(Field Programmable Gate Array,FPGA)、专用集成电路(Application Specific Integrated Circuit,ASIC)、专用标准对象(Application Specific Standard Parts,ASSP)、芯片上系统的系统(System on Chip,SOC)、复杂可编程逻辑设备(Complex Programmable Logic Device,CPLD)、计算机硬件、固件、软件、和/或它们的组合中实现。这些多种实施方式可以包括:实施在一个或者多个计算机程序中,该一个或者多个计算机程序可在包括至少一个可编程处理器的可编程系统上执行和/或解释,该可编程处理器可以是专用或者通用可编程处理器,可以从存储系统、至少一个输入装置、和至少一个输出装置接收数据和指令,并且将数据和指令传输至该存储系统、该至少一个输入装置、和该至少一个输出装置。
用于实施本申请的方法的程序代码可以采用一个或多个编程语言的任何组合来编写。这些程序代码可以提供给通用计算机、专用计算机或其他可编程数据处理装置的处理器或控制器,使得程序代码当由处理器或控制器执行时使流程图和/或区域图中所规定的功能/指令被实施。程序代码可以完全在机器上执行、部分地在机器上执行,作为独立软件包部分地在机器上执行且部分地在远程机器上执行或完全在远程机器或服务器上执行。
在本申请的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、RAM、ROM、可擦除可编程只读存储器(Erasable Programmable Read-Only Memory,EPROM)、快闪存储器、光纤、便捷式紧凑盘只读存储器(Compact Disc Read Only Memory,CD-ROM)、光学储存设备、磁储存设备、或上述内容的任何合适组合。存储介质可以是非暂态(non-transitory)存储介质。
为了提供与用户的交互,可以在计算机上实施此处描述的系统和技术,该计算机具有:用于向用户显示信息的显示装置(例如,阴极射线管(Cathode Ray Tube,CRT)或者液晶显示器(Liquid Crystal Display,LCD)监视器);以及键盘和指向装置(例如,鼠标或者轨迹球),用户可以通过该键盘和该指向装置来将输入提供给计算机。其它种类的装置还可以用于提供与用户的交互;例如,提供给用户的反馈可以是任何形式的传感反馈(例如,视觉反馈、听觉反馈、或者触觉反馈);并且可以用任何形式(包括声输入、语音输入或者、触觉输入)来接收来自用户的输入。
可以将此处描述的系统和技术实施在包括后台部件的计算系统(例如,作为数据服务器)、或者包括中间件部件的计算系统(例如,应用服务器)、或者包括前端部件的计算系统(例如,具有图形用户界面或者网络浏览器的用户计算机,用户可以通过该图形用户界面或者该网络浏览器来与此处描述的系统和技术的实施方式交互)、或者包括这种后台部件、中间件部件、或者前端部件的任何组合的计算系统中。可以通过任何形式或者介质的数字数据通信(例如,通信网络)来将系统的部件相互连接。通信网络的示例包括:局域网(Local Area Network,LAN)、广域网(Wide Area Network,WAN)区块链网络和互联网。
计算机系统可以包括客户端和服务器。客户端和服务器一般远离彼此并且通常通过通信网络进行交互。通过在相应的计算机上运行并且彼此具有客户端-服务器关系的计算机程序来产生客户端和服务器的关系。服务器可以是云服务器,又称为云计算服务器或云主机,是云计算服务体系中的一项主机产品,以解决了传统物理主机与虚拟专用服务器(Virtual Private Server,VPS)服务中,存在的管理难度大,业务扩展性弱的缺陷。服务器也可以为分布式系统的服务器,或者是结合了区块链的服务器。
人工智能是研究使计算机来模拟人的一些思维过程和智能行为(如学习、推理、思考、规划等)的学科,既有硬件层面的技术也有软件层面的技术。人工智能硬件技术一般包括如传感器、专用人工智能芯片、云计算、分布式存储、大数据处理等技术;人工智能软件技术主要包括计算机视觉技术、语音识别技术、自然语言处理技术及机器学习/深度学习技术、大数据处理技术、知识图谱技术等几大方向。
云计算(cloud computing),指的是通过网络接入弹性可扩展的共享物理或虚拟资源池,资源可以包括服务器、指令系统、网络、软件、应用和存储设备等,并可以按需、自服务的方式对资源进行部署和管理的技术体系。通过云计算技术,可以为人工智能、区块链等技术应用、模型训练提供高效强大的数据处理能力。
可以使用上面所示的多种形式的流程,重新排序、增加或删除步骤。例如,本发申请中记载的多个步骤可以并行地执行也可以顺序地执行也可以不同的次序执行,只要能够实现本申请提供的技术方案所期望的结果,本文在此不进行限制。
Claims (27)
- 一种图像标注数据生成方法,包括:获取样本图像、所述样本图像中目标识别对象和所述目标识别对象的视觉描述信息;根据所述样本图像和所述目标识别对象,对所述视觉描述信息中词语进行校验,得到所述视觉描述信息的校验结果;根据所述视觉描述信息的校验结果,对所述视觉描述信息中词语进行调整,得到所述目标识别对象的目标描述信息;将所述目标识别对象的目标描述信息,作为所述目标识别对象的标注数据;其中,所述标注数据用于训练视觉语义理解模型,所述视觉语义理解模型用于根据输入的视觉内容,输出所述视觉内容的描述文本。
- 根据权利要求1所述的方法,其中,所述根据所述样本图像和所述目标识别对象,对所述视觉描述信息中词语进行校验,得到所述视觉描述信息的校验结果,包括:将所述视觉描述信息划分为至少一个视觉名词和至少一个视觉形容词;根据所述样本图像和所述目标识别对象,对所述至少一个视觉名词进行校验,得到所述至少一个视觉名词对应的名词校验结果;根据所述样本图像和所述目标识别对象,对所述至少一个视觉形容词进行校验,得到所述至少一个视觉形容词对应的形容词校验结果;将所述名词校验结果和所述形容词校验结果,确定为所述视觉描述信息的校验结果。
- 根据权利要求2所述的方法,其中,所述根据所述样本图像和所述目标识别对象,对所述至少一个视觉名词进行校验,得到所述至少一个视觉名词对应的名词校验结果,包括:根据所述样本图像和所述目标识别对象,获取所述目标识别对象的类别描述信息;针对所述至少一个视觉名词,检测所述目标识别对象的类别描述信息与所述至少一个视觉名词的一致性,得到所述至少一个视觉名词对应的名词校验结果。
- 根据权利要求2所述的方法,其中,所述根据所述样本图像和所述目标识别对象,对所述至少一个视觉形容词进行校验,得到所述至少一个视觉形容词对应的形容词校验结果,包括:针对所述至少一个视觉形容词,根据所述至少一个视觉形容词,以及所述至少一个视觉形容词对应的视觉名词,生成交互对话内容;将所述样本图像和所述目标识别对象作为上下文,以及将所述交互对话内容输入到大语言模型中,得到所述至少一个视觉形容词的存在判断结果;根据所述至少一个视觉形容词的存在判断结果,生成所述至少一个视觉形容词的形容词校验结果。
- 根据权利要求2所述的方法,其中,所述根据所述视觉描述信息的校验结果,对所述视觉描述信息中词语进行调整,得到所述目标识别对象的目标描述信息,包括:根据每个名词校验结果,在所述至少一个视觉名词中筛选,得到至少一个目标名词;根据每个形容词校验结果,在所述至少一个视觉形容词中筛选,得到至少一个目标形容词;根据所述至少一个目标名词和所述至少一个目标形容词,生成至少一个描述词组;根据所述至少一个描述词组,生成所述目标识别对象的目标描述信息。
- 根据权利要求1所述的方法,其中,所述样本图像和所述目标识别对象,通过以下方式得到:获取实时更新的对象类别,并将获取的实时更新的对象类别添加到动态词表中;获取所述样本图像;在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象,得到所述样本图像中目标识别对象。
- 根据权利要求6所述的方法,其中,所述获取实时更新的对象类别,包括:获取实时更新的视觉集;对所述视觉集中视觉数据进行目标检测,得到至少一个第一对象类别;或,对所述视觉集中视觉数据进行语义理解,得到至少一个第二对象类别;或,对所述视觉集中视觉数据进行目标检测,得到至少一个第一对象类别且对所述视觉集中视觉数据进行语义理解,得到至少一个第二对象类别;将所述第一对象类别或所述第二对象类别中至少之一,确定为实时更新的 对象类别。
- 根据权利要求6所述的方法,在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,还包括:检测所述样本图像中多个目标识别对象之间的交并比;根据所述多个目标识别对象之间的交并比,对所述多个目标识别对象进行检测,得到相似的多个目标识别对象;对所述相似的多个目标识别对象进行融合,得到融合结果;根据所述融合结果对所述样本图像中多个目标识别对象进行更新。
- 根据权利要求6所述的方法,在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,还包括:检测所述样本图像中至少一个目标识别对象的像素点数量;根据所述至少一个目标识别对象的像素点数量,对所述样本图像中目标识别对象进行更新。
- 根据权利要求6所述的方法,其中,在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,还包括:检测所述样本图像中至少一个目标识别对象与所述样本图像之间的尺寸占比;在所述目标识别对象的尺寸占比小于预设阈值的情况下,对所述目标识别对象的边界进行扩展,以调大扩展后的目标识别对象与所述样本图像之间的尺寸占比。
- 根据权利要求6所述的方法,在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,还包括:检测所述样本图像中至少一个目标识别对象与所述样本图像之间的尺寸占比;在所述目标识别对象的尺寸占比小于预设阈值的情况下,对所述样本图像进行裁剪,并更新所述样本图像,以调大所述目标识别对象与裁剪后的样本图像之间的尺寸占比。
- 一种视觉语义理解模型的训练方法,包括:获取样本数据,其中,所述样本数据包括样本图像和标注数据,所述标注数据通过如权利要求1~11中任一项所述的图像标注数据生成方法生成;采用所述样本数据,训练视觉语义理解模型。
- 一种图像标注数据生成装置,包括:样本数据获取模块,设置为获取样本图像、所述样本图像中目标识别对象和所述目标识别对象的视觉描述信息;视觉描述校验模块,设置为根据所述样本图像和所述目标识别对象,对所述视觉描述信息中词语进行校验,得到所述视觉描述信息的校验结果;目标描述信息调整模块,设置为根据所述视觉描述信息的校验结果,对所述视觉描述信息中词语进行调整,得到所述目标识别对象的目标描述信息;标注数据生成模块,设置为将所述目标识别对象的目标描述信息,作为所述目标识别对象的标注数据;其中,所述标注数据用于训练视觉语义理解模型,所述视觉语义理解模型用于根据输入的视觉内容,输出所述视觉内容的描述文本。
- 根据权利要求13所述的装置,其中,所述视觉描述校验模块,包括:描述分词单元,设置为将所述视觉描述信息划分为至少一个视觉名词和至少一个视觉形容词;名词校验单元,设置为根据所述样本图像和所述目标识别对象,对所述至少一个视觉名词进行校验,得到所述至少一个视觉名词对应的名词校验结果;形容词校验单元,设置为根据所述样本图像和所述目标识别对象,对所述至少一个视觉形容词进行校验,得到所述至少一个视觉形容词对应的形容词校验结果;校验结果确定单元,设置为将所述名词校验结果和所述形容词校验结果,确定为所述视觉描述信息的校验结果。
- 根据权利要求14所述的装置,其中,所述名词校验单元,包括:类别获取子单元,设置为根据所述样本图像和所述目标识别对象,获取所述目标识别对象的类别描述信息;类别校验子单元,设置为针对所述至少一个视觉名词,检测所述目标识别对象的类别描述信息与所述至少一个视觉名词的一致性,得到所述至少一个视觉名词对应的名词校验结果。
- 根据权利要求14所述的装置,其中,所述形容词校验单元,包括:交互生成子单元,设置为针对所述至少一个视觉形容词,根据所述至少一个视觉形容词,以及所述至少一个视觉形容词对应的视觉名词,生成交互对话内容;形容词语义理解子单元,设置为将所述样本图像和所述目标识别对象作为上下文,以及将所述交互对话内容输入到大语言模型中,得到所述至少一个视觉形容词的存在判断结果;形容词校验子单元,设置为根据所述至少一个视觉形容词的存在判断结果,生成所述至少一个视觉形容词的形容词校验结果。
- 根据权利要求14所述的装置,其中,所述目标描述信息调整模块,包括:名词筛选单元,设置为根据每个名词校验结果,在所述至少一个视觉名词中筛选,得到至少一个目标名词;形容词筛选单元,设置为根据每个形容词校验结果,在所述至少一个视觉形容词中筛选,得到至少一个目标形容词;描述词组生成单元,设置为根据所述至少一个目标名词和所述至少一个目标形容词,生成至少一个描述词组;描述信息生成单元,设置为根据所述至少一个描述词组,生成所述目标识别对象的目标描述信息。
- 根据权利要求13所述的装置,还包括:实时类别获取模块,设置为获取实时更新的对象类别,并将获取的实时更新的对象类别添加到动态词表中;图像获取模块,设置为获取所述样本图像;对象更新获取模块,设置为在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象,得到所述样本图像中目标识别对象。
- 根据权利要求18所述的装置,其中,所述实时类别获取模块,包括:视觉集获取单元,设置为获取实时更新的视觉集;目标检测单元或语义理解单元中至少之一:其中,目标检测单元,设置为对所述视觉集中视觉数据进行目标检测,得到至少一个第一对象类别;语义理解单元,设置为对所述视觉集中视觉数据进行语义理解,得到至少一个第二对象类别;类别更新单元,设置为将所述第一对象类别或所述第二对象类别中至少之一,确定为实时更新的对象类别。
- 根据权利要求18述的装置,还包括:交并比检测模块,设置为在所述样本图像中识别所述动态词表中每个词语 对应的目标识别对象之后,检测所述样本图像中多个目标识别对象之间的交并比;相似对象检测模块,设置为根据所述多个目标识别对象之间的交并比,对所述多个目标识别对象进行检测,得到相似的多个目标识别对象;对象融合模块,设置为对所述相似的多个目标识别对象进行融合,得到融合结果;冗余筛选模块,设置为根据所述融合结果对所述样本图像中多个目标识别对象进行更新。
- 根据权利要求18所述的装置,还包括:交并比检测模块,设置为在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,检测所述样本图像中至少一个目标识别对象的像素点数量;像素点数量筛选模块,设置为根据所述至少一个目标识别对象的像素点数量,对所述样本图像中目标识别对象进行更新。
- 根据权利要求18所述的装置,还包括:尺寸占比检测模块,设置为在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,检测所述样本图像中至少一个目标识别对象与所述样本图像之间的尺寸占比;对象边界扩展模块,设置为在所述目标识别对象的尺寸占比小于预设阈值的情况下,对所述目标识别对象的边界进行扩展,以调大扩展后的目标识别对象与所述样本图像之间的尺寸占比。
- 根据权利要求18所述的装置,还包括:尺寸占比检测模块,设置为在所述样本图像中识别所述动态词表中每个词语对应的目标识别对象之后,检测所述样本图像中至少一个目标识别对象与所述样本图像之间的尺寸占比;图像裁剪模块,设置为在所述目标识别对象的尺寸占比小于预设阈值的情况下,对所述样本图像进行裁剪,并更新所述样本图像,以调大所述目标识别对象与裁剪后的样本图像之间的尺寸占比。
- 一种视觉语义理解模型的训练装置,包括:样本数据获取模块,设置为获取样本数据,其中,所述样本数据包括样本图像和标注数据,所述标注数据通过如权利要求1~11中任一项所述的图像标注数据生成方法生成;模型训练模块,设置为采用所述样本数据,训练视觉语义理解模型。
- 一种电子设备,包括:至少一个处理器;以及与所述至少一个处理器通信连接的存储器;其中,所述存储器存储有可被所述至少一个处理器执行的指令,所述指令被所述至少一个处理器执行,以使所述至少一个处理器能够执行权利要求1~11中任一项所述的图像标注数据生成方法,或权利要求12所述的视觉语义理解模型的训练方法。
- 一种存储有计算机指令的非瞬时计算机可读存储介质,其中,所述计算机指令用于使计算机执行根据权利要求1~11中任一项所述的图像标注数据生成方法,或权利要求12所述的视觉语义理解模型的训练方法。
- 一种计算机程序产品,包括计算机程序,所述计算机程序在被处理器执行时实现根据权利要求1~11中任一项所述的图像标注数据生成方法,或权利要求12所述的视觉语义理解模型的训练方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202410268967.1A CN118097694A (zh) | 2024-03-08 | 2024-03-08 | 图像标注数据生成、模型训练方法、装置、设备及介质 |
| CN202410268967.1 | 2024-03-08 |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| WO2025185014A1 true WO2025185014A1 (zh) | 2025-09-12 |
| WO2025185014A8 WO2025185014A8 (zh) | 2025-10-02 |
Family
ID=91142073
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/100320 Pending WO2025185014A1 (zh) | 2024-03-08 | 2024-06-20 | 图像标注数据生成、模型训练方法、装置、设备及介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN118097694A (zh) |
| WO (1) | WO2025185014A1 (zh) |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN118097694A (zh) * | 2024-03-08 | 2024-05-28 | 北京百度网讯科技有限公司 | 图像标注数据生成、模型训练方法、装置、设备及介质 |
| CN118397643B (zh) * | 2024-06-28 | 2024-09-24 | 浪潮电子信息产业股份有限公司 | 一种图像处理方法、装置、设备及可读存储介质 |
| CN119418350B (zh) * | 2024-11-08 | 2025-10-28 | 长安汽车金融有限公司 | 一种收据识别方法、装置、设备及介质 |
| CN120012832B (zh) * | 2025-04-22 | 2025-08-22 | 杭州海康威视数字技术股份有限公司 | 视觉问答多模态大模型建立方法和装置 |
Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20120219209A1 (en) * | 2011-02-25 | 2012-08-30 | Microsoft Corporation | Image Labeling with Global Parameters |
| CN109635135A (zh) * | 2018-11-30 | 2019-04-16 | Oppo广东移动通信有限公司 | 图像索引生成方法、装置、终端及存储介质 |
| US20220114361A1 (en) * | 2020-10-14 | 2022-04-14 | Adobe Inc. | Multi-word concept tagging for images using short text decoder |
| CN114998908A (zh) * | 2022-05-23 | 2022-09-02 | 北京百度网讯科技有限公司 | 样本图像标注、模型训练方法、装置、设备以及存储介质 |
| CN117078995A (zh) * | 2023-06-16 | 2023-11-17 | 之江实验室 | 一种文本生成的方法、装置、存储介质及电子设备 |
| CN117495949A (zh) * | 2023-12-19 | 2024-02-02 | 北京百度网讯科技有限公司 | 图像标注方法、装置、设备以及存储介质 |
| CN117636352A (zh) * | 2023-12-12 | 2024-03-01 | 阿波罗智联(北京)科技有限公司 | 图像的标注方法、装置、电子设备和介质 |
| CN118097694A (zh) * | 2024-03-08 | 2024-05-28 | 北京百度网讯科技有限公司 | 图像标注数据生成、模型训练方法、装置、设备及介质 |
-
2024
- 2024-03-08 CN CN202410268967.1A patent/CN118097694A/zh active Pending
- 2024-06-20 WO PCT/CN2024/100320 patent/WO2025185014A1/zh active Pending
Patent Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20120219209A1 (en) * | 2011-02-25 | 2012-08-30 | Microsoft Corporation | Image Labeling with Global Parameters |
| CN109635135A (zh) * | 2018-11-30 | 2019-04-16 | Oppo广东移动通信有限公司 | 图像索引生成方法、装置、终端及存储介质 |
| US20220114361A1 (en) * | 2020-10-14 | 2022-04-14 | Adobe Inc. | Multi-word concept tagging for images using short text decoder |
| CN114998908A (zh) * | 2022-05-23 | 2022-09-02 | 北京百度网讯科技有限公司 | 样本图像标注、模型训练方法、装置、设备以及存储介质 |
| CN117078995A (zh) * | 2023-06-16 | 2023-11-17 | 之江实验室 | 一种文本生成的方法、装置、存储介质及电子设备 |
| CN117636352A (zh) * | 2023-12-12 | 2024-03-01 | 阿波罗智联(北京)科技有限公司 | 图像的标注方法、装置、电子设备和介质 |
| CN117495949A (zh) * | 2023-12-19 | 2024-02-02 | 北京百度网讯科技有限公司 | 图像标注方法、装置、设备以及存储介质 |
| CN118097694A (zh) * | 2024-03-08 | 2024-05-28 | 北京百度网讯科技有限公司 | 图像标注数据生成、模型训练方法、装置、设备及介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| WO2025185014A8 (zh) | 2025-10-02 |
| CN118097694A (zh) | 2024-05-28 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN117407507B (zh) | 基于大语言模型的事件处理方法、装置、设备及介质 | |
| WO2025185014A1 (zh) | 图像标注数据生成、模型训练方法、装置、设备及介质 | |
| CN110446063B (zh) | 视频封面的生成方法、装置及电子设备 | |
| CN112100438B (zh) | 一种标签抽取方法、设备及计算机可读存储介质 | |
| CN111107422B (zh) | 图像处理方法及装置、电子设备和计算机可读存储介质 | |
| CN103544266B (zh) | 一种搜索建议词生成的方法以及装置 | |
| US11929100B2 (en) | Video generation method, apparatus, electronic device, storage medium and program product | |
| CN111046272A (zh) | 一种基于医疗知识图谱的智能问答系统 | |
| CN114090748B (zh) | 问答结果显示方法、装置、设备及存储介质 | |
| US12596737B2 (en) | Method for generating user interest profile, electronic device and storage medium | |
| CN115022668B (zh) | 基于直播的视频生成方法和装置、设备、介质 | |
| CN114067343A (zh) | 一种数据集的构建方法、模型训练方法和对应装置 | |
| CN120783147A (zh) | 基于优化算法的视觉-语言模型图文对精准评测数据构建方法 | |
| CN113301382A (zh) | 视频处理方法、设备、介质及程序产品 | |
| WO2025129964A1 (zh) | 图像标注方法、装置、设备以及存储介质 | |
| CN116246287A (zh) | 目标对象识别方法、训练方法、装置以及存储介质 | |
| CN110162651B (zh) | 基于语义内容摘要的新闻内容图文不符鉴别系统及鉴别方法 | |
| WO2025246204A1 (zh) | 基于大模型的问答方法及电子设备 | |
| CN117579888A (zh) | 视频字幕提取方法、装置、设备及存储介质 | |
| CN116010545B (zh) | 一种数据处理方法、装置及设备 | |
| CN115934937A (zh) | 文本分类模型的训练方法、文本分类方法及装置 | |
| CN118132699B (zh) | 回复信息生成方法、装置、客户端和存储介质 | |
| CN118051659A (zh) | 信息卡片生成方法及装置 | |
| CN117853109A (zh) | 一种基于多模态知识图谱的风险预测系统 | |
| CN117351116A (zh) | 图像生成方法、装置、电子设备以及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24927947 Country of ref document: EP Kind code of ref document: A1 |