WO2026007670A1 - 图像处理方法及装置 - Google Patents

图像处理方法及装置

Info

Publication number
WO2026007670A1
WO2026007670A1 PCT/CN2025/100832 CN2025100832W WO2026007670A1 WO 2026007670 A1 WO2026007670 A1 WO 2026007670A1 CN 2025100832 W CN2025100832 W CN 2025100832W WO 2026007670 A1 WO2026007670 A1 WO 2026007670A1
Authority
WO
WIPO (PCT)
Prior art keywords
image
information
processed
description information
structured
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2025/100832
Other languages
English (en)
French (fr)
Inventor
阳展韬
冯睿蠡
颜科宇
王志才
张晗
肖杰
吴平禹
朱凯
陈霁璇
谢晨伟
毛超杰
刘宇
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Alibaba China Co Ltd
Original Assignee
Alibaba China Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Alibaba China Co Ltd filed Critical Alibaba China Co Ltd
Publication of WO2026007670A1 publication Critical patent/WO2026007670A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/40Extraction of image or video features
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/10Text processing
    • G06F40/166Editing, e.g. inserting or deleting
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T11/00Two-dimensional [2D] image generation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/764Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
    • G06V10/765Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects using rules for classification or partitioning the feature space

Definitions

  • This disclosure relates to the field of computer technology, and in particular to an image processing method.
  • Training text-image multimodal models requires a large amount of image data and corresponding text data.
  • high-quality text descriptions significantly aid downstream multimodal tasks in understanding image content.
  • high-quality text descriptions are often lengthy and complex.
  • Downstream multimodal models lacking strong language understanding capabilities typically struggle to comprehend such complex textual information. Therefore, a method is urgently needed that provides rich yet easily comprehensible image descriptions for downstream multimodal models.
  • embodiments of this disclosure provide an image processing method.
  • One or more embodiments of this disclosure also relate to an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
  • an image processing method comprising:
  • the image to be processed is input into the image information extraction model to obtain the initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object;
  • the object location information corresponding to each object is determined based on the image to be processed and the initial structured description information;
  • the structured image description information corresponding to the image to be processed is generated based on the initial structured description information and the object location information corresponding to each object.
  • the image to be processed is input into the image information extraction model to obtain the initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object;
  • the object location information corresponding to each object is determined based on the image to be processed and the initial structured description information;
  • the structured image description information corresponding to the image to be processed is generated based on the initial structured description information and the object location information corresponding to each object;
  • an image processing apparatus comprising:
  • the acquisition module is configured to acquire the image to be processed.
  • the information extraction module is configured to input the image to be processed into the image information extraction model to obtain the initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object;
  • the determination module is configured to determine the object location information corresponding to each object based on the image to be processed and the initial structured description information;
  • the memory is used to store computer programs/instructions
  • the processor is used to execute the computer programs/instructions, which, when executed by the processor, implement the steps of the above-described image processing method.
  • a computer-readable storage medium stores a computer program/instructions that, when executed by a processor, implement the steps of the image processing method described above.
  • the image processing method provided in this application utilizes the understanding capabilities of a large language model to analyze the image to be processed. Based on preset structured information, it generates initial structured description information corresponding to the objects to be processed.
  • This initial structured description information includes at least one object and object description information for each object, used to describe the image to be processed, facilitating subsequent use of information within the image.
  • the method locates each object in the image, obtaining object location information for each object. This further enriches the structured image description information of the image to be processed, enabling faster and more accurate retrieval of objects and object description information from the image to be processed, according to the needs of downstream tasks.
  • FIG. 1 is a flowchart of an image processing method provided in an embodiment of this disclosure
  • Figure 2 is a schematic diagram of the composition of initial structured description information provided in an embodiment of this disclosure
  • Figure 3 is a schematic diagram of generating sample structured description information according to an embodiment of this disclosure.
  • Figure 5 is a processing framework diagram of an image processing method provided in an embodiment of this disclosure.
  • Figure 6 is a schematic flowchart of an image processing method applied to a cloud-side device according to an embodiment of this disclosure
  • Figure 7 is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of the present disclosure.
  • Figure 8 is a schematic diagram of an image processing device applied to a cloud-side device according to an embodiment of the present disclosure
  • Figure 9 is an architecture diagram of an image processing system provided in an embodiment of this disclosure.
  • Figure 10 is a structural block diagram of a computing device provided in an embodiment of this disclosure.
  • first, second, etc. may be used to describe various information in one or more embodiments of this disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this disclosure, and similarly, second may also be referred to as first.
  • word “if” as used herein may be interpreted as “when”, “in response to a determination”, or “when...”.
  • the user information including but not limited to user device information, user personal information, etc.
  • data including but not limited to data used for analysis, data stored, data displayed, etc.
  • the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
  • a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters.
  • a large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters.
  • Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.
  • NLP Natural Language Processing
  • VQA Visual Question Answering
  • IC Image Captioning
  • Image Generation computer vision tasks
  • text-based sentiment classification text summarization
  • machine translation text-based sentiment classification, text summarization, and machine translation.
  • the main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
  • Multi-modal small models These models require input of information from multiple modalities (such as images and text) simultaneously, but lack the ability of large language models to understand textual information.
  • Text-image multimodal models require images and their corresponding text as input. Typically, images contain a wealth of information, necessitating lengthy text descriptions for accurate representation. Furthermore, the richer and more accurate the image description, the better it is for completing downstream multimodal tasks. However, most small multimodal models lack the text understanding capabilities of larger language models; even with sufficiently rich text information, they are inapplicable. Therefore, a method is needed to provide rich image descriptions that are easily understood and used by small multimodal models.
  • This disclosure also relates to an image processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
  • Figure 1 shows a flowchart of an image processing method provided according to an embodiment of the present disclosure, which specifically includes the following steps.
  • Step 102 Obtain the image to be processed.
  • the image to be processed can be understood as an image that requires text description.
  • the purpose of the method provided in this disclosure is to provide a text description for the image to be processed, thereby obtaining image description information corresponding to the image to be processed.
  • the image to be processed needs to be acquired.
  • the image to be processed can be acquired from a specified storage location, or it can be uploaded by the user through a specified interactive interface.
  • the specific implementation method for acquiring the image to be processed is not limited in the method provided in this disclosure; the actual application shall prevail.
  • Step 104 Input the image to be processed into the image information extraction model to obtain the initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object.
  • the image information extraction model can be understood as a model that understands the image to be processed and generates corresponding descriptive information.
  • the image information extraction model can be understood as a multimodal large language model, which has the ability to understand and analyze multimodal data.
  • the initial structured description information can be understood as the model extraction result output by the image information extraction model. It is not the final description information for the image to be processed.
  • the initial structured description information is structured description information, including at least one object and object description information for each object.
  • An object can be understood as an object in the image to be processed; for example, elements such as streetlights, people, cars, and the sky in the image can all be understood as objects.
  • the object description information can be understood as the descriptive content of the object.
  • the initial structured description information could specifically include:
  • Building 1 (Inanimate object; background; the building features a unique angular design and interior-lit windows. Color information: black and yellow light.)
  • Building 2 Inanimate object; background; This building is tall with many windows, some of which are lit, and has a cylindrical shape. Color information: white and yellow lit windows.
  • Street Inanimate objects; foreground/background; the street shows multiple lanes, vehicles, and wet ground reflecting lights. Color information: dark gray asphalt, white road markings.)
  • Car 1 Inanimate object; Foreground; A car with its headlights on is driving on a street. Type of vehicle: Sedan. Color: Black.
  • Car 2 inanimate object; foreground; another car following Car 1, also with headlights on, appears to be a hatchback. Color information: silver.
  • the objects and object descriptions in the initial structured description information provided in this embodiment can be displayed in the form of an object list. Users can more intuitively and accurately understand the objects and object descriptions contained in the image to be processed by using the list-style object and object description information.
  • obtaining the initial structured description information output by the image information extraction model includes:
  • the initial graph structure description information includes at least one object node and object description information corresponding to each object node.
  • the image information extraction model outputs initial graph structure description information based on the input image to be processed; specifically, the initial structured description information can be an initial graph structure description information.
  • a graph structure is a non-linear data structure composed of nodes and edges.
  • the initial graph structure description information includes at least one object node and object description information corresponding to each object node.
  • the initial structured description information also includes the overall image description information of the image to be processed and the relationships between the objects.
  • the initial structured description information also includes overall image description information for the image to be processed, which is used to represent the content of the image as a whole.
  • the initial structured description information also includes the relationships between various objects.
  • the overall image description information could include:
  • Style This image is a photograph with a realistic style.
  • the subject of the image is city traffic at dusk.
  • the background consists of a cloudy sky that gradients from deep blue at the top to a lighter blue at the bottom.
  • To the left is a large, modern building called "***”, characterized by its unique angular design and internally lit windows.
  • To the right is a tall building with numerous windows, some emitting a warm glow.
  • Ambient lighting indicates that it is dawn or dusk, and artificial light is beginning to have a significant impact on the scene.
  • the foreground depiction shows a city street scene with multiple lanes.
  • the street is busy, and vehicles with their headlights on indicate poor lighting conditions.
  • the vehicles vary in size and shape, and there is a mix of private and commercial vehicles.
  • the sidewalks along the street appear particularly clean under the wet lights, a typical feature of a city environment at night.
  • the initial structured description information also includes the relationships between the objects, specifically:
  • Building 1 is located on the left side of the street;
  • Building 2 is located on the right side of the street;
  • Car 1 is driving on the street
  • the initial structured description information includes at least an object list, which can be understood as at least one object and object description information for each pair of objects. Furthermore, the initial structured description information may also include overall image description information of the image to be processed, and/or the relationships between the objects.
  • initial structured description information is generated.
  • the initial structured description information includes three parts: the first part is the overall image description information, the second part is the object list, and the third part is the association relationship between each object.
  • the first part is the overall image description information, which includes the overall description information of the image to be processed, such as image style, image subject, image background description, image foreground description, and so on.
  • the second part is the object list, which includes at least one object in the image to be processed and object description information for each object.
  • the third part is the relationship, which includes the relationships between objects in the image to be processed.
  • the initial structured description information is a special format of graph structure description information.
  • the content of model understanding and analysis can be displayed in the form of initial structured description information.
  • the objects and object description information in the initial structured description information can be understood as nodes in the graph structure and node description information.
  • the image information extraction model is a pre-trained large language model, which is trained to output initial structured description information in a fixed format based on the input image.
  • the image information extraction model is trained through the following steps:
  • the sample image is input into the image information extraction model to obtain the predicted structured description information output by the image information extraction model.
  • the model loss value is calculated based on the sample structured description information and the prediction structured description information.
  • the model parameters of the image information extraction model are adjusted based on the model loss value, and the image information extraction model is trained continuously until the model training stops.
  • the training method for the image information extraction model uses supervised training, which includes multiple training sample pairs.
  • Each training sample pair includes a sample image and corresponding sample structured description information.
  • the sample image can be understood as a sample used to train the image information extraction model.
  • the sample structured description information can be understood as structured description information corresponding to the sample image, which includes at least one sample object in the sample image and sample object description information corresponding to each sample object.
  • a template is provided for the image information extraction model based on a preset structured description template, informing the image information extraction model that it needs to output the corresponding structured description information according to the format of the preset structured description template.
  • obtaining a sample image and the corresponding sample structured description information includes:
  • the preset structured description template can be understood as a template pre-set in one or more embodiments of this disclosure, which is used to inform the multimodal large language model that it needs to output the corresponding structured description information according to the format of the preset structured description template.
  • the preset structured description template includes at least one type of preset identifier characters, which represent the information type in the preset structured description template.
  • preset identifier characters can be “%%”, “&&", “ ⁇ >”, “()", ";”, “[]”, etc.
  • Each preset identifier character identifies a type of information in the preset structured template. For example, “%%” is used to distinguish main headings, “&&” is used to distinguish subheadings, “ ⁇ >” is used to identify nouns, "()” is used to identify object attributes, ";” is used to separate object attributes, "[]” is used to identify relationships, etc.
  • the preset structured description template can be divided into predefined information parts, thereby defining the preset string format of the preset structured description template.
  • the method provided in this embodiment uses the ICL method, which provides an instance of the target information type only at the location of the target information type, enabling the multimodal large language model to output the corresponding structured description information in a prescribed format. This shortens the length of the text input to the multimodal large language model, and allows the multimodal large language model to output structured description information using only a single-turn dialogue.
  • Predicting structured descriptive information can be understood as the descriptive information output by an untrained image information extraction model based on input sample images. Its data structure corresponds to a predefined character format.
  • the model training stopping condition for the initial image processing model includes:
  • the training stopping condition for the image information extraction model can be set to the number of training rounds reaching a preset number, such as 10 rounds. When the model has reached 10 training rounds, the training stopping condition is met.
  • the initial structured description information includes at least one object and object description information corresponding to each object.
  • the object to be processed is the object being operated on, and the object description information to be processed is the object description information corresponding to the object to be processed. Specifically, the object to be processed and its corresponding object description information can be selected from the object list.
  • the object to be processed can be any one of the objects.
  • the processing between objects can be parallel or sequential. In the embodiments provided in this disclosure, the processing order between objects is not limited.
  • At least one candidate object bounding box can be determined in the image to be processed based on the object.
  • the candidate object bounding box is a bounding box generated in the image during the detection and localization task, which can mark the position information of the object in the image.
  • candidate object bounding boxes can be generated using a localization model, and at least one candidate object bounding box can be determined in the image to be processed based on the object to be processed, including:
  • the object to be processed and the image to be processed are input into the localization model to obtain at least one candidate object bounding box determined by the localization model in the image to be processed based on the object to be processed.
  • a localization model can be understood as a model that locates and marks objects in an image based on input information. For example, if the object to be processed is "man," the localization model inputs the object and the image to be processed, and then marks candidate object bounding boxes related to "man" in the image.
  • the localization model can only mark at least one candidate bounding box in the image to be processed based on the input object to be processed.
  • the localization model has difficulty distinguishing different individuals of the same category. For example, when the object to be processed is "man", the localization model can only mark multiple "man” objects in the image to be processed, but cannot distinguish the differences between the various "man” objects. Therefore, the method provided in this embodiment of the present disclosure needs to further filter the candidate object bounding boxes determined by the localization model to obtain the final target object bounding box.
  • determining at least one candidate object bounding box in the image to be processed based on the object to be processed includes:
  • the object to be processed and the image to be processed are input into the localization model to obtain at least one initial candidate object bounding box determined by the localization model in the image to be processed based on the object to be processed.
  • the object to be processed and each initial candidate object bounding box are input into a multimodal large language model to obtain at least one candidate object bounding box output by the multimodal large language model, wherein the multimodal large language model filters each initial candidate object bounding box according to the object to be processed.
  • a multimodal large language model can be used to filter the candidate object bounding boxes.
  • the object to be processed and the image to be processed are first input into the localization model.
  • the localization model outputs at least one initial candidate object bounding box.
  • the object to be processed and each initial candidate object bounding box are input into the multimodal large language model for filtering.
  • Initial candidate bounding boxes that are obviously inconsistent with the object to be processed are deleted, and candidate bounding boxes that match the object to be processed are retained.
  • S1066 Determine the target object marker box in at least one candidate object marker box according to the description information of the object to be processed, and determine the marker box position information of the target object marker box.
  • the object description information of the object to be processed is used to select the target object marker box from at least one candidate object marker box.
  • the object description information is a detailed description of the object to be processed. Through this object description information, the selection of the target object marker box from the candidate object marker boxes can be assisted, and the marker box position information of the target object marker box can be determined.
  • determining the target object marker box in at least one candidate object marker box based on the description information of the object to be processed includes:
  • the candidate object bounding box that meets the matching similarity criteria is selected as the target object bounding box.
  • the preset criteria can be understood as the conditions for selecting the target object bounding box from multiple candidate bounding boxes; for example, the preset criteria could be that the candidate object bounding box with the highest matching similarity is the target object bounding box.
  • the corresponding bounding box position information can be determined based on the target object bounding box.
  • the image to be processed is input into the image information extraction model for extraction, and the initial structured description information output by the image information extraction model is obtained.
  • the initial structured description information includes an object list, which includes at least one object included in the image to be processed and object description information of each object.
  • the description information of the object to be processed is "wearing a hat”. Since the localization model can only process object names, input the object to be processed, "person”, into the localization model. The localization model locates the object to be processed, "person”, in the image to be processed and generates four initial candidate bounding boxes.
  • the three candidate bounding boxes are matched with the description information "wearing a hat" of the object to be processed.
  • the candidate bounding box with the highest matching information is selected as the target bounding box, and the target bounding box is marked in the image to be processed.
  • Step 108 Generate structured image description information corresponding to the image to be processed based on the initial structured description information and the object location information corresponding to each object.
  • the final structured image description information generated includes not only the objects and their descriptions, but also the location information corresponding to each object.
  • the structured image description information which includes the location information of each object, can describe the image to be processed from multiple dimensions. This facilitates downstream tasks in selecting appropriate sub-information from the structured image description information according to different task requirements and executing subsequent downstream tasks.
  • the method further includes:
  • image descriptor information is extracted from the structured image description information, and the image descriptor information is sent to the downstream task.
  • Downstream tasks can be understood as other image processing tasks that rely on structured image description information. Different downstream tasks can perform different operations on the image to be processed, and therefore, different downstream tasks obtain different information from the image to be processed.
  • the structured image description information can be filtered according to the data request information in the data request instruction to extract the image descriptor information corresponding to the data request information, and the image descriptor information is sent to the downstream task, so that the multimodal model of the downstream task can perform the corresponding downstream task according to the image descriptor information.
  • the image processing method provided in this disclosure includes: acquiring an image to be processed; inputting the image to be processed into an image information extraction model to obtain initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object; determining object location information corresponding to each object based on the image to be processed and the initial structured description information; and generating structured image description information corresponding to the image to be processed based on the initial structured description information and the object location information corresponding to each object.
  • An image information extraction model generates initial structured description information for the image to be processed according to a preset format. Based on at least one object and its corresponding object description information in the initial structured description information, the object location information of each object in the image to be processed can be determined. Finally, based on the initial structured description information and the object location information, the final structured image description information is determined.
  • This specific method of structured image description describes the content of the image, facilitating the subsequent extraction of relevant information from the structured image description information.
  • the image information extraction model is trained based on a specific structured image description method.
  • ICL method only corresponding target information type instances are provided for the target information type, enabling the image information extraction model to learn the preset string format well, and thus output initial structured description information according to the preset string format. This reduces training costs and simplifies the learning process.
  • each object is located in the image to be processed, determining its positional information. This allows for a more detailed depiction of the objects in the initial structured description information, further improving the ease with which subsequent downstream tasks can extract information from the structured image description information.
  • Figure 5 illustrates a processing framework diagram of an image processing method provided in an embodiment of this disclosure.
  • a preset structured description template is designed using a pre-set string format.
  • the preset structured description template is used to generate sample structured description information through a multimodal large language model.
  • the sample structured description information and sample images are then used to train an image information extraction model, enabling the image information extraction model to output structured description information based on the input image.
  • the initial structured description information here includes at least one object and the object description information corresponding to each object, but does not include the position information of each object in the image to be processed.
  • the structured image description information includes at least one object, its description information, and its position information.
  • Figure 6 shows a flowchart illustrating an image processing method applied to a cloud-side device according to an embodiment of this disclosure.
  • the method, applied to a cloud-side device specifically includes:
  • Step 602 The image to be processed is sent by the receiving end device.
  • Step 604 Input the image to be processed into the image information extraction model to obtain the initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object.
  • Step 606 Determine the object location information corresponding to each object based on the image to be processed and the initial structured description information.
  • Step 608 Generate structured image description information corresponding to the image to be processed based on the initial structured description information and the object location information corresponding to each object.
  • Step 610 Send the structured image description information to the end device.
  • the method further includes:
  • the receiving end device sends information processing instructions for the structured image description information
  • the structured image description information is processed to obtain image description adjustment information
  • the image description adjustment information is sent to the end-side device.
  • the image information extraction model is a multimodal large language model, it requires considerable computing resources to run, and the edge device may not have the corresponding processing capabilities. Therefore, the image information extraction model can be deployed on a cloud-side device, that is, the image processing procedure provided in this embodiment is implemented on the cloud-side device. After obtaining the structured image description information generated by the image information extraction model, the cloud-side device can also send the structured image description information to the edge device.
  • the image processing method provided in this disclosure generates initial structured description information corresponding to the image to be processed according to a preset format using an image information extraction model. Based on at least one object in the initial structured description information and the object description information corresponding to each object, the object location information of each object in the image to be processed can be determined. Finally, based on the initial structured description information and the object location information of each object, the final structured image description information is determined.
  • the image information extraction model is trained based on a specific structured image description method.
  • ICL method only corresponding target information type instances are provided for the target information type, enabling the image information extraction model to learn the preset string format well, and thus output initial structured description information according to the preset string format. This reduces training costs and simplifies the learning process.
  • each object is located in the image to be processed, determining its positional information. This allows for a more detailed depiction of the objects in the initial structured description information, further improving the ease with which subsequent downstream tasks can extract information from the structured image description information.
  • FIG7 shows a schematic structural diagram of an image processing apparatus provided in one embodiment of this disclosure. As shown in FIG7, the apparatus includes:
  • the acquisition module 702 is configured to acquire the image to be processed
  • the information extraction module 704 is configured to input the image to be processed into the image information extraction model to obtain the initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object;
  • the determination module 706 is configured to determine the object location information corresponding to each object based on the image to be processed and the initial structured description information;
  • the generation module 708 is configured to generate structured image description information corresponding to the image to be processed based on the initial structured description information and the object position information corresponding to each object.
  • the information extraction module 704 is configured as follows:
  • the initial graph structure description information includes at least one object node and object description information corresponding to each object node.
  • the device further includes a training module configured to:
  • the sample image is input into the image information extraction model to obtain the predicted structured description information output by the image information extraction model;
  • the model loss value is calculated based on the sample structured description information and the prediction structured description information
  • the model parameters of the image information extraction model are adjusted based on the model loss value, and the image information extraction model is trained continuously until the model training stops.
  • the training module is further configured to:
  • the sample image and the preset structured description template are input into the multimodal large language model to obtain the sample structured description information output by the multimodal large language model.
  • the preset structured description template includes at least one type of preset identifier characters, which represent the information type in the preset structured description template.
  • the preset structured description template includes an instance of the target information type corresponding to the target information type.
  • the determining module 706 is further configured to:
  • the object to be processed from the initial structured description information and the description information of the object to be processed corresponding to the object to be processed, wherein the object to be processed is any one of the at least one object;
  • At least one candidate object bounding box is determined in the image to be processed
  • a target object marker box is determined in at least one candidate object marker box, and the marker box position information of the target object marker box is determined;
  • the position information of the marker box is determined as the position information of the object corresponding to the object to be processed.
  • the object to be processed and the image to be processed are input into the localization model to obtain at least one candidate object bounding box determined by the localization model in the image to be processed based on the object to be processed.
  • the determining module 706 is further configured to:
  • the object to be processed and the image to be processed are input into the localization model to obtain at least one initial candidate object bounding box determined by the localization model in the image to be processed based on the object to be processed.
  • the object to be processed and each initial candidate object bounding box are input into a multimodal large language model to obtain at least one candidate object bounding box output by the multimodal large language model, wherein the multimodal large language model filters each initial candidate object bounding box according to the object to be processed.
  • the determining module 706 is further configured to:
  • Candidate object bounding boxes whose matching similarity meets preset conditions are identified as target object bounding boxes.
  • the device further includes an information extraction module, configured to:
  • image descriptor information is extracted from the structured image description information, and the image descriptor information is sent to the downstream task.
  • the image processing apparatus includes: acquiring an image to be processed; inputting the image to be processed into an image information extraction model to obtain initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object; determining object location information corresponding to each object based on the image to be processed and the initial structured description information; and generating structured image description information corresponding to the image to be processed based on the initial structured description information and the object location information corresponding to each object.
  • Figure 8 shows a schematic diagram of an image processing apparatus applied to a cloud-side device according to an embodiment of this disclosure. As shown in Figure 8, the apparatus is applied to a cloud-side device and includes:
  • the information extraction module 804 is configured to input the image to be processed into the image information extraction model to obtain the initial structured description information output by the image information extraction model, wherein the initial structured description information includes at least one object and object description information of each object;
  • the determination module 806 is configured to determine the object location information corresponding to each object based on the image to be processed and the initial structured description information;
  • the device further includes an adjustment module configured to:
  • the receiving end device sends information processing instructions for the structured image description information
  • the image information extraction model is trained based on a specific structured image description method.
  • ICL method only corresponding target information type instances are provided for the target information type, enabling the image information extraction model to learn the preset string format well, and thus output initial structured description information according to the preset string format. This reduces training costs and simplifies the learning process.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Physics & Mathematics (AREA)
  • Software Systems (AREA)
  • Artificial Intelligence (AREA)
  • Medical Informatics (AREA)
  • General Health & Medical Sciences (AREA)
  • Computing Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Evolutionary Computation (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Computational Linguistics (AREA)
  • Mathematical Physics (AREA)
  • Databases & Information Systems (AREA)
  • Data Mining & Analysis (AREA)
  • Processing Or Creating Images (AREA)

Abstract

本公开实施例提供图像处理方法及装置,其中图像处理方法包括:获取待处理图像;将待处理图像输入至图像信息提取模型,获得图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息;根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息;根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息,通过本方法,定义了一种结构化的图像描述信息,使得图像信息提取模型可以根据该结构化格式获得相应的图像描述信息,便于下游任务根据需求更加快捷准确的从结构化图像描述信息中获取到对应的信息。

Description

图像处理方法及装置
本公开要求申请号为202410904423.X的中国专利申请的优先权,该中国专利申请于2024年07月05日提交中国专利局,申请名称为“图像处理方法及装置”,其全部内容通过引用结合在本公开中。
技术领域
本公开实施例涉及计算机技术领域,特别涉及一种图像处理方法。
背景技术
需要大量的图像,以及对应的文本数据对文本-图像多模态模型进行训练。一般来说,高质量的文本描述会显著帮助多模态下游任务理解图像内容。然而,由于图片信息的丰富性,高质量的文本描述通常会很长很复杂。缺乏大语言理解能力的下游多模态模型,通常很难理解这样复杂的文本信息。因此,亟需一种方法,能提供丰富但容易被下游多模态模型理解的图片描述信息。
发明内容
有鉴于此,本公开实施例提供了一种图像处理方法。本公开一个或者多个实施例同时涉及一种图像处理装置,一种计算设备,一种计算机可读存储介质以及一种计算机程序产品,以解决现有技术中存在的技术缺陷。
根据本公开实施例的第一方面,提供了一种图像处理方法,包括:
获取待处理图像;
将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息;
根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息;
根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息。
根据本公开实施例的第二方面,提供了一种图像处理方法,应用于云侧设备,包括:
接收端侧设备发送的待处理图像;
将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息;
根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息;
根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息;
将所述结构化图像描述信息发送至所述端侧设备。
根据本公开实施例的第三方面,提供了一种图像处理装置,包括:
获取模块,被配置为获取待处理图像;
信息提取模块,被配置为将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息;
确定模块,被配置为根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息;
生成模块,被配置为根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息。
根据本公开实施例的第四方面,提供了一种计算设备,包括:
存储器和处理器;
所述存储器用于存储计算机程序/指令,所述处理器用于执行所述计算机程序/指令,该计算机程序/指令被处理器执行时实现上述图像处理方法的步骤。
根据本公开实施例的第五方面,提供了一种计算机可读存储介质,其存储有计算机程序/指令,该计算机程序/指令被处理器执行时实现上述图像处理方法的步骤。
根据本公开实施例的第六方面,提供了一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现上述图像处理方法的步骤。
本公开一个实施例提供的图像处理方法,包括获取待处理图像;将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息;根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息;根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息。
通过本申请实施例提供的图像处理方法,利用大语言模型的理解能力对待处理图像进行分析,根据预设结构化信息生成待处理对象对应的初始结构化描述信息,在初始结构化描述信息中包括至少一个对象和各对象对应的对象描述信息,用于对待处理图像进行描述,便于后续对待处理图像中的信息进行利用。另外,通过初始结构化描述信息和待处理图像结合对待处理图像中的各对象进行定位,获得各对象的对象位置信息,从而进一步丰富了待处理图像的结构化图像描述信息,可以根据下游任务的需求获取更加快捷准确的从待处理图像中获取到对象和对象描述信息。
附图说明
图1是本公开一个实施例提供的一种图像处理方法的流程图;
图2是本公开一个实施例提供的初始结构化描述信息的组成示意图;
图3是本公开一个实施例提供的生成样本结构化描述信息的示意图;
图4是本公开一个实施例提供的对象定位的示意图;
图5是本公开一个实施例提供的一种图像处理方法的处理框架图;
图6是本公开一个实施例提供的一种应用于云侧设备的图像处理方法的流程示意图;
图7是本公开一个实施例提供的一种图像处理装置的结构示意图;
图8是本公开一个实施例提供的一种应用于云侧设备的图像处理装置的结构示意图;
图9是本公开一个实施例提供的一种图像处理系统的架构图;
图10是本公开一个实施例提供的一种计算设备的结构框图。
具体实施方式
在下面的描述中阐述了很多具体细节以便于充分理解本公开。但是本公开能够以很多不同于在此描述的其它方式来实施,本领域技术人员可以在不违背本公开内涵的情况下做类似推广,因此本公开不受下面公开的具体实施的限制。
在本公开一个或多个实施例中使用的术语是仅仅出于描述特定实施例的目的,而非旨在限制本公开一个或多个实施例。在本公开一个或多个实施例和所附权利要求书中所使用的单数形式的“一种”、“所述”和“该”也旨在包括多数形式,除非上下文清楚地表示其他含义。还应当理解,本公开一个或多个实施例中使用的术语“和/或”是指并包含一个或多个相关联的列出项目的任何或所有可能组合。
应当理解,尽管在本公开一个或多个实施例中可能采用术语第一、第二等来描述各种信息,但这些信息不应限于这些术语。这些术语仅用来将同一类型的信息彼此区分开。例如,在不脱离本公开一个或多个实施例范围的情况下,第一也可以被称为第二,类似地,第二也可以被称为第一。取决于语境,如在此所使用的词语“如果”可以被解释成为“在……时”或“当……时”或“响应于确定”。
需要说明的是,本公开所涉及的用户信息(包括但不限于用户设备信息、用户个人信息等)和数据(包括但不限于用于分析的数据、存储的数据、展示的数据等),均为经用户授权或者经过各方充分授权的信息和数据,并且相关数据的收集、使用和处理需要遵守相关地区的相关法律法规和标准,并提供有相应的操作入口,供用户选择授权或者拒绝。
本公开一个或多个实施例中,大模型是指具有大规模模型参数的深度学习模型,通常包含上亿、上百亿、上千亿、上万亿甚至十万亿以上的模型参数。大模型又可以称为基石模型/基础模型(Foundation Model),通过大规模无标注的语料进行大模型的预训练,产出亿级以上参数的预训练模型,这种模型能适应广泛的下游任务,模型具有较好的泛化能力,例如大规模语言模型(Large Language Model,LLM)、多模态预训练模型(multi-modal pre-training model)等。
大模型在实际应用时,仅需少量样本对预训练模型进行微调即可应用于不同的任务中,大模型可以广泛应用于自然语言处理(Natural Language Processing,简称NLP)、计算机视觉等领域,具体可以应用于如视觉问答(Visual Question Answering,简称VQA)、图像描述(Image Caption,简称IC)、图像生成等计算机视觉领域任务,以及基于文本的情感分类、文本摘要生成、机器翻译等自然语言处理领域任务,大模型主要的应用场景包括数字助理、智能机器人、搜索、在线教育、办公软件、电子商务、智能设计等。
首先,对本公开一个或多个实施例涉及的名词术语进行解释。
多模态小模型(Multi-modal):需要同时输入多个模态的信息(如图片和文本信息),但缺乏大语言模型的理解文本信息能力的模型。
文本-图像多模态模型需要图像,以及图像对应的文本作为输入。通常情况下,图像中包含了大量的信息,需要大段文本才能精确的描述图像。而且图像的描述信息越丰富越准确,越有益于完成多模态下游任务。然而,大多数多模态小模型,缺乏像大语言模型一样的文本理解能力,即使文本信息够丰富,多模态小模型也无法应用。因此,需要提供一种方法,能提供内容丰富,也容易被多模态小模型理解使用的图像描述信息。
基于此,在本公开中,提供了一种图像处理方法,本公开同时涉及一种图像处理装置,一种计算设备,一种计算机可读存储介质以及一种计算机程序产品,在下面的实施例中逐一进行详细说明。
参见图1,图1示出了根据本公开一个实施例提供的一种图像处理方法的流程图,具体包括以下步骤。
步骤102:获取待处理图像。
其中,待处理图像可以理解为需要进行文本描述的图像。在本公开实施例提供的方法中,目的是为待处理图像进行文本描述,获得待处理图像对应的图像描述信息。
在本公开实施例提供的方法中,需要为待处理图像生成对应的图像描述信息。需要首先获取到待处理图像。获取到待处理图像的方式可以是从指定的存储位置获取到待处理图像,也可以是用户在指定的交互界面上传该待处理图像。在本公开实施例提供的方法中,对获取待处理图像的具体实现方式不做限定,以实际应用为准。
获取待处理图像,为后续对待处理图像进行图像描述提供了数据基础。
步骤104:将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息。
其中,图像信息提取模型可以理解为理解待处理图像,并生成器对应描述信息的模型,图像信息提取模型在本公开提供的一个或多个实施例中可以理解为多模态大语言模型,其具备理解、分析多模态数据的能力。
初始结构化描述信息可以理解为图像信息提取模型输出的模型提取结果,初始结构化描述信息并非针对待处理图像的最终描述信息。初始结构化描述信息是具有结构化的描述信息,在初始结构化描述信息中包括有至少一个对象和各对象的对象描述信息,对象可以理解为待处理图像中的对象,例如,待处理图像中的路灯、人、车、天空等元素,均可以理解为待处理图像中的对象,对象描述信息可以理解为对象的描述内容。
例如,以某个待处理图片中包括有天空、两个建筑物,建筑物之间有条街道,街道上有两辆汽车为例。初始结构化描述信息具体可以包括:
“天空(无生命物体;背景;天空呈现出蓝色的渐变,并散布着云朵。颜色信息:蓝色色调。);
建筑1(无生命物体;背景;建筑具有独特的角度设计,内部照明的窗户。颜色信息:黑色和黄色灯光。);
建筑2(无生命物体;背景;这座建筑很高,有许多窗户,其中一些亮着灯,具有圆柱形形状。颜色信息:白色和黄色亮灯窗户。);
街道(无生命物体;前景/背景;街道显示有多个车道,车辆和湿漉漉的地面反射着灯光。颜色信息:深灰色沥青,白色道路标线。);
汽车1(无生命物体;前景;一辆汽车开着前灯在街道上行驶,车身类型为轿车。颜色信息:黑色。);
汽车2(无生命物体;前景;另一辆汽车跟随汽车1,也开着前灯,看起来是掀背车。颜色信息:银色。)”。
其中,天空、建筑1、建筑2、街道、汽车1、汽车2为6个对象,每个对象后面括号中的内容即为对象描述信息。
需要注意的是,在本公开实施例提供的初始结构化描述信息中的对象和对象描述信息可以通过对象列表的形式进行展示。用户通过列表形式的对象和对象描述信息,可以更直观、准确的了解待处理图像中所包含的对象和对象描述信息。
在本公开提供的一具体实施方式中,获得所述图像信息提取模型输出的初始结构化描述信息,包括:
获得所述图像信息提取模型输出的初始图结构描述信息,其中,所述初始图结构描述信息中包括至少一个对象节点和各对象节点对应的对象描述信息。
在本实施方式中,图像信息提取模型会根据输入的待处理图像输出初始图结构描述信息,即初始结构化描述信息具体可以为初始图结构描述信息。图结构是一种非线性数据结构,由节点和边组成。在本公开实施例提供的方法中,初始图结构描述信息中包括至少一个对象节点以及各对象节点对应的对象描述信息。
更进一步的,所述初始结构化描述信息中还包括所述待处理图像的图像整体描述信息、各对象之间的关联关系。
在实际应用中,初始结构化描述信息中还包括有针对待处理图像的图像整体描述信息,用于从整体上体现待处理图像的内容。初始结构化描述信息中还包括有各对象之间的关联关系。
例如,依然以上述某个待处理图片中包括有天空、两个建筑物,建筑物之间有条街道,街道上有两辆汽车为例。其图像整体描述信息可以包括:
“风格:这张图片是一张具有现实主义风格的照片。
主题:图片的主题是黄昏时分的城市交通。
背景描述:背景由一片多云的天空组成,天空呈现出从顶部深蓝色到底部渐蓝色的渐变,左侧是一个名为“***”的大型现代建筑,以其独特的角度设计和内部照明的窗户为特色。右侧是一座高楼,拥有许多窗户,其中一些发出温暖的光芒。更远处,有更多城市建筑,包括一个带有绿色玻璃幕墙的建筑。环境光线表明现在是黎明或黄昏,人造光开始对这个场景产生显著影响。
前景描绘:前景展示了一个城市街道场景,一条街道上有多条车道,街道繁忙,车辆都开着前灯,表明光线条件较差,车辆的大小和形状各异,有私人车辆和商业车辆的混合,沿着街道的人行道在湿漉漉的灯光下显得特别干净,这是晚上城市环境的一个典型特征。”。
另外,除了上述的图像整体描述信息和对象列表之外,初始结构化描述信息中还包括有各对象之间的关联关系,具体为:
“天空覆盖着建筑;
建筑1位于街道的左侧;
建筑2位于街道的右侧;
汽车1行驶在街道上;
汽车2跟随汽车1行驶在街道上。”。
以上,为初始结构化描述信息的具体内容的介绍,在本公开实施例提供的一具体实施方式中,初始结构化描述信息至少包括对象列表,对象列表可以理解为至少一个对象和各对对象的对象描述信息。更进一步的,初始结构化描述信息还可以包括待处理图像的图像整体描述信息,和/或各对象之间的关联关系。
参见图2,图2示出了本公开一实施例提供的初始结构化描述信息的组成示意图,如图2所示,待处理图像在经过图像信息提取模型处理后,生成初始结构化描述信息,初始结构化描述信息中包括三个部分,第一个部分为图像整体描述信息,第二个部分为对象列表,第三个部分为各对象之间的关联关系。
第一个部分为图像整体描述信息,其包括了待处理图像的整体描述信息,例如图像风格、图像主题、图像背景描述、图像前景描述等等信息。
第二个部分为对象列表,其包括了待处理图像中的至少一个对象和各对象的对象描述信息。
第三个部分为关联关系,其包括了待处理图像中各对象之间的关联关系。
初始结构化描述信息为一种特殊格式的图结构描述信息,为了便于模型输出,可以使用初始结构化描述信息的形式对模型理解分析的内容进行展示,初始结构化描述信息中的对象和对象描述信息可以理解为图结构中的节点,以及节点的描述信息。
在本公开实施例提供的方法中,图像信息提取模型是预先被训练好的大语言模型,其被训练为根据输入的图像,输出固定格式的初始结构化描述信息。在本公开提供的一具体实施方式中,所述图像信息提取模型通过下述步骤训练获得:
获取样本图像和所述样本图像对应的样本结构化描述信息,其中,所述样本结构化描述信息通过预设结构化描述模版生成。
将所述样本图像输入至图像信息提取模型,获得所述图像信息提取模型输出的预测结构化描述信息。
根据所述样本结构化描述信息和所述预测结构化描述信息计算模型损失值。
根据所述模型损失值调整所述图像信息提取模型的模型参数,并继续训练所述图像信息提取模型,直至达到模型训练停止条件。
具体的,在本公开实施例提供的图像信息提取模型的训练方法使用的是有监督训练,其包括有多个训练样本对,某个训练样本对包括样本图像和样本图像对应的样本结构化描述信息。其中,样本图像可以理解为用于训练图像信息提取模型的样本。样本结构化描述信息可以理解为与样本图像对应的结构化描述信息,其包括有样本图像中的至少一个样本对象和各样本对象对应的样本对象描述信息。
在实际应用中,尽管多模态大语言模型的能力很强,但是多模态大语言模型也无法直接输出我们希望的结构化描述信息,因此,在本公开实施例提供的方法中,基于预设结构化描述模版为图像信息提取模型提供了模版,告知图像信息提取模型需要根据预设结构化描述模版的格式输出相应的结构化描述信息。
具体的,在本公开提供的一具体实施方式中,获取样本图像和所述样本图像对应的样本结构化描述信息,包括:
获取样本图像和预设结构化描述模版;
将所述样本图像和所述预设结构化描述模版输入至多模态大语言模型,获得所述多模态大语言模型输出的样本结构化描述信息。
其中,预设结构化描述模版可以理解为本公开一个或的多个实施方式中预先设置好的模版,其用于告知多模态大语言模型需要根据预设结构化描述模版的格式输出相应的结构化描述信息。
在实际应用中,所述预设结构化描述模版包括至少一类预设标识字符,预设标识字符表示预设结构化描述模版中的信息类型。例如,预设标识字符可以为“%%”、“&&”、“<>”、“()”、“;”、“[]”等等。每个预设标识字符标识预设结构化模版中的一种信息类型,例如,“%%”用来区分大标题,“&&”用来区分小标题,“<>”用来标识名词、“()”用来标识物体属性、“;”用来分割物体属性、“[]”用来标识关系等等。
通过预设结构化描述模版中的各类预设标识符,可以将预设结构化描述模版划分为预先规定的信息部分,从而规定好预设结构化描述模版的预设字符串格式。
在设计好预设字符串格式之后,为了使得多模态大语言模型可以根据设置好的预设结构化描述模版按照预设字符串格式输出结构化描述信息,在本公开提供的一个或多个实施方式中,还可以在预设结构化描述模版中包括目标信息类型对应的目标信息类型实例。通过In-context learning方法(ICL,仅提供少量示例就能完成自然语言处理任务的方法)为多模态大语言模型提供示例,使得多模态大语言模型根据预设字符格式输出相应的结构化描述信息。
传统的ICL方法需要提供多个示例,以对话的形式输入至多模态大语言模型,但是本公开实施例提供的方法中,预设字符串格式的内容较多,如果对所有的信息均提供示例的成本较高。基于此,在本公开实施例提供的方法中,对目标信息类型提供相应的目标信息类型实例。目标信息类型可以理解为预先设置的信息类型,例如,目标信息类型可以是大标题、小标题、对同类别的不同个体进行编号、物体信息的输出格式、对象之间关联关系等。针对目标信息类型提供相应的目标信息类型实例可以更好的帮助多模态大语言模型按照给定的预设字符串格式进行输出。
在本公开实施例提供的方法中,使用的ICL方法,仅在目标信息类型的位置提供目标信息类型实例,即可使得多模态大语言模型按照规定的格式输出相应的结构化描述信息。缩短了输入到多模态大语言模型的文本长度,仅使用单轮对话的方式,即可使得多模态大语言模型输出结构化描述信息。
参见图3,图3示出了本公开一实施例提供的生成样本结构化描述信息的示意图,如图3所示,将样本图像和预设结构化描述模版输入至多模态大语言模型,多模态大语言模型可以在对样本图像进行理解和分析之后,根据预设结构化描述模版输出相应的样本结构化描述信息,样本结构化描述信息根据预先设置好的字符串格式和示例生成。
如图3所示的样本结构化描述信息,其包括有三部分,第一部分为整体描述部分,其中进一步包括了风格、主题、背景描述、前景描述等信息。第二部分为对象列表,其中进一步包括了图像中的各对象,以及各对象的对象描述信息。第三部分为关联关系,进一步包括图像中各对象之间的关联关系。
样本结构化描述信息为本公开实施例提供的一种预先设置好的图像描述信息的结构化格式,多模态大语言模型在对样本图像进行理解分析后,根据预先设置好的预设字符串格式,输出相应的样本结构化描述信息,通过特定的结构化格式对样本图像进行描述,以便在后续对图像信息提取模型进行训练。使得训练好的图像信息提取模型可以输出指定格式的初始结构化描述信息。
在通过上述步骤获得样本图像对应的样本结构化描述信息之后,即可根据样本图像和样本结构化描述信息对图像信息提取模型进行训练。具体的,将样本图像输入至图像信息提取模型,此时的图像信息提取模型可以理解为是上述的多模态大语言模型,其具备了根据样本图像输出预设格式的能力,样本图像输入到图像信息提取模型中进行处理,图像信息提取模型可以根据样本图像输出预测结构化描述信息。
预测结构化描述信息可以理解为未训练好的图像信息提取模型根据输入的样本图像输出的描述信息。其数据结构与预先规定好的预设字符格式相对应。
此时的图像信息提取模型是还未被训练好的模型,其输出的预测结构化描述信息与样本结构化描述信息之间还存在差异,需要根据预测结构化描述信息和样本结构化描述信息计算模型损失值。在本公开实施例提供的方法中,计算模型损式值的方法有很多,例如交叉熵损失函数、最大损失函数、平均值损失函数等等。在本公开提供的实施方式中,对损失函数的具体方式不做限定,以实际应用为准。
在获得模型损失值后,即可根据模型损失值在图像信息提取模型中进行反向传播,调整图像信息提取模型中的模型参数。随后可以重复上述步骤,继续对图像信息提取模型进行训练,直至达到模型训练停止条件,在实际应用中,初始图像处理模型的模型训练停止条件包括:
模型损失值小于预设阈值,和/或训练轮次达到预设的训练轮次。
具体的,在对图像信息提取模型进行训练的过程中,可以将模型的训练停止条件设置为模型损失值小于预设阈值,即当模型损失值小于预设阈值的情况下,无需再调整图像信息提取模型的模型参数。
也可以将图像信息提取模型的模型训练停止条件设置为训练轮次达到预设的训练轮次,例如预设的训练轮次为10轮,当模型的训练轮次达到了10轮的情况下,即达到了模型训练停止条件。
在本公开提供的方法中,对模型训练停止条件不做限定。当图像信息提取模型达到模型训练停止条件的情况下,则说明图像信息提取模型训练完成,获得最终的图像信息提取模型。
步骤106:根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息。
其中,对象位置信息可以理解为对象在待处理图像中的位置信息,图像信息提取模型虽然也具备一定的对象定位能力,但是其定位效果较差,为了便于下游的任务从待处理图像中提取出相应的图像描述信息,还需要进一步获取到各对象更精准的对象位置信息。
在本实施方式中,在获得了初始结构化描述信息后,根据初始结构化描述信息中的至少一个对象和各对象的对象描述信息可以进一步的在待处理图像中确定各对象对应的对象位置信息。具体的,可以通过标记框的形式标记处各对象在待处理图像中的对象位置信息。
在本公开提供的一具体实施方式中,根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息,包括:
S1062、获取所述初始结构化描述信息中的待处理对象和所述待处理对象对应的待处理对象描述信息,其中,所述待处理对象为所述至少一个对象中的任意一个。
在实际应用中,初始结构化描述信息中包括有至少一个对象和各对象对应的对象描述信息,在本实施方式中,以其中一个对象为例进行解释说明,则待处理对象即为被操作处理的对象,待处理对象描述信息即为待处理对象对应的对象描述信息。具体的,可以从对象列表中选取待处理对象和待处理对象对应的待处理对象描述信息。
待处理对象可以是各对象中的任一个,对象之间的处理可以是并行,也可以是顺次执行,在本公开提供的实施例中,对各对象之间的处理顺序不做限定。
S1064、根据所述待处理对象在所述待处理图像中确定至少一个候选对象标记框。
在确定了待处理对象后,即可根据待处理对象在待处理图像中确定至少一个候选对象标记框,候选对象标记框(bounding box)是检测定位任务中在图像中生成的针对对象的标记框,其可以标记处对象的在图像中的位置信息。
在本公开实施例提供的一具体实施方式中,可以通过定位模型来生成候选对象标记框,根据所述待处理对象在所述待处理图像中确定至少一个候选对象标记框,包括:
将所述待处理对象和所述待处理图像输入至定位模型,获得所述定位模型根据所述待处理对象在所述待处理图像中确定的至少一个候选对象标记框。
定位模型可以理解为根据输入的信息,在图像中进行定位标记的模型,定位模型可以根据输入的信息,在待处理图像中标记处于该信息对应的对象。例如,以待处理对象为“man”为例,将待处理对象和待处理图像输入到定位模型中,定位模型在待处理图像中标记出与“man”相关的候选对象标记框。
此时,定位模型仅能根据输入的待处理对象在待处理图像中标记处至少一个候选标记框,但是定位模型难以区分相同类别的不同个体,例如,当待处理对象为“man”时,定位模型仅能标记处待处理图像中的多个man的对象,但是并无法区分各个man之间的区别。因此,本公开实施例提供的方法,需要再进一步对定位模型确定的候选对象标记框进行筛选,获得最终的目标对象标记框。
在本公开提供的一具体实施方式中,根据所述待处理对象在所述待处理图像中确定至少一个候选对象标记框,包括:
将所述待处理对象和所述待处理图像输入至定位模型,获得所述定位模型根据所述待处理对象在所述待处理图像中确定的至少一个初始候选对象标记框;
将所述待处理对象和各初始候选对象标记框输入至多模态大语言模型,获得所述多模态大语言模型输出的至少一个候选对象标记框,其中,所述多模态大语言模型根据所述待处理对象筛选各初始候选对象标记框。
在实际应用中,候选对象标记框中还有可能会有与待处理对象差距较大的情况,为了便于在后续的处理过程中提升处理效率,还可以利用多模态大语言模型对候选对象标记框进行筛选。
具体的,先将待处理对象和待处理图像输入至定位模型,定位模型输出至少一个初始候选对象标记框,此时的候选对象标记框由于定位模型的精度问题,可能会存在定位错误的标记框,即可再将待处理对象和各初始候选对象标记框输入至多模态大语言模型中进行筛选,将明显与待处理对象不符的初始候选标记框进行删除,保留与待处理对象匹配的候选标记框。
在本公开提供的一个或多个具体实施方式中,可以是将待处理对象和各初始候选标记框一起输入到多模态大语言模型中进行筛选,多模态大语言模型输出与待处理对象匹配的候选标记框;也可以是从多个初始候选对象标记框中选取一个初始候选标记框,并将选取的初始候选标记框和待处理对象输入到多模态大语言模型中进行筛选,多模态大语言模型确定该初始候选标记框与待处理对象之间的匹配度,将匹配度满足预设阈值的初始候选标记框作为候选标记框。
S1066、根据所述待处理对象描述信息在至少一个候选对象标记框中确定目标对象标记框,并确定所述目标对象标记框的标记框位置信息。
在本公开实施例提供的方法中,待处理对象的待处理对象描述信息用于在至少一个候选对象标记框中选取出目标对象标记框。待处理对象描述信息为待处理对象的详细描述内容,通过该待处理对象描述信息,可以协助从候选对象标记框中选取出目标对象标记框,并确定目标对象标记框的标记框位置信息。
具体的,根据所述待处理对象描述信息在至少一个候选对象标记框中确定目标对象标记框,包括:
计算所述待处理对象描述信息和各候选对象标记框的匹配相似度;
确定匹配相似度满足预设条件的候选对象标记框为目标对象标记框。
在本公开提供的一个或多个实施例中,计算待处理对象描述信息与各候选对象标记框的匹配相似度。其中,匹配相似度表征了待处理对象描述信息与候选对象标记框之间的相似度,匹配相似度越高,说明待处理对象与候选对象标记框越匹配,反之,说明待处理对象与候选对象标记框越不匹配。
在计算出各候选对象标记框的匹配相似度后,将满足预设条件的匹配相似度对应的候选对象标记框作为目标对象标记框。预设条件可以理解为从多个候选对象标记框中选取目标对象标记框的条件,例如预设条件可以是匹配相似度最高的候选对象标记框为目标对象标记框。在确定目标对象框后,即可根据目标对象框确定对应的标记框位置信息。
S1068、将所述标记框位置信息确定为所述待处理对象对应的对象位置信息。
在确定了标记框位置信息后,将标记框位置信息确定为待处理对象对应的对象位置信息。
参见图4,图4示出了本申请一实施例提供的对象定位的示意图,如图4所示,将待处理图像输入至图像信息提取模型中进行提取,获得图像信息提取模型输出的初始结构化描述信息,在初始结构化描述信息中包括有对象列表,在对象列表中包括待处理图像中包括的至少一个对象和各对象的对象描述信息。
从对象列表中选取出待处理对象“人”,其对应的待处理对象描述信息为“带着帽子”,由于定位模型只能处理对象名称,则将待处理对象“人”输入到定位模型中,定位模型根据待处理对象“人”在待处理图像中进行定位,生成四个初始候选标记框。
将四个初始候选标记框和待处理对象“人”输入至多模态大语言模型中进行筛选,将明显与待处理对象“人”不匹配的初始候选标记框删除,确定三个候选标记框。
将三个候选标记框分别于待处理对象描述信息“带着帽子”进行匹配,选取匹配信息最高的候选标记框为目标标记框,并在待处理图像中标记该目标标记框。
步骤108:根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息。
在确定了各对象对应的对象位置信息后,即可将各对象位置信息添加到初始结构化描述信息中,与各对象进行一一对应。生成的最终的结构化图像描述信息中,不仅包括各对象和各对象的对象描述信息,还包括各对象对应的对象位置信息。
添加了各对象位置信息的结构化图像描述信息,可以从多个维度对待处理图像进行描述,便于下游任务根据不同的任务需求,从结构化图像描述信息中选取相应的子信息,并执行后续的下游任务。具体的,在本公开提供的又一具体实施方式中,所述方法还包括:
接收下游任务的数据请求指令,其中,所述数据请求指令中携带有数据请求信息;
响应于所述数据请求信息从所述结构化图像描述信息中提取图像描述子信息,并将所述图像描述子信息发送至所述下游任务。
其中,下游任务可以理解为需要依赖于结构化图像描述信息的其他图像处理任务,不同的下游任务可以针对待处理图像执行不同的操作,因此不同的下游任务从待处理图像中获取的信息也不同。在接收到下游任务的数据请求指令后,可以根据数据请求指令中的数据请求信息,在结构化图像描述信息中进行筛选,提取出与数据请求信息相对应的图像描述子信息,并将图像描述子信息发送至下游任务,使得下游任务的多模态模型可以根据图像描述子信息执行相应的下游任务。
本公开实施例提供的图像处理方法,包括获取待处理图像;将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息;根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息;根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息。
通过图像信息提取模型按照预设的格式生成待处理图像对应的初始结构化描述信息,根据初始结构化描述信息中的至少一个对象和各对象对应的对象描述信息,可以在待处理图像中确定出各对象的对象位置信息,并根据初始结构化描述信息和各对象的对象位置信息确定出最终的结构化图像描述信息。通过特定的结构化描述图像的方式,将图像的内容进行描述,便于在后续从结构化图像描述信息中的提取相关信息。
其次,基于特定的结构化描述图像的方式训练图像信息提取模型,使用ICL方法,仅针对目标信息类型提供相应的目标信息类型实例,使得图像信息提取模型可以很好的学习到预设的字符串格式,从而根据预设的字符串格式输出初始结构化描述信息。降低了训练成本,简化了学习流程。
最后,通过初始结构化描述信息和定位模型在待处理图像中对各对象进行定位,确定出各对象对应的位置信息。从而更好的对初始结构化描述信息中的各对象进行更详细的描绘。更进一步的提升后续下游任务从结构化图像描述信息中获取信息的便捷性。
图5示出了本公开一实施例提供的一种图像处理方法的处理框架图,如图5所示,本公开实施例提供的方法中,通过预先设置的字符串格式设计预设结构化描述模版。利用预设结构化描述模版通过多模态大语言模型生成样本结构化描述信息,将样本结构化描述信息和样本图像训练图像信息提取模型,使得图像信息提取模型具备根据输入的图像输出结构化描述信息的能力。
在模型训练完成后,将待处理图像输入至训练好的图像信息提取模型,获得图像信息提取模型输出的初始结构化描述信息,这里的初始结构化描述信息包括至少一个对象和各对象对应的对象描述信息,但是不包括各对象在待处理图像中的位置信息。
再基于对象和对象描述信息通过定位模型的定位处理,获得各对象在待处理图像中的位置信息,将各对象的位置信息添加到初始结构化描述信息中,生成最终的结构化图像描述信息。在结构化图像描述信息中包括有至少一个对象,各对象的对象描述信息和对象位置信息。
图6示出了本公开一实施例提供的一种应用于云侧设备的图像处理方法的流程示意图,该方法应用于云侧设备,具体包括:
步骤602:接收端侧设备发送的待处理图像。
步骤604:将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息。
步骤606:根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息。
步骤608:根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息。
步骤610:将所述结构化图像描述信息发送至所述端侧设备。
在本公开提供的一具体实施方式中,所述方法还包括:
接收端侧设备针对所述结构化图像描述信息发送的信息处理指令;
响应于所述信息处理指令处理所述结构化图像描述信息,获得图像描述调整信息;
将所述图像描述调整信息发送至所述端侧设备。
在实际应用中,由于图像信息提取模型是多模态大语言模型,其在运行时需要较好的计算资源,端侧设备可能不具备相应的处理能力。因此,可以将图像信息提取模型部署在云侧设备,即本公开实施例提供的图像处理过程在云侧设备实现。云侧设备在获得了图像信息提取模型生成的结构化图像描述信息后,还可以将结构化图像描述信息发送至端侧设备。
本公开实施例提供的图像处理方法,通过图像信息提取模型按照预设的格式生成待处理图像对应的初始结构化描述信息,根据初始结构化描述信息中的至少一个对象和各对象对应的对象描述信息,可以在待处理图像中确定出各对象的对象位置信息,并根据初始结构化描述信息和各对象的对象位置信息确定出最终的结构化图像描述信息。通过特定的结构化描述图像的方式,将图像的内容进行描述,便于在后续从结构化图像描述信息中的提取相关信息。
其次,基于特定的结构化描述图像的方式训练图像信息提取模型,使用ICL方法,仅针对目标信息类型提供相应的目标信息类型实例,使得图像信息提取模型可以很好的学习到预设的字符串格式,从而根据预设的字符串格式输出初始结构化描述信息。降低了训练成本,简化了学习流程。
最后,通过初始结构化描述信息和定位模型在待处理图像中对各对象进行定位,确定出各对象对应的位置信息。从而更好的对初始结构化描述信息中的各对象进行更详细的描绘。更进一步的提升后续下游任务从结构化图像描述信息中获取信息的便捷性。
与上述方法实施例相对应,本公开还提供了图像处理装置实施例,图7示出了本公开一个实施例提供的一种图像处理装置的结构示意图。如图7所示,该装置包括:
获取模块702,被配置为获取待处理图像;
信息提取模块704,被配置为将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息;
确定模块706,被配置为根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息;
生成模块708,被配置为根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息。
可选的,信息提取模块704,被配置为:
获得所述图像信息提取模型输出的初始图结构描述信息,其中,所述初始图结构描述信息中包括至少一个对象节点和各对象节点对应的对象描述信息。
可选的,所述装置还包括训练模块,被配置为:
获取样本图像和所述样本图像对应的样本结构化描述信息,其中,所述样本结构化描述信息通过预设结构化描述模版生成;
将所述样本图像输入至图像信息提取模型,获得所述图像信息提取模型输出的预测结构化描述信息;
根据所述样本结构化描述信息和所述预测结构化描述信息计算模型损失值;
根据所述模型损失值调整所述图像信息提取模型的模型参数,并继续训练所述图像信息提取模型,直至达到模型训练停止条件。
可选的,所述训练模块,进一步被配置为:
获取样本图像和预设结构化描述模版;
将所述样本图像和所述预设结构化描述模版输入至多模态大语言模型,获得所述多模态大语言模型输出的样本结构化描述信息。
可选的,所述预设结构化描述模版包括至少一类预设标识字符,预设标识字符表示预设结构化描述模版中的信息类型。
可选的,所述预设结构化描述模版中包括目标信息类型对应的目标信息类型实例。
可选的,所述确定模块706,进一步被配置为:
获取所述初始结构化描述信息中的待处理对象和所述待处理对象对应的待处理对象描述信息,其中,所述待处理对象为所述至少一个对象中的任意一个;
根据所述待处理对象在所述待处理图像中确定至少一个候选对象标记框;
根据所述待处理对象描述信息在至少一个候选对象标记框中确定目标对象标记框,并确定所述目标对象标记框的标记框位置信息;
将所述标记框位置信息确定为所述待处理对象对应的对象位置信息。
可选的,所述确定模块706,进一步被配置为:
将所述待处理对象和所述待处理图像输入至定位模型,获得所述定位模型根据所述待处理对象在所述待处理图像中确定的至少一个候选对象标记框。
可选的,所述确定模块706,还被配置为:
将所述待处理对象和所述待处理图像输入至定位模型,获得所述定位模型根据所述待处理对象在所述待处理图像中确定的至少一个初始候选对象标记框;
将所述待处理对象和各初始候选对象标记框输入至多模态大语言模型,获得所述多模态大语言模型输出的至少一个候选对象标记框,其中,所述多模态大语言模型根据所述待处理对象筛选各初始候选对象标记框。
可选的,所述确定模块706,进一步被配置为:
计算所述待处理对象描述信息和各候选对象标记框的匹配相似度;
确定匹配相似度满足预设条件的候选对象标记框为目标对象标记框。
可选的,所述初始结构化描述信息中还包括所述待处理图像的图像整体描述信息、各对象之间的关联关系。
可选的,所述装置还包括信息提取模块,被配置为:
接收下游任务的数据请求指令,其中,所述数据请求指令中携带有数据请求信息;
响应于所述数据请求信息从所述结构化图像描述信息中提取图像描述子信息,并将所述图像描述子信息发送至所述下游任务。
本公开实施例提供的图像处理装置,包括获取待处理图像;将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息;根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息;根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息。
通过图像信息提取模型按照预设的格式生成待处理图像对应的初始结构化描述信息,根据初始结构化描述信息中的至少一个对象和各对象对应的对象描述信息,可以在待处理图像中确定出各对象的对象位置信息,并根据初始结构化描述信息和各对象的对象位置信息确定出最终的结构化图像描述信息。通过特定的结构化描述图像的方式,将图像的内容进行描述,便于在后续从结构化图像描述信息中的提取相关信息。
其次,基于特定的结构化描述图像的方式训练图像信息提取模型,使用ICL方法,仅针对目标信息类型提供相应的目标信息类型实例,使得图像信息提取模型可以很好的学习到预设的字符串格式,从而根据预设的字符串格式输出初始结构化描述信息。降低了训练成本,简化了学习流程。
最后,通过初始结构化描述信息和定位模型在待处理图像中对各对象进行定位,确定出各对象对应的位置信息。从而更好的对初始结构化描述信息中的各对象进行更详细的描绘。更进一步的提升后续下游任务从结构化图像描述信息中获取信息的便捷性。
上述为本实施例的一种图像处理装置的示意性方案。需要说明的是,该图像处理装置的技术方案与上述的图像处理方法的技术方案属于同一构思,图像处理装置的技术方案未详细描述的细节内容,均可以参见上述图像处理方法的技术方案的描述。
与上述方法实施例相对应,本公开还提供了图像处理装置实施例,图8示出了本公开一个实施例提供的一种应用于云侧设备的图像处理装置的结构示意图。如图8所示,该装置应用于云侧设备,包括:
接收模块802,被配置为接收端侧设备发送的待处理图像;
信息提取模块804,被配置为将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息;
确定模块806,被配置为根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息;
生成模块808,被配置为根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息;
发送模块810,被配置为将所述结构化图像描述信息发送至所述端侧设备。
可选的,所述装置还包括调整模块,被配置为:
接收端侧设备针对所述结构化图像描述信息发送的信息处理指令;
响应于所述信息处理指令处理所述结构化图像描述信息,获得图像描述调整信息;
将所述图像描述调整信息发送至所述端侧设备。
本公开实施例提供的图像处理装置,通过图像信息提取模型按照预设的格式生成待处理图像对应的初始结构化描述信息,根据初始结构化描述信息中的至少一个对象和各对象对应的对象描述信息,可以在待处理图像中确定出各对象的对象位置信息,并根据初始结构化描述信息和各对象的对象位置信息确定出最终的结构化图像描述信息。通过特定的结构化描述图像的方式,将图像的内容进行描述,便于在后续从结构化图像描述信息中的提取相关信息。
其次,基于特定的结构化描述图像的方式训练图像信息提取模型,使用ICL方法,仅针对目标信息类型提供相应的目标信息类型实例,使得图像信息提取模型可以很好的学习到预设的字符串格式,从而根据预设的字符串格式输出初始结构化描述信息。降低了训练成本,简化了学习流程。
最后,通过初始结构化描述信息和定位模型在待处理图像中对各对象进行定位,确定出各对象对应的位置信息。从而更好的对初始结构化描述信息中的各对象进行更详细的描绘。更进一步的提升后续下游任务从结构化图像描述信息中获取信息的便捷性。
上述为本实施例的一种图像处理装置的示意性方案。需要说明的是,该图像处理装置的技术方案与上述的图像处理方法的技术方案属于同一构思,图像处理装置的技术方案未详细描述的细节内容,均可以参见上述图像处理方法的技术方案的描述。
参见图9,图9示出了本公开一个实施例提供的一种图像处理系统的架构图,图像处理系统可以包括客户端100和服务端200;
客户端100,用于向服务端200发送待处理图像;
服务端200,用于将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息;根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息;根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息;向客户端100发送结构化图像描述信息;
客户端100,还用于接收服务端200发送的结构化图像描述信息。
图像处理系统可以包括多个客户端100以及服务端200,其中,客户端100可以称为端侧设备,服务端200可以称为云侧设备。多个客户端100之间通过服务端200可以建立通信连接,在图像处理场景中,服务端200即用来在多个客户端100之间提供图像处理服务,多个客户端100可以分别作为发送端或接收端,通过服务端200实现通信。
用户通过客户端100可与服务端200进行交互以接收其它客户端100发送的数据,或将数据发送至其它客户端100等。在图像处理场景中,可以是用户通过客户端100向服务端200发布数据流,服务端200根据该数据流生成结构化图像描述信息,并将结构化图像描述信息推送至其他建立通信的客户端中。
其中,客户端100与服务端200之间通过网络建立连接。网络为客户端100与服务端200之间提供了通信链路的介质。网络可以包括各种连接类型,例如有线、无线通信链路或者光纤电缆等等。客户端100所传输的数据可能需要经过编码、转码、压缩等处理之后才发布至服务端200。
客户端100可以为浏览器、APP(Application,应用程序)、或网页应用如H5(HyperText Markup Language5,超文本标记语言第5版)应用、或轻应用(也被称为小程序,一种轻量级应用程序)或云应用等,客户端100可以基于服务端200提供的相应服务的软件开发工具包(SDK,Software Development Kit),如基于实时通信(RTC,Real Time Communication)SDK开发获得等。客户端100可以部署在电子设备中,需要依赖设备运行或者设备中的某些APP而运行等。电子设备例如可以具有显示屏并支持信息浏览等,如可以是个人移动终端如手机、平板电脑、个人计算机等。在电子设备中通常还可以配置各种其它类应用,例如人机对话类应用、模型训练类应用、文本处理类应用、网页浏览器应用、购物类应用、搜索类应用、即时通信工具、邮箱客户端、社交平台软件等。
服务端200可以包括提供各种服务的服务器,例如为多个客户端提供通信服务的服务器,又如为客户端上使用的模型提供支持的用于后台训练的服务器,又如对客户端发送的数据进行处理的服务器等。需要说明的是,服务端200可以实现成多个服务器组成的分布式服务器集群,也可以实现成单个服务器。服务器也可以为分布式系统的服务器,或者是结合了区块链的服务器。服务器也可以是云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、内容分发网络(CDN,Content Delivery Network)以及大数据和人工智能平台等基础云计算服务的云服务器,或者是带人工智能技术的智能云计算服务器或智能云主机。
值得说明的是,本公开实施例中提供的图像处理方法一般由服务端执行,但是,在本公开的其它实施例中,客户端也可以与服务端具有相似的功能,从而执行本公开实施例所提供的图像处理方法。在其它实施例中,本公开实施例所提供的图像处理方法还可以是由客户端与服务端共同执行。
图10示出了根据本公开一个实施例提供的一种计算设备1000的结构框图。该计算设备1000的部件包括但不限于存储器1010和处理器1020。处理器1020与存储器1010通过总线1030相连接,数据库1050用于保存数据。
计算设备1000还包括接入设备1040,接入设备1040使得计算设备1000能够经由一个或多个网络1060通信。这些网络的示例包括公用交换电话网(PSTN,Public Switched Telephone Network)、局域网(LAN,Local Area Network)、广域网(WAN,Wide Area Network)、个域网(PAN,Personal Area Network)或诸如因特网的通信网络的组合。接入设备1040可以包括有线或无线的任何类型的网络接口(例如,网络接口卡(NIC,network interface controller))中的一个或多个,诸如IEEE802.11无线局域网(WLAN,Wireless Local Area Network)无线接口、全球微波互联接入(Wi-MAX,Worldwide Interoperability for Microwave Access)接口、以太网接口、通用串行总线(USB,Universal Serial Bus)接口、蜂窝网络接口、蓝牙接口、近场通信(NFC,Near Field Communication)。
在本公开的一个实施例中,计算设备1000的上述部件以及图10中未示出的其他部件也可以彼此相连接,例如通过总线。应当理解,图10所示的计算设备结构框图仅仅是出于示例的目的,而不是对本公开范围的限制。本领域技术人员可以根据需要,增添或替换其他部件。
计算设备1000可以是任何类型的静止或移动计算设备,包括移动计算机或移动计算设备(例如,平板计算机、个人数字助理、膝上型计算机、笔记本计算机、上网本等)、移动电话(例如,智能手机)、可佩戴的计算设备(例如,智能手表、智能眼镜等)或其他类型的移动设备,或者诸如台式计算机或个人计算机(PC,Personal Computer)的静止计算设备。计算设备1000还可以是移动式或静止式的服务器。
其中,处理器1020用于执行如下计算机程序/指令,该计算机程序/指令被处理器执行时实现上述图像处理方法的步骤。
本公开中的各个实施例均采用递进的方式描述,各个实施例之间相同相似的部分互相参见即可,每个实施例重点说明的都是与其他实施例的不同之处。尤其,对于计算设备实施例而言,由于其基本相似于图像处理方法实施例,所以描述的比较简单,相关之处参见图像处理方法实施例的部分说明即可。
本公开一实施例还提供一种计算机可读存储介质,其存储有计算机程序/指令,该计算机程序/指令被处理器执行时实现上述图像处理方法的步骤。
本公开中的各个实施例均采用递进的方式描述,各个实施例之间相同相似的部分互相参见即可,每个实施例重点说明的都是与其他实施例的不同之处。尤其,对于计算机可读存储介质实施例而言,由于其基本相似于图像处理方法实施例,所以描述的比较简单,相关之处参见图像处理方法实施例的部分说明即可。
本公开一实施例还提供一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现上述图像处理方法的步骤。
上述为本实施例的一种计算机程序产品的示意性方案。需要说明的是,该计算机程序产品的技术方案与上述的图像处理方法的技术方案属于同一构思,计算机程序产品的技术方案未详细描述的细节内容,均可以参见上述图像处理方法的技术方案的描述。
上述对本公开特定实施例进行了描述。其它实施例在所附权利要求书的范围内。在一些情况下,在权利要求书中记载的动作或步骤可以按照不同于实施例中的顺序来执行并且仍然可以实现期望的结果。另外,在附图中描绘的过程不一定要求示出的特定顺序或者连续顺序才能实现期望的结果。在某些实施方式中,多任务处理和并行处理也是可以的或者可能是有利的。
所述计算机指令包括计算机程序代码,所述计算机程序代码可以为源代码形式、对象代码形式、可执行文件或某些中间形式等。所述计算机可读介质可以包括:能够携带所述计算机程序代码的任何实体或装置、记录介质、U盘、移动硬盘、磁碟、光盘、计算机存储器、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、电载波信号、电信信号以及软件分发介质等。需要说明的是,所述计算机可读介质包含的内容可以根据专利实践的要求进行适当的增减,例如在某些地区,根据专利实践,计算机可读介质不包括电载波信号和电信信号。
需要说明的是,上述对本公开特定实施例进行了描述。其它实施例在所附权利要求书的范围内。在一些情况下,在权利要求书中记载的动作或步骤可以按照不同于实施例中的顺序来执行并且仍然可以实现期望的结果。另外,在附图中描绘的过程不一定要求示出的特定顺序或者连续顺序才能实现期望的结果。在某些实施方式中,多任务处理和并行处理也是可以的或者可能是有利的。其次,本领域技术人员也应该知悉,本公开中所描述的实施例均属于优选实施例,所涉及的动作和模块并不一定都是本公开实施例所必须的。
在上述实施例中,对各个实施例的描述都各有侧重,某个实施例中没有详述的部分,可以参见其它实施例的相关描述。
以上公开的本公开优选实施例只是用于帮助阐述本公开。可选实施例并没有详尽叙述所有的细节,也不限制该发明仅为所述的具体实施方式。显然,根据本公开实施例的内容,可作很多的修改和变化。本公开选取并具体描述这些实施例,是为了更好地解释本公开实施例的原理和实际应用,从而使所属技术领域技术人员能很好地理解和利用本公开。本公开仅受权利要求书及其全部范围和等效物的限制。

Claims (18)

  1. 一种图像处理方法,包括:
    获取待处理图像;
    将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息;
    根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息;
    根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息。
  2. 如权利要求1所述的方法,获得所述图像信息提取模型输出的初始结构化描述信息,包括:
    获得所述图像信息提取模型输出的初始图结构描述信息,其中,所述初始图结构描述信息中包括至少一个对象节点和各对象节点对应的对象描述信息。
  3. 如权利要求1或2所述的方法,所述图像信息提取模型通过下述步骤训练获得:
    获取样本图像和所述样本图像对应的样本结构化描述信息,其中,所述样本结构化描述信息通过预设结构化描述模版生成;
    将所述样本图像输入至图像信息提取模型,获得所述图像信息提取模型输出的预测结构化描述信息;
    根据所述样本结构化描述信息和所述预测结构化描述信息计算模型损失值;
    根据所述模型损失值调整所述图像信息提取模型的模型参数,并继续训练所述图像信息提取模型,直至达到模型训练停止条件。
  4. 如权利要求3所述的方法,获取样本图像和所述样本图像对应的样本结构化描述信息,包括:
    获取样本图像和预设结构化描述模版;
    将所述样本图像和所述预设结构化描述模版输入至多模态大语言模型,获得所述多模态大语言模型输出的样本结构化描述信息。
  5. 如权利要求3或4所述的方法,所述预设结构化描述模版包括至少一类预设标识字符,预设标识字符表示预设结构化描述模版中的信息类型。
  6. 如权利要求3-5任意一项所述的方法,所述预设结构化描述模版中包括目标信息类型对应的目标信息类型实例。
  7. 如权利要求1-6任意一项所述的方法,根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息,包括:
    获取所述初始结构化描述信息中的待处理对象和所述待处理对象对应的待处理对象描述信息,其中,所述待处理对象为所述至少一个对象中的任意一个;
    根据所述待处理对象在所述待处理图像中确定至少一个候选对象标记框;
    根据所述待处理对象描述信息在至少一个候选对象标记框中确定目标对象标记框,并确定所述目标对象标记框的标记框位置信息;
    将所述标记框位置信息确定为所述待处理对象对应的对象位置信息。
  8. 如权利要求7所述的方法,根据所述待处理对象在所述待处理图像中确定至少一个候选对象标记框,包括:
    将所述待处理对象和所述待处理图像输入至定位模型,获得所述定位模型根据所述待处理对象在所述待处理图像中确定的至少一个候选对象标记框。
  9. 如权利要求7所述的方法,根据所述待处理对象在所述待处理图像中确定至少一个候选对象标记框,包括:
    将所述待处理对象和所述待处理图像输入至定位模型,获得所述定位模型根据所述待处理对象在所述待处理图像中确定的至少一个初始候选对象标记框;
    将所述待处理对象和各初始候选对象标记框输入至多模态大语言模型,获得所述多模态大语言模型输出的至少一个候选对象标记框,其中,所述多模态大语言模型根据所述待处理对象筛选各初始候选对象标记框。
  10. 如权利要求7-9任意一项所述的方法,根据所述待处理对象描述信息在至少一个候选对象标记框中确定目标对象标记框,包括:
    计算所述待处理对象描述信息和各候选对象标记框的匹配相似度;
    确定匹配相似度满足预设条件的候选对象标记框为目标对象标记框。
  11. 如权利要求1-10任意一项所述的方法,所述初始结构化描述信息中还包括所述待处理图像的图像整体描述信息、各对象之间的关联关系。
  12. 如权利要求1-11任意一项所述的方法,还包括:
    接收下游任务的数据请求指令,其中,所述数据请求指令中携带有数据请求信息;
    响应于所述数据请求信息从所述结构化图像描述信息中提取图像描述子信息,并将所述图像描述子信息发送至所述下游任务。
  13. 一种图像处理方法,应用于云侧设备,包括:
    接收端侧设备发送的待处理图像;
    将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息;
    根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息;
    根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息;
    将所述结构化图像描述信息发送至所述端侧设备。
  14. 如权利要求13所述的方法,还包括:
    接收端侧设备针对所述结构化图像描述信息发送的信息处理指令;
    响应于所述信息处理指令处理所述结构化图像描述信息,获得图像描述调整信息;
    将所述图像描述调整信息发送至所述端侧设备。
  15. 一种图像处理装置,包括:
    获取模块,被配置为获取待处理图像;
    信息提取模块,被配置为将所述待处理图像输入至图像信息提取模型,获得所述图像信息提取模型输出的初始结构化描述信息,其中,所述初始结构化描述信息包括至少一个对象和各对象的对象描述信息;
    确定模块,被配置为根据所述待处理图像和所述初始结构化描述信息确定各对象对应的对象位置信息;
    生成模块,被配置为根据所述初始结构化描述信息和各对象对应的对象位置信息生成所述待处理图像对应的结构化图像描述信息。
  16. 一种计算设备,包括:
    存储器和处理器;
    所述存储器用于存储计算机程序/指令,所述处理器用于执行所述计算机程序/指令,该计算机程序/指令被处理器执行时实现权利要求1至14任意一项所述方法的步骤。
  17. 一种计算机可读存储介质,其存储有计算机程序/指令,该计算机程序/指令被处理器执行时实现权利要求1至14任意一项所述方法的步骤。
  18. 一种计算机程序产品,包括计算机程序/指令,该计算机程序/指令被处理器执行时实现权利要求1至14任意一项所述方法的步骤。
PCT/CN2025/100832 2024-07-05 2025-06-13 图像处理方法及装置 Pending WO2026007670A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202410904423.X 2024-07-05
CN202410904423.XA CN118968081A (zh) 2024-07-05 2024-07-05 图像处理方法及装置

Publications (1)

Publication Number Publication Date
WO2026007670A1 true WO2026007670A1 (zh) 2026-01-08

Family

ID=93404603

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2025/100832 Pending WO2026007670A1 (zh) 2024-07-05 2025-06-13 图像处理方法及装置

Country Status (2)

Country Link
CN (1) CN118968081A (zh)
WO (1) WO2026007670A1 (zh)

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118968081A (zh) * 2024-07-05 2024-11-15 阿里巴巴(中国)有限公司 图像处理方法及装置
CN119444934A (zh) * 2024-11-19 2025-02-14 北京百度网讯科技有限公司 基于大模型的图像生成方法、装置、电子设备以及存储介质
CN119886362A (zh) * 2025-03-27 2025-04-25 阿里巴巴(中国)有限公司 模型训练方法、数据处理方法、系统及存储介质

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113449801A (zh) * 2021-07-08 2021-09-28 西安交通大学 一种基于多级图像上下文编解码的图像人物行为描述生成方法
CN114627353A (zh) * 2022-03-21 2022-06-14 北京有竹居网络技术有限公司 一种图像描述生成方法、装置、设备、介质及产品
CN115359323A (zh) * 2022-08-31 2022-11-18 北京百度网讯科技有限公司 图像的文本信息生成方法和深度学习模型的训练方法
CN116050496A (zh) * 2023-01-28 2023-05-02 Oppo广东移动通信有限公司 图片描述信息生成模型的确定方法及装置、介质、设备
US11978271B1 (en) * 2023-10-27 2024-05-07 Google Llc Instance level scene recognition with a vision language model
CN118968081A (zh) * 2024-07-05 2024-11-15 阿里巴巴(中国)有限公司 图像处理方法及装置

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113449801A (zh) * 2021-07-08 2021-09-28 西安交通大学 一种基于多级图像上下文编解码的图像人物行为描述生成方法
CN114627353A (zh) * 2022-03-21 2022-06-14 北京有竹居网络技术有限公司 一种图像描述生成方法、装置、设备、介质及产品
CN115359323A (zh) * 2022-08-31 2022-11-18 北京百度网讯科技有限公司 图像的文本信息生成方法和深度学习模型的训练方法
CN116050496A (zh) * 2023-01-28 2023-05-02 Oppo广东移动通信有限公司 图片描述信息生成模型的确定方法及装置、介质、设备
US11978271B1 (en) * 2023-10-27 2024-05-07 Google Llc Instance level scene recognition with a vision language model
CN118968081A (zh) * 2024-07-05 2024-11-15 阿里巴巴(中国)有限公司 图像处理方法及装置

Also Published As

Publication number Publication date
CN118968081A (zh) 2024-11-15

Similar Documents

Publication Publication Date Title
WO2026007670A1 (zh) 图像处理方法及装置
KR20210153009A (ko) 광고를 자동 생성하는 방법, 장치, 기기 및 컴퓨터 판독가능 저장매체
CN113704388A (zh) 多任务预训练模型的训练方法、装置、电子设备和介质
CN110379020B (zh) 一种基于生成对抗网络的激光点云上色方法和装置
CN117079299A (zh) 数据处理方法、装置、电子设备及存储介质
CN118052907B (zh) 一种文本配图生成方法和相关装置
CN114581543B (zh) 一种图像描述方法、装置、设备、存储介质
CN115114395A (zh) 内容检索及模型训练方法、装置、电子设备和存储介质
CN113987167B (zh) 基于依赖感知图卷积网络的方面级情感分类方法及系统
CN118397643B (zh) 一种图像处理方法、装置、设备及可读存储介质
CN116932788B (zh) 封面图像提取方法、装置、设备及计算机存储介质
CN115168609A (zh) 一种文本匹配方法、装置、计算机设备和存储介质
CN118429658B (zh) 信息抽取方法以及信息抽取模型训练方法
CN120336859A (zh) 模型训练方法、电子设备及计算机可读存储介质
CN117390224A (zh) 视觉语音问答模型的训练方法、装置、交互方法及系统
WO2026001219A1 (zh) 视频生成、虚拟对象的运动视频生成、视频编辑、视频生成模型训练及基于视频生成模型的信息处理方法
WO2025163427A1 (zh) 任务处理方法、自动问答方法以及图像处理方法
CN116758558B (zh) 基于跨模态生成对抗网络的图文情感分类方法及系统
CN117350301A (zh) 一种基于多模态融合大模型的舆情事件监测方法
CN117216710A (zh) 多模态自动标注方法、标注模型的训练方法及相关设备
CN116978007A (zh) 图像内容的识别方法、装置、电子设备、存储介质及产品
CN112861474B (zh) 一种信息标注方法、装置、设备及计算机可读存储介质
WO2025260927A1 (zh) 信息抽取方法及装置
CN118864042A (zh) 物品信息分析模型的生成方法、生成装置和电子设备
WO2025026012A1 (zh) 视频检索方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25832024

Country of ref document: EP

Kind code of ref document: A1