WO2023084833A1 - 画像処理装置、画像処理方法、及びプログラム - Google Patents
画像処理装置、画像処理方法、及びプログラム Download PDFInfo
- Publication number
- WO2023084833A1 WO2023084833A1 PCT/JP2022/025412 JP2022025412W WO2023084833A1 WO 2023084833 A1 WO2023084833 A1 WO 2023084833A1 JP 2022025412 W JP2022025412 W JP 2022025412W WO 2023084833 A1 WO2023084833 A1 WO 2023084833A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- image
- text
- data
- feature amount
- unit
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
- G06V10/44—Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/50—Information retrieval; Database structures therefor; File system structures therefor of still image data
- G06F16/56—Information retrieval; Database structures therefor; File system structures therefor of still image data having vectorial format
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/50—Information retrieval; Database structures therefor; File system structures therefor of still image data
- G06F16/58—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
- G06F16/583—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/74—Image or video pattern matching; Proximity measures in feature spaces
- G06V10/761—Proximity, similarity or dissimilarity measures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/80—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
- G06V10/806—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level of extracted features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
- G06V30/19—Recognition using electronic means
- G06V30/191—Design or setup of recognition systems or techniques; Extraction of features in feature space; Clustering techniques; Blind source separation
- G06V30/19173—Classification techniques
Definitions
- the present disclosure relates to an image processing device, an image processing method, and a program.
- This image classification is, for example, classifying from some image (medium) whether the image or a specific object in the image is a pigeon or a swallow.
- Non-Patent Document 1 an image of a pigeon and text data that is a sentence describing the pigeon in the image are used.
- the present invention has been made in view of the above points, and an object of the present invention is to extract a multimodal feature amount compared to the conventional art.
- the invention according to claim 1 is an image processing apparatus for extracting a feature amount of image data, comprising: an image understanding unit for extracting an image feature amount by vectorizing an image pattern of the image data; a text understanding unit for vectorizing a text pattern of accompanying text data attached to the image data to extract a text feature quantity; a feature amount mixing unit that projects the text feature amount into the same vector space and mixes the image feature amount and the text feature amount to generate a mixed feature amount as the feature amount. It is a device.
- FIG. 1 is a schematic diagram of a communication system according to an embodiment
- FIG. 2 is a hardware configuration diagram of an image classification device and a communication terminal
- FIG. 1 is a functional configuration diagram of an image classification device according to an embodiment of the present invention
- FIG. 3 is a detailed functional configuration diagram of a feature extraction unit in the image classification device
- FIG. 4 is a detailed functional configuration diagram of a text generation unit in the feature extraction unit
- FIG. 4 is a flow chart showing processing executed by an image classification device in a training (learning) phase
- 4 is a flowchart showing detailed processing executed by a feature extraction unit
- 4 is a flow chart showing processing performed by an image classification device in an inference phase
- FIG. 9 is a diagram showing experimental results.
- FIG. 1 is a schematic diagram of a communication system according to an embodiment of the invention.
- a communication system 1 of this embodiment is constructed by an image classification device 3 and a communication terminal 5 .
- the communication terminal 5 is managed and used by the user Y.
- the image classification device 3 and the communication terminal 5 can communicate via a communication network 100 such as the Internet.
- the connection form of the communication network 100 may be either wireless or wired.
- the image classification device 3 is composed of one or more computers. When the image classification device 3 is composed of a plurality of computers, it may be indicated as “image classification device” or "image classification system”.
- the image classification device 3 is a device that classifies images by AI (Artificial Intelligence). This image classification is, for example, classifying from some image (medium) whether the image or a specific object in the image is a pigeon or a swallow. Then, the image classification device 3 outputs classification result data as a result of image classification. As an output method, by transmitting the classification result data to the communication terminal 5, the communication terminal 5 can display or print a graph or the like related to the classification result data, or a display connected to the image classification device 3 can be used to display or print the above graph or the like. or printing the graph or the like with a printer or the like connected to the image classification device 3 .
- AI Artificial Intelligence
- the communication terminal 5 is a computer, and although a notebook computer is shown as an example in FIG. 1, the communication terminal 5 is not limited to a node type and may be a desktop computer. Also, the communication terminal may be a smart phone or a tablet terminal. In FIG. 1 , user Y operates communication terminal 5 .
- FIG. 2 is a hardware configuration diagram of an image classification device and a communication terminal.
- the image classification device 3 has a processor 301 , a memory 302 , an auxiliary storage device 303 , a connection device 304 , a communication device 305 and a drive device 306 .
- Each piece of hardware constituting the image classification device 3 is interconnected via a bus 307 .
- the processor 301 serves as a control unit that controls the entire image classification device 3, and has various arithmetic devices such as a CPU (Central Processing Unit).
- the processor 301 reads various programs onto the memory 302 and executes them.
- the processor 301 may include a GPGPU (General-purpose computing on graphics processing units).
- the memory 302 has main storage devices such as ROM (Read Only Memory) and RAM (Random Access Memory).
- the processor 301 and the memory 302 form a so-called computer, and the processor 301 executes various programs read onto the memory 302, thereby realizing various functions of the computer.
- the auxiliary storage device 303 stores various programs and various information used when the various programs are executed by the processor 301 .
- the connection device 304 is a connection device that connects an external device (for example, the display device 310 and the operation device 311 ) and the image classification device 3 .
- a communication device 305 is a communication device for transmitting and receiving various information to and from another device.
- a drive device 306 is a device for setting a recording medium 330 .
- the recording medium 330 here includes media for optically, electrically, or magnetically recording information such as CD-ROMs (Compact Disc Read-Only Memory), flexible discs, and magneto-optical discs.
- the recording medium 330 may also include a semiconductor memory that electrically records information, such as a ROM (Read Only Memory) and a flash memory.
- auxiliary storage device 303 Various programs to be installed in the auxiliary storage device 303 are installed by, for example, setting the distributed recording medium 330 in the drive device 306 and reading the various programs recorded in the recording medium 330 by the drive device 306. be done. Alternatively, various programs installed in the auxiliary storage device 303 may be installed by being downloaded from the network via the communication device 305 .
- FIG. 2 shows the hardware configuration of the communication terminal 5, but since each configuration is the same except that the reference numerals are changed from the 300s to the 500s, description thereof will be omitted.
- FIG. 3 is a functional configuration diagram of the image classification device according to the embodiment of the present invention.
- the image classification device 3 has an input unit 30, a reading unit 31, a selection unit 32, a feature extraction unit 33, a similarity calculation unit 34, a loss calculation unit 35, a parameter update unit 36, and an output unit 39. ing. These units are functions realized by instructions from the processor 301 in FIG. 2 based on programs.
- learning models A and B are stored in the memory 302 or the auxiliary storage device 303 in FIG.
- the learning model A is constructed from a large number of image similarity parameters described later.
- the learning model B is constructed from a large number of text generation probability parameters, which will be described later.
- the memory 302 or the auxiliary storage device 303 in FIG. 2 stores a large number of image data that are candidates for support data as teacher data. Text data indicating the content of the image is attached to each of the large number of image data. That is, one pair of support data consists of image data and accompanying text data, and a large amount of pairs of support data are stored in the memory 302 or the auxiliary storage device 303 in FIG.
- one pair of support data includes image data of a pigeon and text data accompanying this image data, which is a sentence describing the pigeon appearing in the image.
- the text data attached to this image data will be referred to as "associated text data”.
- “accompanying” includes the case where text data is added to image data, and the case where text data and image data are separately input or output and associated with each other.
- Text data accompanying image data may be generated based on the image data by the image classification device 3 (generated text data) and added to the image data.
- the input unit 30 inputs image data, which is query data as classification target (evaluation target) data for training or inference.
- image data which is query data as classification target (evaluation target) data for training or inference.
- the input unit 30 inputs query data transmitted from the communication terminal 5 by the user Y to the image classification device 3 to the image classification device.
- Associated text data accompanies the image data, which is the query data. That is, one pair of query data is composed of the image data and the accompanying text data.
- the accompanying text data is always accompanied, but in the case of the inference phase, the accompanying text data may not be accompanying.
- As a method of accompanying the accompanying text data there are cases where it is captioned in the image data and cases where it is manually input by the user Y.
- FIG. In many machine learning models, humans cannot intervene in image classification inference, but by allowing user Y to input text data, user Y can intervene in image classification inference. .
- the reading unit 31 reads, from the memory 302 or the auxiliary storage device 303 in FIG. 2, a group of support data candidates (M types and j pairs for each type) to be compared with the query data.
- M is 100 and j is 60.
- a total of 6000 pairs will be read.
- M is 100 and j is 60 is an example, M may be more than 100 or less than 100, and j may be more than 60 or less than 60.
- the selection unit 32 randomly selects N types of k pairs of support data to be compared with the query data from the group of support data candidates.
- the following description will be made on the assumption that, for example, support data with 5 types of N and 1 pair of k (a total of 5 pairs) are randomly selected.
- This method of selecting one pair of each of the five types of support data is generally performed, but the selection unit 32 does not necessarily need to select one pair of each of the five types of support data. For example, there may be 2 pairs of 10 types (20 pairs in total).
- the training support data is given information indicating the type of subject (also referred to as "class") in the image of the image data. For example, if the image is an image of a bird, it indicates the type of bird such as "pigeon", "hawk", "swallow”.
- the feature extraction unit 33 extracts an image feature amount from image data in one pair, and further extracts a text feature amount from text data in the same pair. Furthermore, the feature extraction unit 33 mixes the image feature amount and the text feature amount to generate a mixed feature amount. The feature extraction unit 33 also generates text data from the image feature amount.
- the text data generated from the image feature amount will be referred to as "generated text data”. That is, the generated text data is image-derived text data, and is different in type from text-derived accompanying text data.
- FIG. 4 is a detailed functional configuration diagram of the feature extraction unit in the image classification device.
- the feature extraction section 33 has an image understanding section 41 , a text generation section 42 , a text understanding section 43 and a feature quantity mixing section 44 .
- Arbitrary neural networks can be used for the image understanding unit 41, the text generation unit 42, the feature amount mixing unit 44, and the similarity calculation unit .
- the image understanding unit 41 uses a four-layer CNN (Convolutional Neural Network). By pre-learning the text generation unit 42 and the text understanding unit 43, the text generation ability and the text understanding ability are improved.
- CNN Convolutional Neural Network
- the image understanding unit 41 acquires image data (an example of first image data) from the query data from the input unit 30, and acquires from the selection unit 32 a specific one pair out of five types of one pair. image data (an example of second image data) in the support data of . Then, the image understanding unit 41 vectorizes the image pattern of the image data of the query data to extract the image feature amount for the query, and vectorizes the image pattern of the image data of the support data to extract the image feature amount for the support. Extract.
- An image feature is a vector
- the text generator 42 can use any neural network, and RNN (Recurrent Neural Network) and Transformer, which use the image feature as initial values, are common.
- the text generation unit 42 projects the query image feature amount extracted by the image understanding unit 41 onto the vector space of the text data and decodes it to generate generated text data for the query derived from the image. By projecting the image feature amount for support extracted by the unit 41 onto the vector space of the text data and decoding it, generated text data for support derived from the image is generated.
- FIG. 5 is a detailed functional block diagram of the text generator.
- the text generator 42 has a linear transformation layer 421 and a decoder 422. Further, the linear transformation layer 421 holds linear transformation layer parameters 421p, and the decoder 422 holds decoder parameters 422p. The linear transformation layer parameters 421p and the decoder parameters 422p are included in the learning model B shown in FIG.
- the linear transformation layer 421 uses the linear transformation layer parameter 421p to project the image feature quantity acquired from the image understanding unit 41 onto the vector space of the accompanying text data, thereby extracting the feature quantity derived from the image.
- the decoder 422 uses the decoder parameter 422p to generate image-derived generated text data from the feature amount acquired from the linear transformation layer 421 .
- a language model having an encoder-decoder type structure is disclosed in Reference 1, for example.
- An encoder-decoder type structure is a structure in which text is first given as an input, converted into features by the encoder, the features are input to the decoder, and the decoder generates text.
- the existing language model Encoder in Reference 1 is not used, and an arbitrary neural network such as a linear transformation layer is used before the Decoder. Add to With this configuration, it is possible to convert the image feature quantity into a feature quantity suitable for the language model, input it to the Decoder, and generate text.
- the text understanding unit 43 acquires accompanying text data from the query data from the input unit 30, and from the selection unit 32, one specific pair of support data out of one pair of five types. Gets the accompanying text data. Then, the text understanding unit 43 vectorizes the text pattern of the accompanying text data of the query data to extract the text feature amount for the query, and vectorizes the text pattern of the accompanying text data of the support data to extract the text feature for support. Extract quantity.
- the text understanding unit 43 converts text data into vectors using an existing language model such as BERT (Bidirectional Encoder Representations from Transformers).
- BERT Bidirectional Encoder Representations from Transformers
- accompanying text data is attached to image data in the training phase, but accompanying text data may not be attached to image data in the inference phase.
- the text understanding unit 43 treats (deems) the image-derived query text data generated by the text generation unit 42 as accompanying text data. Extract features of data.
- the feature amount mixing unit 44 projects the query image feature amount extracted by the image understanding unit 41 and the query text feature amount extracted by the text understanding unit 43 onto the same vector space, By mixing the image feature amount for query and the text feature amount for query, a mixed feature amount as a feature amount for query is generated.
- the feature amount mixing unit 44 projects the image feature amount for support extracted by the image understanding unit 41 and the text feature amount for support extracted by the text understanding unit 43 into the same vector space, and By mixing the image feature amount for support and the text feature amount for support, a mixed feature amount as a feature amount for support is generated.
- the vector space of one feature amount is projected onto the other feature amount, and where the other feature amount is projected onto a third vector space different from each other.
- the feature amount mixing unit 44 can reflect both the image feature amount and the text feature amount in the similarity calculation.
- the feature mixing unit 44 can use any neural network that accepts both image features and text features as inputs.
- the following model is used as the feature quantity mixing unit 44 .
- ximage be the image feature quantity
- xLang be the text feature quantity output by the text understanding unit 43 .
- MLP Multilayer perceptron
- Linear be a linear transformation layer to two dimensions.
- [ ; ] be an operation to connect vectors vertically.
- the vector h output by the feature quantity mixing unit 44 is represented by (Equation 1), (Equation 2), and (Equation 3) as follows.
- the feature amount mixing unit 44 projects the text feature amount output by BERT by MLP into the same space as the image feature amount (z Lang ), using (Formula 1).
- the feature amount mixing unit 44 dynamically determines the importance of the image feature amount and the text feature amount from ⁇ image and ⁇ Lang using (Formula 2).
- ⁇ image and ⁇ Lang are guaranteed to be non-negative numbers summing to 1 by the softmax operation.
- the degree to which the accompanying text data attached to the image data affects the classification result is ⁇ image and ⁇ Lang are dynamically determined to increase.
- the user can manually change the degree to which the text entered by the user is reflected in the classification results.
- Linear is the operation of multiplying the weight matrix from the left and adding the bias vector. The weight matrix and bias vector in the Linear operation are included in the learning model A's image similarity parameter and the learning model B's text generation probability parameter.
- the feature amount mixing unit 44 determines the feature amount to be output by a weighted sum according to the degree of importance using (Formula 3).
- the image similarity parameter of the learning model A is used when the image understanding unit 41, the text understanding unit 43, and the feature amount mixing unit 44 execute each process.
- the text generation probability parameter of learning model B is used when the image understanding unit 41 and the text generation unit 42 execute each process.
- the text generation probability parameter of learning model B is not used.
- the text generation probability parameter of learning model B is used and updated by training (learning). This is done so that the text generator 42 can generate the generated text data even when the accompanying text data is not attached to the image data in the inference phase. This is also because training (learning) the learning model B has a positive effect of improving the comprehension ability of the image understanding unit 41 that uses the text generation probability parameter.
- the similarity calculation unit 34 compares the mixed feature amount for query and the mixed feature amount for support to calculate the image similarity.
- this image similarity is output to the output unit 39 and used as classification result data for image classification.
- this image similarity is output to the loss calculator 35 .
- the similarity calculator 34 is a bilinear layer. Now consider N-way k-shot image classification. The similarity calculation unit 34 first gives k supporting feature amounts (vectors) for each class. A vector obtained by averaging these is used as a class feature. Let X be a matrix in which N class feature values (vectors) are arranged. Let y be the feature value of the query data and W be the learnable parameter. At this time, the classification score for each class of query data is expressed as follows.
- Each component of this vector indicates the probability that the query data belongs to each class.
- a loss calculator 35 calculates a loss function value from the image similarity. Further, the loss calculation unit 35 calculates a loss function value from the generated text data of the query data/support data, the generation probability distribution of the query data/support data, and the accompanying text data of the query data/support data.
- the loss function calculated by the loss calculator 35 can use the classification score of the similarity calculator 34 or any loss related to text generation.
- Cross-Entropy Loss and negative log-likelihood function are typically used.
- the parameter update unit 36 updates the neural network of the feature extraction unit 33 and the similarity calculation unit 34 based on the loss function value calculated by the loss calculation unit 35 from the image similarity calculated by the similarity calculation unit 34.
- the image similarity parameter of learning model A is updated.
- the loss calculation unit 35 performs learning so that the degree of similarity between the image data of the support data and the image data of the query data is reduced, and further, the degree of similarity with the incorrect image is increased.
- the parameter updating unit 36 updates the text generation probability parameter of the learning model B of the neural network constituting the feature extracting unit 33 and the similarity calculating unit 34 based on the loss function value calculated by the loss calculating unit 35.
- the loss calculator 35 performs learning so as to increase the probability that the generated text data is similar to the accompanying text data.
- the parameter updater 56 calculates the slope of the loss based on the loss calculated by the loss calculator 35 and updates the parameters.
- FIG. In addition, it divides into a training (learning) phase and an inference phase, and demonstrates.
- FIG. 6 is a flow chart showing the processing performed by the image classification device in the training (learning) phase.
- the input unit 30 inputs training teacher data (query data) (S10).
- the reading unit 31 reads out a candidate group of teacher data (support data) for training (S11).
- the selection unit 32 randomly selects one pair of five types of support data (image data and accompanying text data) as teacher data from the candidate group (S12).
- the selection unit 32 also selects an arbitrary number of pairs from the same five types as query data.
- the selection unit 32 defines the same type of support data as the correct answer for the query data, and defines different types of support data as the incorrect answer for the query data. By defining , the data defining the correct or incorrect answer is added to the support data.
- the support data indicating "pigeon” is defined as the correct answer
- the support data indicating the other types (classes) is defined as the incorrect answer. It should be noted that the correct answer or the incorrect answer may be defined by the reading unit 31 .
- the feature extraction unit 33 generates a mixed feature amount for query based on the query data acquired from the input unit 30, and extracts the support data of 5 types and 1 pair (a total of 5 pairs) selected by the selection unit 32.
- a mixed feature quantity for support is generated based on a predetermined one of the support data (S13).
- the feature extraction unit 33 receives defined set data of correct or incorrect answers (query data, support data, and definition data of correct or incorrect answers), and extracts the query data and support data included in the set data. is calculated and output to the similarity calculation unit.
- a vector obtained by averaging the image feature amounts of the image data of each pair may be used as the image feature amount of the support data.
- FIG. 7 is a flowchart showing detailed processing executed by the feature extraction unit.
- the image understanding unit 41 extracts each image feature amount (image feature amount for query, image feature amount for support) based on each image data of query data and support data. (S131).
- the text generation unit 42 generates each generation text data based on each image feature amount (S132).
- steps S133 and S135, which will be described later, are not executed, and subsequently, the text understanding unit 43 acquires each text feature quantity (text feature quantity for query, text feature quantity for text) is extracted (S134).
- the feature amount mixing unit 44 mixes the image feature amount for query and the text feature amount for query to generate a mixed feature amount for query, and mixes the image feature amount for support and the text feature amount for support. Then, a mixed feature quantity for support is generated (S136).
- the similarity calculation unit 34 compares the mixed feature amount for query (an example of the first mixed feature amount) and the mixed feature amount for support (an example of the second mixed feature amount). Then, the image similarity is calculated (S14). At this time, the similarity calculation unit 34 calculates the similarity of each pair of query data and support data included in the set data, and passes it to the loss calculation unit.
- the feature extracting unit 33 determines whether or not the calculation of the similarities for all five pairs out of the five types of one pair of support data (five pairs in total) selected by the selecting unit 32 has been completed (S15). ). Then, when the feature extraction unit 33 determines that the calculation of the similarities for all the five pairs of support data has not been completed (S15; NO), the process returns to step S13, and the calculation of the similarities has not been completed. Step S13 and subsequent steps are performed on the support data. As for the query data acquired from the input unit 30, since the mixed feature amount has already been generated, the reprocessing after step S13 is not performed.
- the loss calculation unit 35 calculates the loss (S16). .
- the loss calculation unit 35 calculates the loss based on the similarity of each pair of query data and support data included in each set data, and the definition data of correct or incorrect answers for each pair of support data with respect to the query data. do. Note that this degree of similarity includes the degree of similarity between images and the degree of similarity between accompanying texts.
- the parameter update unit 36 uses L cntr to determine the text feature amounts of the “attached text data in the teacher data” and the “generated text data generated by the text generation unit” attached to images of the same class.
- the loss function L is a value obtained by summing the above four Ls according to (Equation 5).
- the parameter updating unit 36 calculates the gradient of loss, and updates (trains) the image similarity parameter of learning model A and the text generation probability parameter of learning model B (S17). At this time, the parameter updating unit 36 updates the parameters so as to minimize the loss.
- the selection unit 32 determines whether or not a specified number of selections (for example, 20 times) has been completed (S18). For example, when the selection unit 32 selects 20 times as the prescribed number of times, 5 pairs of support data are selected in one selection, and thus 100 pairs of support data are selected in total. However, since the selection unit 32 randomly selects one pair of five types of support data (five pairs in total) from the candidate group, the same support data may be selected multiple times.
- a specified number of selections for example, 20 times
- 5 pairs of support data are selected in one selection, and thus 100 pairs of support data are selected in total.
- the selection unit 32 randomly selects one pair of five types of support data (five pairs in total) from the candidate group, the same support data may be selected multiple times.
- step S18 when the selection unit 32 determines that the specified number of selections has not been completed (S18; NO), the process returns to step S12, and the selection unit 32 selects a new random candidate from the candidate group. 1 pair of 5 types (total 5 pairs) of support data are selected, and then the processing from step S13 onwards is performed.
- step S18 when the selection unit 32 determines that the specified number of selections has been completed (S18; YES), the processing of the training phase shown in FIG. 6 ends.
- FIG. 8 is a flow chart showing the processing performed by the image classifier in the inference phase.
- the input unit 30 inputs query data, which is data to be classified for inference (S30).
- the reading unit 31 reads support data for inference (S31).
- the feature extraction unit 33 generates a mixed feature amount for query based on the query data, which is the classification target data acquired from the input unit 30, and selects one pair of five types selected by the selection unit 32 (five pairs in total). ), a mixed feature amount for support is generated based on a predetermined one of the support data (S32).
- FIG. 7 is a flowchart showing detailed processing executed by the feature extraction unit.
- the image understanding unit 41 extracts each image feature amount (image feature amount for query, image feature amount for support) based on each image data of query data and support data. (S131).
- the text generation unit 42 generates each generation text data based on each image feature amount (S132). In the inference phase, steps S133 and S135, which will be described later, are executed.
- ⁇ Supplement C> Note that in the inference phase, generated text data is generated using beam-search. Therefore, if the beam width is 10, for example, the top 10 tokens at each time can be used as generation candidates. However, beam-search is computationally heavy and cannot be used in the training phase. Therefore, in order to suppress the divergence between the training phase and the inference phase, the text generation unit 42 performs "language generation" processing.
- the process of "language generation” is the process of generating generated text data by combining the greedy method and random sampling in the training phase.
- the text generator 42 selects the highest token at each time and generates generated text data.
- the text generation unit 42 selects a token from a predetermined number (for example, top 20) of tokens at each time by sampling to generate generated text data.
- the text generating section 42 passes the generated text data to the text understanding section 43.
- FIG. At this time, the text generation unit 42 performs a beam search with a length penalty of 0.5 and a beam width of 5, and generates generated text data that maximizes the classification score that the query data belongs to each class. It is passed to the text understanding unit 43 .
- the text understanding unit 43 determines whether both the query data and the support data include accompanying text data, that is, whether both the query data image data and the support data image data are accompanied by accompanying text data. It is determined whether there is (S133). Then, the text understanding unit 43 determines that both the query data and the support data include accompanying text data, that is, both the image data of the query data and the image data of the support data are accompanied by the accompanying text data. If it is determined that there is (S133; YES), the text understanding unit 43 calculates each text feature quantity (text feature quantity for query, text feature for text amount) is extracted (S134).
- step S133 if the text understanding unit 43 determines that both the query data and the support data do not include accompanying text data, that is, if both the query data image data and the support data image data When it is determined that accompanying text data is not attached (S133; NO), the text understanding section 43 performs the following processing.
- the text understanding unit 43 extracts the text feature amount based on the accompanying text of the query data, Based on this, the text feature quantity is extracted (S135).
- the text understanding unit 43 extracts the text feature amount based on the accompanying text of the support data, and extracts the generated text of the query data. (S135).
- the text understanding unit 43 performs , extracts the respective text features (S135).
- the text understanding unit 43 first maps (projects) x image onto the same vector space as the text feature quantity, according to (Equation 6).
- LayerNorm indicates Layer Normalization (Reference 2). ⁇ Reference 2> Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
- the text understanding unit 43 passes the resulting representation to the text generation unit 42 as a sequence of length 1. Then, the text generation unit 42 autoregressively generates the j-th token tj according to the probability pj shown in (Equation 7) below. That is, the probability pj indicates how likely the predetermined (j-th) token tj associated with the generated text data generated by the text generation unit 42 is to be correct (probability).
- the text understanding unit 43 extracts the text feature quantity xLang shown in (Equation 8). perform pooling.
- the feature amount mixing unit 44 mixes the query image feature amount and the query text feature amount to generate a mixed feature amount for query, and also generates a mixed feature amount for query. and text features for support are mixed to generate a mixed feature for support (S136).
- the similarity calculation unit 34 compares the mixed feature amount for query (an example of the first mixed feature amount) and the mixed feature amount for support (an example of the second mixed feature amount). Then, the image similarity is calculated (S33).
- the feature extraction unit 33 determines whether the comparison of all five pairs of support data out of the five pairs of support data selected by the selection unit 32 (five pairs in total) has been completed ( S34). Then, when the feature extraction unit 33 determines that the comparison of all five pairs of support data has not been completed (S35; NO), the process returns to step S32, and five types of one pair of support data (five pairs in total) are extracted. Step S32 and subsequent steps are performed for the support data for which the comparison of . As for the query data, which is the classification target data acquired from the input unit 30, since the mixed feature amount has already been generated, the reprocessing after step S32 is not performed.
- step S34 when the feature extraction unit 33 determines that the comparison of all five pairs of support data has been completed (S34; YES), the output unit 39 outputs a , and outputs classification result data indicating the classification result (S35).
- the image related to the classification target data is an image of a pigeon, and there is a 90% chance that it is a pigeon image and a 10% chance that it is another bird image. It is shown.
- the result of executing all the proposed methods according to this embodiment is the value of LIDE in the first row of the first row.
- the experimental results when the above ⁇ Supplement A>, ⁇ Supplement B>, ⁇ Supplement C>, and ⁇ Supplement D> are not performed, respectively, are the 2nd row, 1st row, the 2nd row, 2nd row, 3 It is shown in the fourth line of the third column and the third line of the third column.
- the experimental results when all of the above ⁇ Supplement A>, ⁇ Supplement B>, ⁇ Supplement C>, and ⁇ Supplement D> are not performed are shown in the fifth line of the second row.
- all the four elements of the proposed method of ⁇ Supplement A>, ⁇ Supplement B>, ⁇ Supplement C>, and ⁇ Supplement D> contribute to performance improvement.
- the image classification device 3 mixes the image feature amount of the image data and the text feature amount of the accompanying text data attached to the image data to obtain the mixed feature amount. Generate.
- the image classification device 3 as a feature extraction device, can extract multimodal feature quantities compared to simply comparing feature quantities between image data and comparing text data. Effective.
- the image classification device 3 extracts feature amounts related to image data with higher accuracy, thereby achieving the effect of being able to perform image classification with higher accuracy.
- images can be classified with high accuracy by the following processing, and the user can intervene in the classification result by inputting accompanying text data in the inference phase. .
- ⁇ Supplement A> when text data is used to supplement information in a small number of cases image classification task, the loss function of image classification is used together to suppress the divergence regarding the text data that can be used in the training phase and the inference phase.
- the learning of the text understanding unit 43 progresses so as to output feature amounts that capture minute differences in text data through contrast learning.
- ⁇ Supplement C> the divergence between the text generation methods in the training phase and the inference phase is suppressed by generating generated text data by random sampling.
- the text generation unit 42 performs learning considering the performance improvement of image classification by pooling using the text generation score (generated score for each token).
- the present invention is not limited to the above-described embodiments, and may be configured or processed (operations) as described below.
- the image classification device 3 can be implemented by a computer and a program, and the program can be recorded on a (non-temporary) recording medium or provided via the communication network 100 .
- the image classification device 3 is shown, but if the feature extraction unit 33 is specialized, it can be expressed as a feature extraction device. Further, both the image classification device 3 and the feature extraction device can be expressed as image processing devices.
- the number of data can be inflated by performing rule-based paraphrasing of accompanying text data to be input.
- rephrasing there is a rephrasing of "This bird is large” by rephrasing "big” in "This bird is big” to "large”.
- An image processing device for extracting a feature amount of image data, an image understanding step of vectorizing the image pattern of the image data and extracting an image feature quantity; a text understanding step of vectorizing a text pattern of accompanying text data attached to the image data to extract a text feature quantity; By projecting the image feature amount extracted by the image understanding step and the text feature amount extracted by the text understanding step into the same vector space and mixing the image feature amount and the text feature amount, the A feature quantity mixing step of generating a mixed feature quantity as a feature quantity; An image processing device that executes
- the image understanding step, the text understanding step, and the feature amount mixing step are each realized by a neural network, and the image understanding step, the text understanding step, and the feature amount mixing step are based on model parameters of the neural network. 2.
- the image processing device according to additional item 2, The processor a text generating step of generating generated text data by projecting the image feature amount extracted by the image understanding step onto a vector space of the accompanying text data; a parameter update step of updating text generation probability parameters included in the model parameters based on the generated text data generated by the text generation step and the accompanying text data; An image processing device that executes
- the parameter updating step includes processing for updating the text generation probability parameter based on the loss based on the image feature quantity and the accompanying text data, and the loss based on the image feature quantity and the generated text data.
- Item 5 The image processing apparatus according to item 4.
- the parameter updating step updates the text generation probability parameter so that the text feature amount of the accompanying text data and the text feature amount of the generated text data for image data of the same class are close to each other; 5.
- the image processing apparatus according to item 4 further comprising a process of updating the text generation probability parameter so that the text feature amount of the accompanying text data and the text feature amount of the generated text data become distant.
- the text generating step includes a process of generating the generated text data by random sampling from a predetermined number of high-order tokens at each time and generating the generated text data at normal time in combination. image processing device.
- the text understanding step includes a process of extracting the text feature amount by performing weighted pooling using a probability indicating the probability of a predetermined token related to the generated text data, 5.
- the image processing device according to additional item 4.
- An image processing method executed by an image processing device for extracting a feature amount of image data The image processing device is an image understanding step of vectorizing the image pattern of the image data and extracting an image feature quantity; a text understanding step of vectorizing a text pattern of accompanying text data attached to the image data to extract a text feature quantity; Projecting the image feature amount extracted by the image understanding step and the text feature amount extracted by the text understanding step into the same vector space, and mixing the image feature amount and the text feature amount, A feature quantity mixing step of generating a mixed feature quantity as a feature quantity; An image processing method that performs
- Appendix 11 A non-transitory recording medium recording a program that causes a computer to execute the method according to claim 10.
- Image classification device an example of image processing device
- communication terminal 30 input section 31 reading unit 32 selection part 33 feature extractor 34 similarity calculator 35 loss calculator 36 parameter update unit 39
- Output section 41
- Image understanding unit 42
- text generator 43
- Text Comprehension 44
- feature quantity mixing unit 422 decoder
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Multimedia (AREA)
- Evolutionary Computation (AREA)
- Databases & Information Systems (AREA)
- Artificial Intelligence (AREA)
- Computing Systems (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Software Systems (AREA)
- Data Mining & Analysis (AREA)
- General Engineering & Computer Science (AREA)
- Library & Information Science (AREA)
- Image Analysis (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
Description
まず、図1を用いて、本実施形態の通信システム1の構成の概略について説明する。図1は、本発明の実施形態に係る通信システムの概略図である。
次に、図2を用いて、画像分類装置3及び通信端末5のハードウェア構成を説明する。図2は、画像分類装置及び通信端末のハードウェア構成図である。
次に、図3を用いて、画像分類装置の機能構成について説明する。図3は、本発明の実施形態に係る画像分類装置の機能構成図である。
ここで、図4を用いて、画像分類装置における特徴抽出部を詳細に説明する、図4は、画像分類装置における特徴抽出部の詳細な機能構成図である。
ここで、図5を用いて、テキスト生成部42について、更に詳細に説明する。図5は、テキスト生成部の詳細な機能ブロック図である。
<参考文献1>Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Encoder-Decoder型の構造とは、まずテキストを入力として与えられてEncoderによって特徴量に変換し、Decoderにその特徴量を入力し、Decoderがテキストを生成する構造のことをいう。本実施形態においては、テキスト生成部62に画像特徴量が入力されるため、参考文献1における既存の言語モデルのEncoderを使用せず、代わりに線形変換層などの任意のニューラルネットワークをDecoderの前に追加する。この構成によって、画像特徴量を言語モデルに適した特徴量に変換し、Decoderに入力し、テキストを生成することが可能になる。
続いて、図6乃至図8を用いて、本実施形態の処理又は動作について詳細に説明する。なお、訓練(学習)フェーズと推論フェーズに分けて説明する。
まずは、図6及び図7を用いて、訓練フェーズについて説明する。図6は、訓練(学習)フェーズにおいて画像分類装置が実行する処理を示すフローチャートである。
なお、教師データ中の付随テキストデータが利用可能な訓練フェーズでは、教師データ中の付随テキストデータは、テキスト理解部43に入力されている。しかし、推論フェーズでは、テキスト生成部42が生成した生成テキストデータが入力される可能性があるため、訓練フェーズと推論フェーズで乖離が発生してしまう。そこで、本実施形態では、画像特徴量及び付随テキストデータから計算されたクロスエントロピー損失Lclass,goldだけでなく、画像特徴量及びテキスト生成部42が生成した生成テキストデータとから計算されたクロスエントロピー損失Lclass,genも併用して学習する。この処理により、訓練フェーズと推論フェーズの乖離を抑えることが可能になる。
また、(式4)に示す対照学習の損失Lcntrを利用することでモデルがテキストデータの微細な違いを捉えて特徴量を獲得することも可能である。
は、損失計算部35が、同じようにテキスト生成部42が生成した生成テキストデータの入力に基づいて計算したベクトルである。なお、対照学習(contrastive learning)とは、正例と負例を区別し、入力と正例が近づくように、入力と負例が遠ざかるように行う学習である。ここでは、パラメータ更新部36が、Lcntrを用いて、同じクラスの画像に付随する「教師データ中の付随テキストデータ」と「テキスト生成部が生成した生成テキストデータ」の各テキスト特徴量
は遠ざかるようにテキスト生成確率パラメータを更新する。このように、対照学習を行うことで、テキスト理解部43が両テキストデータの微細な違いを捉えた特徴量を出力するように学習を進めることができるという効果が生じる。
次に、図7及び図8を用いて、訓練フェーズについて説明する。図8は、推論フェーズにおいて画像分類装置が実行する処理を示すフローチャートである。
なお、推論フェーズでは、beam-searchを用いて生成テキストデータが生成される。そのため、例えばbeam幅を10とすると、各時刻の上位10個のトークンを生成候補とすることができる。しかし、beam-searchは計算量が重いため、訓練フェーズに用いることができない。そこで、訓練フェーズと推論フェーズの乖離を抑えるために、テキスト生成部42は、「言語生成」の処理を行う。
なお、テキスト生成部42が、生成テキストデータの生成の損失として、teacher-forcingとクロスエントロピー損失で計算されたLtextのみで学習すると、「教師データ中の付随テキストデータを再現する」ことが学習の目的となり、「画像分類に資する生成テキストデータを生成する」ことは、考慮できていない。これは、テキスト生成部42の離散的処理により勾配グラフが壊れることが原因である。つまり、画像分類によって得られる損失(Lclass,gold,Lclass,gen)から得られる勾配が、誤差逆伝播法によりテキスト生成部42にまで伝播しないことが原因である。これを改善するために、テキスト理解部43が以下に示す処理を行なった。
ここで、LayerNorm はLayer Normalizationを示す(参考文献2)。
<参考文献2>Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
テキスト理解部43は、得られた表現をテキスト生成部42に長さ1の系列として渡す。そして、テキスト生成部42は、自己回帰的にj番目のトークンtjを以下の(式7)で示す確率pjに従い生成する。即ち、確率pjは、テキスト生成部42によって生成された生成テキストデータに係る所定番目(j番目)のトークンtjが、どのくらいの確率で正しそうか(確からしさ)を示す。
続いて、実験設定及び実験結果について説明する。
データセットCaltech-UCSD Birds (CUB)(参考文献3,4)を5-way 1-shot 分類問題として用いた。本データは200種の鳥の品種をクラスとしており、1種に対して40-60枚の画像がある。200種の画像のうち、100種が訓練用、50種が開発用、50種がテスト用である。
<参考文献3>CatherineWah, Steve Branson, PeterWelinder, Pietro Perona, and Serge Belongie. The caltech-ucsd birds-200-2011
dataset. 2011.
<参考文献4>Scott Reed, Zeynep Akata, Honglak Lee, and Bernt Schiele. Learning deep representations of fine-grained visual descriptions. In CVPR, pp. 49-58, 2016.
<実験結果>
図9は、実験結果を示す図である。図9において、本実施形態による提案手法を全て実行した結果が1段目1行目のLIDEの値である。そこから、上述の<補足A>、<補足B>、<補足C>、及び<補足D>をそれぞれ行わない場合の実験結果が、2段目1行目、2段目2行目、3段目4行目、3段目3行目に示されている。また、上述の<補足A>、<補足B>、<補足C>、及び<補足D>の全てを行わない場合の実験結果が、2段目5行目に示されている。これにより、<補足A>、<補足B>、<補足C>、及び<補足D>による提案手法の4要素全てが性能向上に貢献することが確認できた。
以上説明したように本実施形態によれば、画像分類装置3は、画像データの画像特徴量と、画像データに付随している付随テキストデータのテキスト特徴量を混合することで、混合特徴量を生成する。これにより、画像分類装置3は、特徴抽出装置として、単に、画像像データ同士の特徴量の比較、及びテキストデータ同士の比較の場合に比べて、マルチモーダルな特徴量を抽出することができるという効果を奏する。また、画像分類装置3は、より高精度な画像データに関する特徴量を抽出することで、より高精度な画像分類を行うことができるという効果を奏する。
本発明は上述の実施形態に限定されるものではなく、以下に示すような構成又は処理(動作)であってもよい。
上述の実施形態には、以下に示す発明としても表すことができる。
画像データの特徴量を抽出する画像処理装置であって、
前記画像データの画像パターンをベクトル化して画像特徴量を抽出する画像理解ステップと、
前記画像データに付随している付随テキストデータのテキストパターンをベクトル化してテキスト特徴量を抽出するテキスト理解ステップと、
前記画像理解ステップによって抽出された前記画像特徴量と前記テキスト理解ステップによって抽出された前記テキスト特徴量を同じベクトル空間に射影して、前記画像特徴量と前記テキスト特徴量を混合することで、前記特徴量としての混合特徴量を生成する特徴量混合ステップと、
を実行する画像処理装置。
前記画像理解ステップ、前記テキスト理解ステップ、及び前記特徴量混合ステップは、それぞれニューラルネットワークで実現され、前記画像理解ステップ、前記テキスト理解ステップ、及び前記特徴量混合ステップは前記ニューラルネットワークのモデルパラメータに基づいて処理を行う、付記項1に記載の画像処理装置。
付記項2に記載の画像処理装置であって、
前記プロセッサは、
前記特徴量混合ステップによって生成された第1の画像データに係る第1の混合特徴量、及び前記特徴量混合ステップによって生成された第2の画像データに係る第2の混合特徴量の画像類似度を計算する類似度計算ステップと、
前記類似度計算ステップによって計算された前記画像類似度に基づいて、前記モデルパラメータに含まれる画像類似度パラメータを更新するパラメータ更新ステップと、
を実行する画像処理装置。
付記項2に記載の画像処理装置であって、
前記プロセッサは、
前記画像理解ステップによって抽出された前記画像特徴量を前記付随テキストデータのベクトル空間に射影することで生成テキストデータを生成するテキスト生成ステップと、
前記テキスト生成ステップによって生成された前記生成テキストデータと前記付随テキストデータに基づいて、前記モデルパラメータに含まれるテキスト生成確率パラメータを更新するパラメータ更新ステップと、
を実行する画像処理装置。
付記項1に記載の画像処理装置であって、
前記プロセッサは、
前記画像理解ステップによって抽出された前記画像特徴量を前記付随テキストデータのベクトル空間に射影することで生成テキストデータを生成するテキスト生成ステップを実行し、
前記画像データに前記付随テキストデータが付随していない場合には、前記テキスト理解ステップは、前記テキスト生成ステップによって生成された前記生成テキストデータを前記付随テキストデータとすることで、前記テキスト特徴量を抽出する処理を含む、画像処理装置。
前記パラメータ更新ステップは、前記画像特徴量及び前記付随テキストデータに基づく損失、並びに、前記画像特徴量及び前記生成テキストデータに基づく損失に基づいて、前記テキスト生成確率パラメータを更新する処理を含む、付記項4に記載の画像処理装置。
前記パラメータ更新ステップは、同じクラスの画像データに関する前記付随テキストデータのテキスト特徴量及び前記生成テキストデータのテキスト特徴量が近づくように前記テキスト生成確率パラメータを更新すると共に、異なるクラスの画像データに関する前記付随テキストデータのテキスト特徴量及び前記生成テキストデータのテキスト特徴量が遠ざかるように前記テキスト生成確率パラメータを更新する処理を含む、付記項4に記載の画像処理装置。
前記テキスト生成ステップは、各時刻で上位の所定数のトークンからのランダムサンプリングによる前記生成テキストデータの生成と、通常時の前記生成テキストデータの生成とを併用する処理を含む、付記項4に記載の画像処理装置。
前記テキスト理解ステップは、前記生成テキストデータに係る所定番目のトークンの確からしさを示す確率を用いた重み付きプーリングを行うことで、前記テキスト特徴量を抽出する処理を含む、
付記項4に記載の画像処理装置。
画像データの特徴量を抽出する画像処理装置が実行する画像処理方法であって、
前記画像処理装置は、
前記画像データの画像パターンをベクトル化して画像特徴量を抽出する画像理解ステップと、
前記画像データに付随している付随テキストデータのテキストパターンをベクトル化してテキスト特徴量を抽出するテキスト理解ステップと、
前記画像理解ステップによって抽出された前記画像特徴量と前記テキスト理解ステップによって抽出された前記テキスト特徴量を同じベクトル空間に射影して、前記画像特徴量と前記テキスト特徴量を混合することで、前記特徴量としての混合特徴量を生成する特徴量混合ステップと、
を実行する画像処理方法。
コンピュータに、付記項10に記載の方法を実行させるプログラムを記録した非一時的記録媒体。
本特許出願は2021年11月12日に出願した国際出願PCT/JP2021/041801に基づきその優先権を主張するものであり、国際出願PCT/JP2021/041801の全内容を本願に援用する。
3 画像分類装置(画像処理装置の一例)
5 通信端末
30 入力部
31 読出部
32 選択部
33 特徴抽出部
34 類似度計算部
35 損失計算部
36 パラメータ更新部
39 出力部
41 画像理解部
42 テキスト生成部
43 テキスト理解部
44 特徴量混合部
422 デコーダ
Claims (12)
-
画像データの特徴量を抽出する画像処理装置であって、
前記画像データの画像パターンをベクトル化して画像特徴量を抽出する画像理解部と、
前記画像データに付随している付随テキストデータのテキストパターンをベクトル化してテキスト特徴量を抽出するテキスト理解部と、
前記画像理解部によって抽出された前記画像特徴量と前記テキスト理解部によって抽出された前記テキスト特徴量を同じベクトル空間に射影して、前記画像特徴量と前記テキスト特徴量を混合することで、前記特徴量としての混合特徴量を生成する特徴量混合部と、
を有する画像処理装置。
-
前記画像理解部、前記テキスト理解部、及び前記特徴量混合部は、それぞれニューラルネットワークで構成され、前記画像理解部、前記テキスト理解部、及び前記特徴量混合部は前記ニューラルネットワークのモデルパラメータに基づいて処理を行う、請求項1に記載の画像処理装置。
-
請求項2に記載の画像処理装置であって、
前記特徴量混合部によって生成された第1の画像データに係る第1の混合特徴量、及び前記特徴量混合部によって生成された第2の画像データに係る第2の混合特徴量の画像類似度を計算する類似度計算部と、
前記類似度計算部によって計算された前記画像類似度に基づいて、前記モデルパラメータに含まれる画像類似度パラメータを更新するパラメータ更新部と、
を有する画像処理装置。
-
請求項2に記載の画像処理装置であって、
前記画像理解部によって抽出された前記画像特徴量を前記付随テキストデータのベクトル空間に射影することで生成テキストデータを生成するテキスト生成部を有し、
前記テキスト生成部によって生成された前記生成テキストデータと前記付随テキストデータに基づいて、前記モデルパラメータに含まれるテキスト生成確率パラメータを更新するパラメータ更新部と、
を有する画像処理装置。
-
請求項1に記載の画像処理装置であって、
前記画像理解部によって抽出された前記画像特徴量を前記付随テキストデータのベクトル空間に射影することで生成テキストデータを生成するテキスト生成部を有し、
前記画像データに前記付随テキストデータが付随していない場合には、前記テキスト理解部は、前記テキスト生成部によって生成された前記生成テキストデータを前記付随テキストデータとすることで、前記テキスト特徴量を抽出する画像処理装置。
-
前記パラメータ更新部は、前記画像特徴量及び前記付随テキストデータに基づく損失、並びに、前記画像特徴量及び前記生成テキストデータに基づく損失に基づいて、前記テキスト生成確率パラメータを更新する、請求項4に記載の画像処理装置。
-
前記パラメータ更新部は、同じクラスの画像データに関する前記付随テキストデータのテキスト特徴量及び前記生成テキストデータのテキスト特徴量が近づくように前記テキスト生成確率パラメータを更新すると共に、異なるクラスの画像データに関する前記付随テキストデータのテキスト特徴量及び前記生成テキストデータのテキスト特徴量が遠ざかるように前記テキスト生成確率パラメータを更新する、請求項4に記載の画像処理装置。
-
前記テキスト生成部は、各時刻で上位の所定数のトークンからのランダムサンプリングによる前記生成テキストデータの生成と、通常時の前記生成テキストデータの生成とを併用する、請求項4に記載の画像処理装置。
-
前記テキスト理解部は、前記生成テキストデータに係る所定番目のトークンの確からしさを示す確率を用いた重み付きプーリングを行うことで、前記テキスト特徴量を抽出する、
請求項4に記載の画像処理装置。
-
請求項2乃至9のいずれか一項に記載の画像処理装置と、
通信ネットワークを介して前記画像処理装置に前記画像データを送信し、前記通信ネットワークを介して前記画像処理装置から前画像類似度に基づく画像の分類結果データを受信する通信端末と、
を有する通信システム。
-
画像データの特徴量を抽出する画像処理装置が実行する画像処理方法であって、
前記画像処理装置は、
前記画像データの画像パターンをベクトル化して画像特徴量を抽出する画像理解ステップと、
前記画像データに付随している付随テキストデータのテキストパターンをベクトル化してテキスト特徴量を抽出するテキスト理解ステップと、
前記画像理解ステップによって抽出された前記画像特徴量と前記テキスト理解ステップによって抽出された前記テキスト特徴量を同じベクトル空間に射影して、前記画像特徴量と前記テキスト特徴量を混合することで、前記特徴量としての混合特徴量を生成する特徴量混合ステップと、
を実行する画像処理方法。
-
コンピュータに、請求項11に記載の方法を実行させるプログラム。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2023559416A JP7647919B2 (ja) | 2021-11-12 | 2022-06-24 | 画像処理装置、画像処理方法、及びプログラム |
| US18/708,562 US20250005913A1 (en) | 2021-11-12 | 2022-06-24 | Image processing apparatus, image processing method, and program |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2021/041801 WO2023084759A1 (ja) | 2021-11-12 | 2021-11-12 | 画像処理装置、画像処理方法、及びプログラム |
| JPPCT/JP2021/041801 | 2021-11-12 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023084833A1 true WO2023084833A1 (ja) | 2023-05-19 |
Family
ID=86335445
Family Applications (2)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2021/041801 Ceased WO2023084759A1 (ja) | 2021-11-12 | 2021-11-12 | 画像処理装置、画像処理方法、及びプログラム |
| PCT/JP2022/025412 Ceased WO2023084833A1 (ja) | 2021-11-12 | 2022-06-24 | 画像処理装置、画像処理方法、及びプログラム |
Family Applications Before (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2021/041801 Ceased WO2023084759A1 (ja) | 2021-11-12 | 2021-11-12 | 画像処理装置、画像処理方法、及びプログラム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US20250005913A1 (ja) |
| JP (1) | JP7647919B2 (ja) |
| WO (2) | WO2023084759A1 (ja) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2025105245A1 (ja) * | 2023-11-15 | 2025-05-22 | 株式会社日立製作所 | 物体検出方法及び物体検出システム |
| WO2025257891A1 (ja) * | 2024-06-10 | 2025-12-18 | Ntt株式会社 | 音声理解装置、音声理解方法、及びプログラム |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN117710995A (zh) * | 2023-12-26 | 2024-03-15 | 上海合合信息科技股份有限公司 | 一种文档图像的多模态分类方法及装置 |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2020052463A (ja) * | 2018-09-21 | 2020-04-02 | 株式会社マクロミル | 情報処理方法および情報処理装置 |
| US20200311467A1 (en) * | 2019-03-29 | 2020-10-01 | Microsoft Technology Licensing, Llc | Generating multi modal image representation for an image |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP3500097B2 (ja) | 1999-08-26 | 2004-02-23 | 日本電信電話株式会社 | 複合メディア検索方法および複合メディア検索用プログラム記録媒体 |
| JP6553776B1 (ja) | 2018-07-02 | 2019-07-31 | 日本電信電話株式会社 | テキスト類似度算出装置、テキスト類似度算出方法、及びプログラム |
-
2021
- 2021-11-12 WO PCT/JP2021/041801 patent/WO2023084759A1/ja not_active Ceased
-
2022
- 2022-06-24 US US18/708,562 patent/US20250005913A1/en active Pending
- 2022-06-24 WO PCT/JP2022/025412 patent/WO2023084833A1/ja not_active Ceased
- 2022-06-24 JP JP2023559416A patent/JP7647919B2/ja active Active
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2020052463A (ja) * | 2018-09-21 | 2020-04-02 | 株式会社マクロミル | 情報処理方法および情報処理装置 |
| US20200311467A1 (en) * | 2019-03-29 | 2020-10-01 | Microsoft Technology Licensing, Llc | Generating multi modal image representation for an image |
Non-Patent Citations (1)
| Title |
|---|
| WATANABE YASUHIKO, NAGAO, MAKOTO: "Image Analysis Using Natural Language Information Extracted from Explanation Text.", JOURNAL OF THE JAPANESE SOCIETY FOR ARTIFICIAL INTELLIGENCE, vol. 13, no. 1, 1 January 1998 (1998-01-01), pages 66 - 74, XP093065527 * |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2025105245A1 (ja) * | 2023-11-15 | 2025-05-22 | 株式会社日立製作所 | 物体検出方法及び物体検出システム |
| WO2025257891A1 (ja) * | 2024-06-10 | 2025-12-18 | Ntt株式会社 | 音声理解装置、音声理解方法、及びプログラム |
Also Published As
| Publication number | Publication date |
|---|---|
| JPWO2023084833A1 (ja) | 2023-05-19 |
| US20250005913A1 (en) | 2025-01-02 |
| WO2023084759A1 (ja) | 2023-05-19 |
| JP7647919B2 (ja) | 2025-03-18 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN111291183B (zh) | 利用文本分类模型进行分类预测的方法及装置 | |
| Zhou et al. | Deep semantic dictionary learning for multi-label image classification | |
| JP7120433B2 (ja) | 回答生成装置、回答学習装置、回答生成方法、及び回答生成プログラム | |
| US11132512B2 (en) | Multi-perspective, multi-task neural network model for matching text to program code | |
| Vasilev et al. | Python deep learning | |
| CN108733792B (zh) | 一种实体关系抽取方法 | |
| Li et al. | Visualizing and understanding neural models in NLP | |
| Zhang et al. | Exploring region relationships implicitly: Image captioning with visual relationship attention | |
| US11900250B2 (en) | Deep learning model for learning program embeddings | |
| JP7647919B2 (ja) | 画像処理装置、画像処理方法、及びプログラム | |
| JP6772213B2 (ja) | 質問応答装置、質問応答方法及びプログラム | |
| CN114328931B (zh) | 题目批改方法、模型的训练方法、计算机设备及存储介质 | |
| JP2019533259A (ja) | 逐次正則化を用いた同時多タスクニューラルネットワークモデルのトレーニング | |
| CN110111864A (zh) | 一种基于关系模型的医学报告生成模型及其生成方法 | |
| CN112926655A (zh) | 一种图像内容理解与视觉问答vqa方法、存储介质和终端 | |
| US20250292021A1 (en) | Classification using a grammar-constrained generative language model | |
| CN114332565A (zh) | 一种基于分布估计的条件生成对抗网络文本生成图像方法 | |
| US12530377B2 (en) | Additional searching based on confidence in a classification performed by a generative language machine learning model | |
| Mustapha et al. | Convolution neural network and deep learning | |
| JP2021136025A (ja) | 学習装置、学習方法、学習プログラム、評価装置、評価方法、および評価プログラム | |
| Vasilev | Python Deep Learning: Understand how deep neural networks work and apply them to real-world tasks | |
| JP7815154B2 (ja) | プログラム、情報処理装置および情報処理方法 | |
| Bharadi | Random net implementation of MLP and LSTMs using averaging ensembles of deep learning models | |
| US12632471B2 (en) | Query clarification based on confidence in a classification performed by a generative language machine learning model | |
| US20260111792A1 (en) | Systems and methods for promoting diversity of machine learning training data sets through application of an embedding function |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 22892336 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2023559416 Country of ref document: JP |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18708562 Country of ref document: US |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 22892336 Country of ref document: EP Kind code of ref document: A1 |













