WO2019007041A1 - 基于多视图联合嵌入空间的图像-文本双向检索方法 - Google Patents
基于多视图联合嵌入空间的图像-文本双向检索方法 Download PDFInfo
- Publication number
- WO2019007041A1 WO2019007041A1 PCT/CN2018/074408 CN2018074408W WO2019007041A1 WO 2019007041 A1 WO2019007041 A1 WO 2019007041A1 CN 2018074408 W CN2018074408 W CN 2018074408W WO 2019007041 A1 WO2019007041 A1 WO 2019007041A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- picture
- sentence
- view
- phrase
- text
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/33—Querying
- G06F16/3331—Query processing
- G06F16/334—Query execution
- G06F16/3344—Query execution using natural language analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/205—Parsing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/33—Querying
- G06F16/338—Presentation of query results
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/36—Creation of semantic tools, e.g. ontology or thesauri
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/50—Information retrieval; Database structures therefor; File system structures therefor of still image data
- G06F16/58—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
- G06F16/583—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/211—Selection of the most significant subset of features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/213—Feature extraction, e.g. by transforming the feature space; Summarisation; Mappings, e.g. subspace methods
- G06F18/2132—Feature extraction, e.g. by transforming the feature space; Summarisation; Mappings, e.g. subspace methods based on discrimination criteria, e.g. discriminant analysis
- G06F18/21322—Rendering the within-class scatter matrix non-singular
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/213—Feature extraction, e.g. by transforming the feature space; Summarisation; Mappings, e.g. subspace methods
- G06F18/2135—Feature extraction, e.g. by transforming the feature space; Summarisation; Mappings, e.g. subspace methods based on approximation criteria, e.g. principal component analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/214—Generating training patterns; Bootstrap methods, e.g. bagging or boosting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/22—Matching criteria, e.g. proximity measures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/25—Fusion techniques
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/25—Fusion techniques
- G06F18/253—Fusion techniques of extracted features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/253—Grammatical analysis; Style critique
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/80—Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/20—Scenes; Scene-specific elements in augmented reality scenes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V30/00—Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
- G06V30/10—Character recognition
- G06V30/26—Techniques for post-processing, e.g. correcting the recognition result
- G06V30/262—Techniques for post-processing, e.g. correcting the recognition result using context analysis, e.g. lexical, syntactic or semantic context
- G06V30/274—Syntactic or semantic context, e.g. balancing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/213—Feature extraction, e.g. by transforming the feature space; Summarisation; Mappings, e.g. subspace methods
- G06F18/2132—Feature extraction, e.g. by transforming the feature space; Summarisation; Mappings, e.g. subspace methods based on discrimination criteria, e.g. discriminant analysis
- G06F18/21322—Rendering the within-class scatter matrix non-singular
- G06F18/21324—Rendering the within-class scatter matrix non-singular involving projections, e.g. Fisherface techniques
Definitions
- the invention relates to the field of computer vision, in particular to an image-text bidirectional retrieval method based on multi-view joint embedding space, and obtains a multi-view joint embedding space by learning to realize a bidirectional image-text retrieval task.
- the method uses different views to observe the data, obtains the semantic relationship of the data under different granularities, and fuses through the fusion sorting method to obtain more precise semantic associations, which makes the retrieval results more accurate.
- the current mainstream method is to transform the two kinds of data to obtain isomorphic features, thereby embedding the data into a common space, so that they can be directly compared.
- researchers have done a lot of work on what features are used to represent the data and which methods are used to embed the data in a common space.
- the present invention provides an image-text bidirectional retrieval method based on multi-view joint embedding space, which combines global and local level information to complete an image-text retrieval task, from a picture-sentence view. Observe the data under the area-phrase view, obtain the semantic relationship between the global level and the local level, and then obtain a precise semantic understanding by combining the two semantic associations.
- the principle of the present invention is that existing methods for acquiring and understanding image and text semantic information have been well represented in some aspects, but the existing method of observing semantic association from a single perspective cannot obtain two methods.
- the invention completes the image-text retrieval task by combining the information of the global level and the local level, that is, the semantic association obtained by the above two perspectives is obtained, and a more comprehensive and accurate semantic association relationship is obtained, which is beneficial to comprehensively understanding the semantic information of the data and Get accurate search results.
- the invention proposes that the multi-view joint spatial learning framework is divided into three parts: a picture-sentence embedding module, a region-phrase embedding module, and a multi-view fusion module.
- a complete sentence in a frame of picture or text is regarded as a basic unit, and two modal data (ie, image data and text data) are obtained by the existing pre-trained model.
- the characteristics of semantic information using a bifurcated neural network to learn the isomorphic features of two sets of features that enable image and text data to be mapped into a common space.
- each picture and each sentence are respectively divided into regions and phrases, and then the existing feature extraction methods are used to extract the features of the local data containing semantic information, and send them to another
- the isomorphic feature that maintains the semantic similarity and a local level subspace embedded with fine-grained data and can directly calculate the distance are also obtained, so as to explore the relationship between these fine-grained data.
- a two-branch neural network in both views, and one branch is used to process one modal data to make it isomorphic.
- the neural network of each branch consists of two layers of fully connected layers.
- the two-branch neural network structure in the two views is the same, but after training for different training data, the two neural networks can extract features for the picture-sentence data and the region-phrase data separately and retain the semantic association relationship under different views.
- the present invention uses the inner product between data vectors to calculate the distance, and thereby represents the similarity between the data, and during training, in order to preserve the semantic association information, a series of constraints are used to ensure that the semantically related data is This common space has a more adjacent positional relationship.
- the multi-view fusion module we calculate the distance between the image and the text data in the multi-view joint space by proportionally combining the distances calculated in the first two views. This final distance can more accurately show the semantic similarity between the data and can be used as a basis for sorting the search tasks. It should be noted that in the process of performing the retrieval task, the semantic association relationship of one view may be used for searching separately, or the semantic association relationship after multi-view fusion may be used for retrieval, and in the subsequent experimental results, multiple views may be explained. The semantic similarity obtained after fusion can more accurately reflect the semantic relationship between data.
- ⁇ , each document D i in the data set includes a picture I i and a related text T i , D i (I i , T i ), each piece of text consists of multiple sentences, each sentence is independent Describe the matching picture; in the picture-sentence view, set f i to represent a picture of the training image I i , ⁇ s i1 , s i2 ,..., s ik ⁇ represents the sentence in T i The set, k is the number of sentences in the text T i ; in the region-phrase view, r im is set to represent the mth region extracted by the picture f i , and p in represents the nth extracted from the sentence in the text T i
- step 2) The two sets of features (CNN features and FV features) obtained in step 1) are respectively sent into two branches of the bifurcated neural network, and the isomorphic features of the picture and sentence data are obtained through training, and the pictures and sentences are mapped at this time. Go to the global level subspace and obtain the semantic association information of the image and text data under the picture-sentence view;
- step 3) The two sets of features obtained in step 3) (dependencies of RCNN features and phrases) are respectively sent to the two branches of another bifurcated neural network, and trained to obtain isomorphic features of the region and phrase data. And phrases are mapped to the local level subspace to obtain semantic association information of the image and text data in the region-phrase view;
- the invention provides an image-text bidirectional retrieval method based on multi-view joint embedding space, and obtains a multi-view joint embedding space by learning to realize a bidirectional image-text retrieval task.
- the method uses different views to observe the data, obtains the semantic relationship of the data under different granularities, and fuses through the fusion sorting method to obtain more precise semantic associations, which makes the retrieval results more accurate.
- the present invention uses a bifurcated neural network to respectively map image and text data to a global level subspace of a picture-sentence view and a local level subspace of a region-phrase view, and obtain a semantic relationship between the global level and the local level, according to These two sets of semantic associations can accomplish the image-text two-way retrieval task alone, but the obtained retrieval results are not comprehensive.
- the invention proposes a multi-view fusion sorting method, which fuses the semantic association relations under the two views and jointly calculates the distance between the data in the multi-view joint space, and the distance relationship between the obtained data can more accurately reflect the data.
- the semantic similarity makes the retrieval result more accurate.
- the present invention has the following technical advantages:
- the present invention observes high-level semantic association relationships between different modal data from multiple views, and fuses these association relationships to form a semantic association relationship under multi-view, and the existing image-text retrieval method does not have this consideration.
- the method of the invention can learn the semantic information of the data at different granularities, thereby effectively extracting more accurate semantic information and obtaining higher-accuracy search results;
- the present invention uses the fusion sorting method to fuse the semantic associations of different views, so that the relative distance of the data in the multi-view joint space can well synthesize the semantic relations in different views, and finally obtain more precise semantics. Similarity
- the present invention adopts a bifurcated neural network, and its function is that the data of different modalities are heterogeneous, and it is impossible to directly compare or calculate the distance.
- Each branch processes a modal data through a bifurcated neural network. Deform the data of different modalities into isomorphic features, so that the heterogeneous data exists in a common space at the same time, and the distance can be directly calculated;
- the present invention adopts a series of constraints.
- the data is transformed by the dual-branch neural network, but in the conversion process, the original semantic relationship of the data needs to be retained; using the interval-based random loss function Separating the distance between the semantic related data and the distance between the semantically independent data to ensure that the semantic similarity information between the data in the common space is preserved.
- FIG. 1 is a flow chart of image-text bidirectional retrieval based on multi-view joint embedding space in the present invention.
- FIG. 2 is a schematic diagram of a multi-view joint embedding space learning process in an embodiment of the present invention
- VGG is a 19-layer VGG model, extracting the CNN features of the picture
- HGLMM is the Fisher Vector feature of the Hybrid Gaussian-Laplacian mixture model
- RCNN is the RCNN of the Faster RCNN model extraction region.
- Parser is a dependent triplet of the Standford CoreNLP parser to extract phrases.
- the neural network consists of two layers of fully connected layers, and the two neural networks in each view form a bifurcated neural network.
- Figure 3 is a schematic diagram of the consistency between modalities and the consistency within the modal
- (a) is the inter-modal consistency representation
- (b) is the intra-modal consistency representation
- FIG. 4 is an image-text retrieval result obtained by using the method of the present invention under the Pascal 1K data set according to an embodiment of the present invention.
- the invention provides an image-text bidirectional retrieval method based on multi-view joint embedding space, and combines the semantic association relationship between the global level and the local level to perform retrieval; firstly obtain global and local respectively from the picture-sentence view and the region-phrase view.
- the semantic association of the level, in the picture-sentence view, the picture and the sentence are embedded in the subspace of the global level to obtain the semantic association information of the picture and the sentence; in the area-phrase view, each area of the picture and the sentence are extracted.
- Each phrase is embedded in the subspace of the local level to obtain the semantic association information of the region and the phrase; in both views, the data is processed by the bifurcated neural network to obtain the isomorphic feature embedded in the common space, and the constraint is used to retain the data in the training.
- the original semantic relationship then through the multi-view fusion sorting method to fuse the two semantic associations to obtain more precise semantic similarity between the data.
- each document includes a picture and a related piece of text.
- D i (I i , T i )
- each text consists of several sentences, each of which independently describes the matching picture.
- f i to represent a picture of the training image I i and use ⁇ s i1 , s i2 ,..., s ik ⁇ to represent the set of sentences in T i (k is the text T i The number of sentences).
- all sentences can be divided into two sets, one set contains all the sentences matching the training picture, and one set contains all the sentences that do not match the training picture.
- the specific mathematical expression is as shown in equation (1):
- d(f i , s ix ) represents the distance between the picture f i and the sentence s ix in its matching set; d(f i , s jy ) indicates that the picture f i does not match The distance between the sentences s jy in the collection.
- d(f i , s ix ) represents the distance between the sentence s ix and the picture f i in its matching set; d(f j , s ix ) indicates that the sentence s ix does not match it The distance between the pictures f j in the set.
- d(s ix , s iy ) represents the distance between the sentences s ix and s iy describing the same picture f i ;
- d(s ix , s jz ) represents the sentence s ix describing the picture f i The distance from the sentence s jz describing the picture f j .
- each relationship R has a separate weight W R and an offset b R , and the number of phrases extracted for each sentence is different.
- ⁇ region-phrase ⁇ i,j,x,y ⁇ ij max[0,1- ⁇ ij ⁇ d(r ix ,p jy )] (6)
- d(r ix , p jy ) represents the distance between the region r ix and the phrase p jy .
- the constant ⁇ ij is used to normalize according to the number of positive and negative ⁇ ij .
- d multi-view (I i , T j ) d frame-sentence (I i ,T j )+ ⁇ d region-phrase (I i ,T j ) (7)
- d frame-sentence (I i , T j ) represents the distance between the image I i and the text T j in the picture-sentence view
- d region-phrase (I i , T j ) represents the image I
- d multi-view (I i , T j ) represents the distance between the image I i and the text T j after multi-view fusion.
- Figure 1 is a flow chart of a multi-view joint embedded spatial learning framework.
- Figure 2 shows a schematic diagram of the learning framework proposed by the present invention.
- the framework is divided into three parts.
- In the picture-sentence view we embed the pictures and sentences into the subspace of the global level.
- In the region-phrase view we extract small components from pictures and sentences and embed them into subspaces at the local level.
- In each view we use a two-branch neural network to process the data so that they become isomorphic features embedded in the common space.
- a multi-view fusion sorting method is proposed to fuse the semantic correlation information obtained by two view analysis to obtain the final distance relationship between the data.
- Figure 3 shows a schematic of modal consistency (left) and modal consistency (right) in the picture-sentence view.
- the square represents the picture
- the circle represents the sentence
- the image of the same color represents the same semantic information.
- Inter-modal consistency means that in the picture-sentence view, the distance between a picture (black square) and the sentence it matches (black circle) must be greater than the picture (black square) and the sentence that does not match it (gray The distance between the circles is small, and this constraint applies equally to sentences.
- Intramodal consistency means that in a picture-sentence view, the distance between a sentence (black square) and a sentence (other black squares) that is semantically similar must be greater than the sentence (black square) and the sentence unrelated to its semantics ( The distance between the gray squares is small.
- Figure 4 shows an actual case of image-text retrieval, which gives the results of the picture-sentence view, the region-phrase view and the multi-view in the first five sentences returned for the upper left image, and the correct search results are added. Thick representation.
- the picture-sentence view can only retrieve the sentence after the global level is understood, but since it can not distinguish the content in the area, it will have the correct matching sentence and some have similar meanings but include Incorrect individual sentences are confused.
- the region-phrase view it returns some sentences that contain the correct individuals but the relationships between the individuals are inaccurate.
- the third sentence in this view which can recognize ‘a young girl’, but misunderstands that the relationship between the girl and the bicycle is ‘riding’, and finally returns a wrong sentence.
- the multi-view after fusion it can capture the semantic relationship of the global level and the semantic relationship of the local level at the same time, so the multi-view retrieval effect is the most accurate.
- Tables 1 and 2 show the results of experimental verification of the present invention on image-text retrieval and text-image retrieval on Pascal1K and Flickr8K, respectively.
- the figure shows the comparison of the effects of the present invention with other existing advanced algorithms, including SDT-RNN (Semantic Dependency Trees-Recurrent Neural Networks), kCCA (kernel Canonical Correlation Analysis).
- DeViSE Deep Visual-Semantic Embedding
- DCCA Deep Canonical Correlation Analysis
- VQA-A VehicleQuestion Answering-agnostic
- DFE Deep Fragment Embedding
- the method of the present invention is more effective than other comparative methods.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Data Mining & Analysis (AREA)
- General Engineering & Computer Science (AREA)
- Artificial Intelligence (AREA)
- Evolutionary Computation (AREA)
- Computational Linguistics (AREA)
- General Health & Medical Sciences (AREA)
- Health & Medical Sciences (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Life Sciences & Earth Sciences (AREA)
- Software Systems (AREA)
- Computing Systems (AREA)
- Databases & Information Systems (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Mathematical Physics (AREA)
- Biomedical Technology (AREA)
- Biophysics (AREA)
- Molecular Biology (AREA)
- Multimedia (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Library & Information Science (AREA)
- Medical Informatics (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
本发明公布了一种基于多视图联合嵌入空间的图像-文本双向检索方法,通过结合全局层面和局部层面的语义关联关系进行检索;先从画面-句子视图和区域-短语视图下分别获得全局和局部层面的语义关联关系,在画面-句子视图中获取画面和句子全局层面子空间中的语义关联信息;在区域-短语视图中获取区域和短语局部层面子空间中的语义关联信息;两个视图中均通过双分支的神经网络处理数据得到同构特征嵌入共同空间,在训练中使用约束条件保留数据原有的语义关系;再通过多视图融合排序方法融合两种语义关联关系得到数据之间更精准的语义相似度,使得检索结果准确度更高。
Description
本发明涉及计算机视觉领域,尤其涉及一种基于多视图联合嵌入空间的图像-文本双向检索方法,通过学习得到多视图联合嵌入空间,实现双向图像-文本检索任务。该方法利用不同视图观察数据,获得数据在不同粒度下的语义关联关系,并通过融合排序方法进行融合,获得更精确的语义关联关系,使得检索结果更加准确。
随着计算机视觉领域研究的不断发展,出现了大批与图像-文本相关的任务,如图片说明(imagecaption),图片描述(densecaption),以及视觉问题回答(visualquestionanswering)等。所有这些任务都要求计算机充分理解图像和文本的语义信息,并通过学习能够将一种模态数据的语义信息翻译到另一个模态,因此,该类任务的核心问题是如何弥合两种不同模态在语义层面上的差距,找到沟通两种模态数据的方法。不同模态的数据存在于异构的空间中,因而想要直接通过计算数据之间的距离来度量它们的语义相似度是不现实的。为了解决这一问题,目前的主流方法是对两种数据进行变形得到同构的特征,从而将数据嵌入到一个共同的空间,这样它们可以直接被比较。针对使用何种特征表示数据和使用何种方法将数据嵌入到共同空间中,研究者们做了大量的工作。
考虑到一些由深度学习得到的特征在许多计算机视觉领域的任务上取得了很好的成绩,大量的研究者使用这些特征来表示图像或文本,并将它们转化成同构形式从而可以将数据映射到共同空间来完成图像-文本检索任务。然而,这些学习得到的特征仅仅提供了数据全局层面的信息而缺少局部信息的描述,因此只用这些特征来表示数据是无法挖掘更细粒度的数据之间的关联性的,比如图像的区域和文本中的短语之间的关联关系。另一类方法是将图像和文本都切分成小的部分,将它们投影到一个共同的空间,以便于从细粒度的数据中捕捉到局部语义关联信息。
尽管上述方法在某些方面已经取得了很好的表现,但仅仅从单一的视角,即局部视角或全局视角,观察语义关联是无法获得两种模态数据之间全面完整的关联关系的。亦即,同时获得上述两种视角观察得到的语义关联并加以合理利用有利于综合理解数据的语义信息并获得精确的检索结果。但目前尚无能同时获得不同视角的异构数据关联关系,并对这些关系加以融合得到数据之间最终的语义相似度的方法。
发明内容
为了克服上述现有技术的不足,本发明提供一种基于多视图联合嵌入空间的图像-文本双向检索方法,通过结合全局层面和局部层面的信息来完成图像-文本检索任务,从画面-句子视图和区域-短语视图下观察数据,获得全局层面和局部层面的语义关联关系,然后通过融合这两种语义关联关系得到一个精准的语义理解。
本发明的原理是:现有的图像和文本语义信息的获取和理解的方法在某些方面已经取得了很好的表现,但是,现有仅仅从单一的视角观察语义关联的方法是无法获得两种模态数据之间全面完整的关联关系的。本发明通过结合全局层面和局部层面的信息来完成图像-文本检索任务,即融合上述两种视角观察得到的语义关联,获得更全面准确的语义关联关系,将有利于综合理解数据的语义信息并获得精确的检索结果。本发明提出基于多视图联合空间学习框架分成三个部分:画面-句子嵌入模块、区域-短语嵌入模块,以及多视图融合模块。在画面-句子视图嵌入模块中,将一帧图片或文本中的一个完整句子视为基本单位,通过已有的预先训练好的模型获得两种模态数据(即图像数据和文本数据)富含语义信息的特征,利用双分支神经网络学习两组特征的同构特征,这些特征向量使得图像和文本数据能够被映射到共同的空间中。与此同时,在对双分支神经网络进行训练的过程中,需要保持数据之间原有的关联关系,即语义相似度,之后尝试将这种关联关系在全局层面的子空间中通过可计算得到的距离保存下来。在区域-短语视图嵌入模块中,将每一幅画面和每一个句子分别切分成区域和短语,然后使用已有的特征提取方法提取这些局部数据包含语义信息的特征,将它们送入到另一个双分支神经网络中,同样获得保持了语义相似度的同构特征和一个嵌入了细粒度数据并可以直接计算距离的局部层面子空间,以便于探索这些细粒度数据之间的关联关系。
在这两个模块中,为了将异构数据嵌入到共同空间中,两个视图中均采用双分支的神经网络,分别用一个分支处理一个模态的数据,使其变成同构的可进行比较的特征。每一分支的神经网络分别由两层全连接层构成。两个视图中的双分支神经网络结构相同,但是针对不同训练数据训练之后,两个神经网络可以分别针对画面-句子数据,和区域-短语数据提取特征并保留不同视图下的语义关联关系。本发明使用数据向量之间的内积来计算距离,并以此表示数据之间的相似度,并且在训练期间,为了保存语义关联信息,使用了一系列的约束条件来保证语义相关的数据在这一共同空间中有着更邻近的位置关系。
在多视图融合模块中,我们通过按比例结合前两个视图中计算得到的距离来计算图像和文本数据之间在多视图联合空间中的距离。这一最终的距离能够更精确的显示出数据之间的 语义相似度,并且可以作为检索任务的排序依据。需要注意的是,在执行检索任务的过程中,可以单独利用一个视图的语义关联关系进行检索,也可以使用多视图融合后的语义关联关系进行检索,在后续的实验结果中,可以说明多视图融合后得到的语义相似度更能准确的反应数据之间的语义关联关系。
本发明提供的技术方案是:
一种基于多视图联合嵌入空间的图像-文本双向检索方法,通过结合全局层面和局部层面的信息进行图像-文本双向检索;针对数据集D={D
1,D
2,...,D
|D|},数据集中每一个文档D
i包括一张图片I
i和一段相关的文本T
i,D
i=(I
i,T
i),每一段文本由多个句子组成,每一个句子都独立地对相匹配的图片进行描述;在画面-句子视图中,设定f
i代表训练图像I
i的一幅画面,{s
i1,s
i2,...,s
ik}代表T
i中的句子集合,k是文本T
i中句子的个数;在区域-短语视图中,设定r
im代表画面f
i提取出的第m个区域,p
in代表文本T
i中的句子提取出的第n个短语;本发明方法首先从画面-句子视图和区域-短语视图下观察数据,分别获得全局层面和局部层面的语义关联关系,然后通过融合这两种语义关联关系得到一个精准的语义理解;具体包括如下步骤:
1)分别提取所有的图像的画面和所有文本中的句子,使用已有的19层VGG(Visual Geometry Group提出的神经网络结构)模型提取画面数据的CNN(Convolutional Neural Network)特征,使用已有的混合高斯-拉普拉斯混合模型(HGLMM)提取句子数据的FV(Fisher vector)特征;
2)将步骤1)得到的两组特征(CNN特征和FV特征)分别送入双分支神经网络的两个分支中,经过训练得到画面和句子数据的同构特征,此时画面和句子被映射到全局层面子空间,并获得画面-句子视图下图像和文本数据的语义关联信息;
3)使用已有的Faster RCNN模型(Faster Region-based Convolutional Network)提取所有画面的区域RCNN特征,使用已有的Standford CoreNLP语法分析器提取所有句子的短语依赖关系(dependency triplet),保留含有关键信息的区域和短语特征;
4)将步骤3)得到的两组特征(RCNN特征和短语的依赖关系)分别送入另一个双分支神经网络的两个分支中,经过训练得到区域和短语数据的同构特征,此时区域和短语被映射到局部层面子空间,得到区域-短语视图下图像和文本数据的语义关联信息;
5)使用融合排序方法,将步骤2)和步骤4)得到的不同视图下图像和文本数据的语义关联信息融合起来,计算得到多视图下图像和文本数据在多视图联合空间内的距离,该距离用于度量语义相似度,在检索过程中将其作为排序标准;
6)对检索请求计算得到该检索请求数据在多视图联合空间中与数据集D中另一模态数据(图像或文本)之间的距离(即多视图下图像和文本数据在多视图联合空间内的距离),根据距离对检索结果进行排序;
由此实现基于多视图联合嵌入空间的图像-文本双向检索。
与现有技术相比,本发明的有益效果是:
本发明提供一种基于多视图联合嵌入空间的图像-文本双向检索方法,通过学习得到多视图联合嵌入空间,实现双向图像-文本检索任务。该方法利用不同视图观察数据,获得数据在不同粒度下的语义关联关系,并通过融合排序方法进行融合,获得更精确的语义关联关系,使得检索结果更加准确。具体地,本发明使用双分支神经网络分别将图像和文本数据映射到画面-句子视图的全局层面子空间和区域-短语视图的局部层面子空间,获得全局层面和局部层面的语义关联关系,根据这两组语义关联关系可以单独完成图像-文本双向检索任务,但所得到的检索结果并不全面。本发明提出多视图融合排序方法,将两个视图下的语义关联关系融合在一起,共同计算数据之间在多视图联合空间中的距离,得到的数据之间的距离关系能够更精准的反应数据的语义相似度,使得检索结果的准确度更高。具体地,本发明具有如下技术优势:
(一)本发明从多个视图观察不同模态数据之间的高层语义关联关系,并将这些关联关系融合在一起形成多视图下的语义关联关系,而现有图像-文本检索方法没有此考虑;采用本发明方法可以学习到数据在不同粒度下存在的语义信息,从而能有效的提取更精确的语义信息,获得准确度更高的检索结果;
(二)本发明使用融合排序方法能够将不同视图的语义关联关系融合在一起,使得数据在多视图联合空间中的相对距离能够很好地综合不同视图下的语义关系,最终获得更精确的语义相似度;
(三)本发明采用双分支神经网络,其作用在于不同模态的数据其特征是异构的,无法直接进行比较或计算距离,通过双分支神经网络,每一分支处理一个模态的数据,将不同模态的数据变形成为同构特征,使异构数据同时存在于一个共同空间,可以直接计算距离;
(四)本发明采用一系列约束条件,为了获得同构特征,利用双分支神经网络对数据进行转换,但是在转换过程中,数据原有的语义关系需要得到保留;利用基于间隔的随机损失函数将语义相关数据之间的距离与语义无关数据之间的距离范围拉开间隔,确保在共同空间 中数据之间的语义相似度信息得到保留。
图1是本发明中基于多视图联合嵌入空间进行图像-文本双向检索的流程框图。
图2是本发明实施例中多视图联合嵌入空间学习过程的示意图;
其中,VGG是19层VGG模型,提取画面的CNN特征,HGLMM是混合高斯-拉普拉斯混合模型(Hybrid Gaussian-Laplacian mixture model)提取句子的Fisher Vector特征,RCNN是Faster RCNN模型提取区域的RCNN特征,Parser是Standford CoreNLP语法分析器提取短语的依赖三元组。神经网络均由两层全连接层组成,每个视图下的两个神经网络组成双分支神经网络。
图3模态间一致性与模态内一致性示意图;
其中,(a)为模态间一致性表示,(b)为模态内一致性表示。
图4是本发明实施例提供的采用本发明方法在Pascal1K数据集下得到的图像-文本检索结果。
下面结合附图,通过实施例进一步描述本发明,但不以任何方式限制本发明的范围。
本发明提供一种基于多视图联合嵌入空间的图像-文本双向检索方法,通过结合全局层面和局部层面的语义关联关系进行检索;先从画面-句子视图和区域-短语视图下分别获得全局和局部层面的语义关联关系,在画面-句子视图中,将画面和句子嵌入到全局层面的子空间中获取画面和句子的语义关联信息;在区域-短语视图中,提取画面的每一个区域和句子的每一个短语,嵌入到局部层面的子空间中获取区域和短语的语义关联信息;两个视图中均通过双分支的神经网络处理数据得到同构特征嵌入共同空间,在训练中使用约束条件保留数据原有的语义关系;再通过多视图融合排序方法融合两种语义关联关系得到数据之间更精准的语义相似度。
我们利用数据集D={D
1,D
2,...,D
|D|}对图像-文本检索问题加以描述,在这一数据集中,每一个文档包括一张图片和一段相关的文本,如D
i=(I
i,T
i),每一个文本由若干个句子组成,每一个句子都独立地对相匹配的图片进行描述。在画面-句子视图中,我们用f
i代表训练图像I
i的一幅画面,用{s
i1,s
i2,...,s
ik}代表T
i中的句子集合(k是文本T
i中句子的个数)。在区域-短语视图中,我们用r
im代表画面f
i提取出的第m个区域,用p
in代表文本T
i中的句子提取出的第n 个短语。接下来我们将详细描述两个视图中使用的约束条件,以及最后多视图融合模块的融合排序方法。
1、画面-句子视图
我们将画面和句子分别送入双分支神经网络中,并且获得在全局层面子空间的同构特征。在训练神经网络的过程中,为了保留模态间一致性和模态内一致性,我们提出了基于间隔的随机损失函数。为了将数据放入画面-句子视图中进行处理,我们做了如下特征提取:
对于图像,我们使用19层VGG模型提取得到的4096维CNN特征(向量)作为画面的原始特征;对于文本,我们使用混合高斯-拉普拉斯混合模型(Hybrid Gaussian-Laplacian mixture model,HGLMM)提取得到的Fisher vector(FV)特征(向量)作为句子的原始特征,为了计算方便,我们用PCA(Principal Components Analysis,主成分分析)将最初的18000维FV特征向量降维至4999维。
1)模态间一致性
对于训练画面f
i,所有的句子可以分成两个集合,一个集合包含了所有与训练画面匹配的句子,一个集合包含了所有与训练画面不匹配的句子。我们可以推测出一个合理的一致性要求,即在画面-句子视图中,画面f
i和在匹配集合中的句子之间的距离必须比画面f
i和在不匹配集合中的句子之间的距离小,且距离的差距需要大于间隔m。具体数学表示如式(1):
d(f
i,s
ix)+m<d(f
i,s
jy)if i≠j (1)
式(1)中,d(f
i,s
ix)表示画面f
i与在其匹配集合中的句子s
ix之间的距离;d(f
i,s
jy)表示画面f
i与在其不匹配集合中的句子s
jy之间的距离。
类似的约束可以应用在训练句子s
ix上:
d(f
i,s
ix)+m<d(f
j,s
ix)if i≠j (2)
式(2)中,d(f
i,s
ix)表示句子s
ix与在其匹配集合中的画面f
i之间的距离;d(f
j,s
ix)表示句子s
ix与在其不匹配集合中的画面f
j之间的距离。
2)模态内一致性
在训练过程中,除了考虑到数据的模态间一致性,我们还需要针对数据集中同一画面伴随的若干个句子做一些约束,我们称之为模态内一致性。具体来说,对于共享相同含义即描述相同画面的句子而言,他们需要紧密的联系在一起,并且能够与其他的句子区分开来。
为了实现模态内一致性,我们使用了如下约束:
d(s
ix,s
iy)+m<d(s
ix,s
jz)if i≠j (3)
式(3)中,d(s
ix,s
iy)表示描述同一个画面f
i的句子s
ix与s
iy之间的距离;d(s
ix,s
jz)表示描述画面f
i的句子s
ix与描述画面f
j的句子s
jz之间的距离。
尽管我们应该对画面作出类似公式(3)的约束,即描述同一个句子的画面之间的距离应该更近,但目前我们使用的数据集中,还难以确定若干画面是否描述了相同的句子,因此我们不采用这一项约束,而只对描述同一个画面的句子之间的距离进行公式(3)的约束。
结合上述的约束条件,我们最终总结出了在画面-句子视图上的损失函数:
这里,间隔m可以根据所采用的距离作出调整,为了便于优化,在这里我们将其固定为m=0.1并将它应用在所有的训练样本中,与此同时,通过实验我们发现当λ
1=2且λ
2=0.2时我们能够获得最好的实验结果。
2、区域-短语视图
在这一视图中,我们希望能够挖掘存在于区域和短语之间的细粒度语义关联关系。我们通过已有的模型提取得到区域和短语特征,对于区域,我们提取画面中得分最高的19个区域的4096维RCNN特征;对于短语,我们利用语法分析器得到依赖树结构,并选择包含关键语义信息的短语。我们用1-of-k编码向量w来代表每一个单词,并将p
jy用依赖关系三元组(R,w
1,w
2)表示的短语映射到嵌入空间中,如式(5):
式(5)中,W
e是一个400000×d的矩阵,用于将1-of-k向量编码成一个d维的单词向量,其中400000是字典的单词数目,这里我们设定d=200。注意每一个关系R都有单独的权重W
R和偏移量b
R,并且每一个句子提取的短语数是不一样的。
在这个视图中,我们利用双分支神经网络将图像和文本数据射到区域-短语视图的局部子空间中。在训练网络的过程中,我们要求在匹配的图像文本对中的区域和短语之间的距离要比在不匹配对中的区域和短语之间的距离小。在计算这一视图下的数据映射到局部层面子空间的损失函数表示如下:
ψ
region-phrase=∑
i,j,x,yκ
ijmax[0,1-η
ij×d(r
ix,p
jy)] (6)
式(6)中,d(r
ix,p
jy)表示区域r
ix和短语p
jy之间的距离,我们定义η
ij在i=j的时候等于+1,在i≠j时等于-1,常量κ
ij用于根据η
ij的正负个数进行归一化。
3、多视图融合模块
在画面-句子嵌入模块和区域-短语嵌入模块分别通过学习得到各自的嵌入空间之后,我们可以借助这两个空间中的信息来获得多视图下的数据间距离。为了获得图像I
i和文本T
j之间更精确的语义相似度,我们按比例结合前两个视图下计算得到的距离作为多视图联合空间中两个数据之间的最后距离:
d
multi-view(I
i,T
j)=d
frame-sentence(I
i,T
j)+λd
region-phrase(I
i,T
j) (7)
式(7)中,d
frame-sentence(I
i,T
j)表示图像I
i和文本T
j之间在画面-句子视图中的距离,d
region-phrase(I
i,T
j)表示图像I
i和文本T
j之间在区域-短语视图中的距离,d
multi-view(I
i,T
j)表示图像I
i和文本T
j之间在多视图融合后的距离。权重λ用于平衡画面-句子视图和区域-短语视图距离的比例,经过实验,我们发现λ=0.6能够产生很好的效果。
图1为多视图联合嵌入空间学习框架流程图。
图2展示了本发明提出的学习框架示意图。该框架分三个部分,画面-句子视图中,我们将画面和句子嵌入到全局层面的子空间中。在区域-短语视图中,我们提取画面和句子中小的成分,将这些成分嵌入到局部层面的子空间中。每一个视图中,我们都用双分支的神经网络处理数据以使它们变成同构特征嵌入共同空间。在多视图融合模块,提出了多视图融合排序方法来融合两个视图分析得到的语义关联信息得到数据之间最终的距离关系。
图3展示了在画面-句子视图下模态间一致性(左)和模态内一致性(右)的示意图。正方形代表画面,圆形代表句子,相同颜色的图像表示的是同一个语义信息。模态间一致性含义为在画面-句子视图中,一个画面(黑色正方形)和与其匹配的句子(黑色圆形)之间的距离必须比该画面(黑色正方形)和与其不匹配的句子(灰色圆形)之间的距离小,此约束对句子同样适用。模态内一致性含义为在画面-句子视图中,一个句子(黑色正方形)和与其语 义相近的句子(其他黑色正方形)之间的距离必须比该句子(黑色正方形)和与其语义无关的句子(灰色正方形)之间的距离小。
图4展示了一个图像-文本检索的实际案例,分别给出了画面-句子视图,区域-短语视图和多视图在针对左上角图像返回的前五个句子的检索结果,正确的检索结果用加粗表示。在这个例子中,我们可以看出,画面-句子视图仅仅能检索到全局层面理解后的句子,但是由于它不能辨别区域中的内容,所以会把正确匹配的句子和一些有相似含义但并包括不正确个体的句子混淆。对于区域-短语视图,它返回了一些包含有正确的个体但个体之间关系不准确的句子。比如这一视图下的第三个句子,它可以辨认出‘a young girl’,但是却误解了女孩和自行车之间的关系是‘riding’,最终返回了一个错误的句子。但是在融合后的多视图中,它可以同时捕捉到全局层面的语义关联关系和局部层面的语义关联关系,因而多视图的检索效果是准确度最高的。
表1.Pascal1K数据集下实施例中图像-文本双向检索结果
表2.Flickr8K数据集下实施例中图像-文本双向检索结果
表1和表2分别给出了本发明在Pascal1K和Flickr8K上通过图像-文本检索和文本-图像检索进行实验验证的结果。为了评价检索效果,我们遵循了标准的排序度量标准,使用Recall@K,即正确匹配的数据排在前K(K=1,5,10)个检索结果中的概率,来对检索准确性进行度量。图中列出了本发明与其他现有先进算法的效果比较,包括SDT-RNN(Semantic Dependency Trees-Recurrent Neural Networks,语义依赖树-循环神经网络),kCCA(kernel Canonical Correlation Analysis,核典型相关分析),DeViSE(Deep Visual-Semantic Embedding,深度视觉语义嵌入),DCCA(Deep Canonical Correlation Analysis,深度典型相关分析),VQA-A(VisualQuestion Answering-agnostic,视觉问题回答不可知论),DFE(Deep Fragment Embedding,深度片段嵌入)。
从表1和表2我们可以看出,本发明方法和其他对比方法相比较而言效果更好。此外,我们还分别展示了两种单独视图的检索效果和多视图融合后的检索效果,从数据中可以看出,结合两个视图之后的多视图融合方法检索效果更好。这一结果证明了单独的两个视图之间彼此是互补的关系,所以将二者结合起来之后可以获得更精准全面的语义关联,检索效果会更好。
需要注意的是,公布实施例的目的在于帮助进一步理解本发明,但是本领域的技术人员可以理解:在不脱离本发明及所附权利要求的精神和范围内,各种替换和修改都是可能的。因此,本发明不应局限于实施例所公开的内容,本发明要求保护的范围以权利要求书界定的范围为准。
Claims (8)
- 一种基于多视图联合嵌入空间的图像-文本双向检索方法,通过结合全局层面和局部层面的语义关联关系进行图像-文本双向检索;针对数据集D={D 1,D 2,…,D |D|},数据集中每一个文档D i包括一张图片I i和一段相关的文本T i,表示为D i=(I i,T i),每一段文本由多个句子组成,每一个句子均独立描述相匹配的图片;在画面-句子视图中,设定f i代表训练图像I i的一幅画面,{s i1,s i2,…,s ik}代表T i中的句子集合,k是文本T i中句子的个数;在区域-短语视图中,设定r im代表画面f i提取出的第m个区域,p in代表文本T i中的句子提取出的第n个短语;所述双向检索方法首先从画面-句子视图和区域-短语视图下分别获得全局层面和局部层面的语义关联关系,然后通过融合两种语义关联关系得到一个精准的语义理解;具体包括如下步骤:1)分别提取所有图像的画面和所有文本中的句子,分别送入模型中提取数据的特征,得到画面的CNN特征和句子的FV特征;2)将步骤1)得到的画面的CNN特征和句子的FV特征分别送入双分支神经网络的两个分支中,经过训练得到画面和句子数据的同构特征,此时画面和句子被映射到全局层面子空间,并获得画面-句子视图下图像和文本数据的语义关联信息;3)使用RCNN模型和语法分析器分别提取所有画面的区域RCNN特征和所有句子的短语依赖关系,保留含有关键信息的区域和短语的特征;4)将步骤3)得到的区域和短语的特征分别送入另一个双分支神经网络的两个分支中,经过训练得到区域和短语数据的同构特征,此时区域和短语被映射到局部层面子空间,得到区域-短语视图下图像和文本数据的语义关联信息;5)使用融合排序方法,将步骤2)和步骤4)得到的不同视图下图像和文本数据的语义关联信息进行融合,计算得到多视图下图像和文本数据在多视图联合空间内的距离,该距离用于度量语义相似度,在检索过程中作为排序标准;6)对检索请求计算得到该检索请求数据在多视图联合空间中与数据集中另一模态数据之间的距离,根据距离对检索结果进行排序;由此实现基于多视图联合嵌入空间的图像-文本双向检索。
- 如权利要求1所述图像-文本双向检索方法,其特征是,步骤1)提取特征具体是:对于图像,使用19层VGG模型提取得到4096维CNN特征向量,作为画面的原始特征;对于文本,使用混合高斯-拉普拉斯混合模型提取得到FV特征向量,作为句子的原始特征;并通 过PCA将特征向量由18000维降维至4999维。
- 如权利要求1所述图像-文本双向检索方法,其特征是,步骤2)将特征分别送入双分支神经网络的两个分支中进行训练,得到画面和句子数据的同构特征,训练过程中设定约束条件以保留模态间一致性和模态内一致性,采用基于间隔的随机损失函数;具体包括:A.训练画面f i:将所有句子分成匹配集合和不匹配集合,匹配集合包含所有与训练画面匹配的句子,不匹配集合包含所有与训练画面不匹配的句子;设定一致性约束要求为:在画面-句子视图中,画面f i和在匹配集合中的句子之间的距离必须比画面f i和在不匹配集合中的句子之间的距离小,且距离的差距需要大于间隔m,表示如式(1):d(f i,s ix)+m<d(f i,s jy)if i≠j (1)式(1)中,d(f i,s ix)表示画面f i与在其匹配集合中的句子s ix之间的距离;d(f i,s jy)表示画面f i与在其不匹配集合中的句子s jy之间的距离;B.将式(2)的约束应用在训练句子s ix上:d(f i,s ix)+m<d(f j,s ix)if i≠j (2)式(2)中,d(f i,s ix)表示句子s ix与在其匹配集合中的画面f i之间的距离;d(f j,s ix)表示句子s ix与在其不匹配集合中的画面f j之间的距离;C.针对数据集中同一画面伴随的多个句子设定约束条件,表示为式(3):d(s ix,s iy)+m<d(s ix,s jz)if i≠j (3)式(3)中,d(s ix,s iy)表示描述同一个画面f i的句子s ix与s iy之间的距离;d(s ix,s jz)表示描述画面f i的句子s ix与描述画面f j的句子s jz之间的距离;D.建立在画面-句子视图上的损失函数如式(4):其中,m为间隔,可根据所采用的距离作出调整。
- 如权利要求3所述图像-文本双向检索方法,其特征是,间隔m取值为0.1;参数λ 1取值为2;参数λ 2取值为0.2。
- 如权利要求1所述图像-文本双向检索方法,其特征是,步骤4)将区域和短语的特征分别送入双分支神经网络的两个分支中进行训练,得到区域和短语数据的同构特征;具体地,在训练过程中,设定条件为:在匹配的图像文本对中的区域和短语之间的距离比在不匹配对中的区域和短语之间的距离小;通过式(6)的损失函数计算区域-短语视图下的数据并映射到局部层面子空间:ψ region-phrase=Σ i,j,x,yκ ijmax[0,1-η ij×d(r ix,p jy)] (6)式(6)中,d(r ix,p jy)表示区域r ix和短语p jy之间的距离;设定η ij在i=j的时候等于+1,在i≠j时等于-1,常量κ ij根据η ij的正负个数进行归一化。
- 如权利要求1所述图像-文本双向检索方法,其特征是,步骤5)通过式(7)按比例结合两个视图下计算得到的距离,作为多视图联合空间中两个数据之间的距离:d multi-view(I i,T j)=d frame-sentence(I i,T j)+λd region-phrase(I i,T j) (7)式(7)中,d frame-sentence(I i,T j)表示图像I i和文本T j之间在画面-句子视图中的距离;d region-phrase(I i,T j)表示图像I i和文本T j之间在区域-短语视图中的距离;d multi-view(I i,T j)表示图像I i和文本T j之间在多视图融合后的距离;权重λ用于平衡画面-句子视图和区域-短语视图距离的比例。
- 如权利要求7所述图像-文本双向检索方法,其特征是,权重λ取值为0.6。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US16/622,570 US11106951B2 (en) | 2017-07-06 | 2018-01-29 | Method of bidirectional image-text retrieval based on multi-view joint embedding space |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201710545632.X | 2017-07-06 | ||
| CN201710545632.XA CN107330100B (zh) | 2017-07-06 | 2017-07-06 | 基于多视图联合嵌入空间的图像-文本双向检索方法 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2019007041A1 true WO2019007041A1 (zh) | 2019-01-10 |
Family
ID=60195963
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2018/074408 Ceased WO2019007041A1 (zh) | 2017-07-06 | 2018-01-29 | 基于多视图联合嵌入空间的图像-文本双向检索方法 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US11106951B2 (zh) |
| CN (1) | CN107330100B (zh) |
| WO (1) | WO2019007041A1 (zh) |
Cited By (15)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109857895A (zh) * | 2019-01-25 | 2019-06-07 | 清华大学 | 基于多环路视图卷积神经网络的立体视觉检索方法与系统 |
| CN110197521A (zh) * | 2019-05-21 | 2019-09-03 | 复旦大学 | 基于语义结构表示的视觉文本嵌入方法 |
| CN110298395A (zh) * | 2019-06-18 | 2019-10-01 | 天津大学 | 一种基于三模态对抗网络的图文匹配方法 |
| CN110489551A (zh) * | 2019-07-16 | 2019-11-22 | 哈尔滨工程大学 | 一种基于写作习惯的作者识别方法 |
| CN111814658A (zh) * | 2020-07-07 | 2020-10-23 | 西安电子科技大学 | 基于语义的场景语义结构图检索方法 |
| CN112860847A (zh) * | 2021-01-19 | 2021-05-28 | 中国科学院自动化研究所 | 视频问答的交互方法及系统 |
| CN115033670A (zh) * | 2022-06-02 | 2022-09-09 | 西安电子科技大学 | 多粒度特征融合的跨模态图文检索方法 |
| CN116484878A (zh) * | 2023-06-21 | 2023-07-25 | 国网智能电网研究院有限公司 | 电力异质数据的语义关联方法、装置、设备及存储介质 |
| CN116842212A (zh) * | 2023-05-26 | 2023-10-03 | 西安电子科技大学 | 基于边界框提取和语义一致性约束的文本-行人检索方法 |
| CN117009569A (zh) * | 2023-07-10 | 2023-11-07 | 浙江工业大学 | 一种基于导向视觉语义对齐的遥感图文检索方法 |
| CN119011969A (zh) * | 2024-08-02 | 2024-11-22 | 浙江大学 | 一种基于检索的文生视频方法 |
| CN119149769A (zh) * | 2024-11-18 | 2024-12-17 | 中国海洋大学 | 基于自适应视角匹配的海洋遥感图文检索方法及系统 |
| CN120216612A (zh) * | 2025-05-28 | 2025-06-27 | 星际空间(天津)科技发展有限公司 | 知识驱动的地下空间信息检索方法、系统及设备 |
| CN120316300A (zh) * | 2025-06-10 | 2025-07-15 | 西北工业大学 | 基于生成式模型的文本信息引导的自进化目标检索方法 |
| CN121093956A (zh) * | 2025-11-11 | 2025-12-09 | 云南电网有限责任公司 | 一种基于图结构节点影响力的多文档关键短语提取方法 |
Families Citing this family (44)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN107330100B (zh) * | 2017-07-06 | 2020-04-03 | 北京大学深圳研究生院 | 基于多视图联合嵌入空间的图像-文本双向检索方法 |
| CN108288067B (zh) * | 2017-09-12 | 2020-07-24 | 腾讯科技(深圳)有限公司 | 图像文本匹配模型的训练方法、双向搜索方法及相关装置 |
| US11461383B2 (en) * | 2017-09-25 | 2022-10-04 | Equifax Inc. | Dual deep learning architecture for machine-learning systems |
| CN109255047A (zh) * | 2018-07-18 | 2019-01-22 | 西安电子科技大学 | 基于互补语义对齐和对称检索的图像-文本互检索方法 |
| CN109189930B (zh) * | 2018-09-01 | 2021-02-23 | 网易(杭州)网络有限公司 | 文本特征提取及提取模型优化方法以及介质、装置和设备 |
| CN109284414B (zh) * | 2018-09-30 | 2020-12-04 | 中国科学院计算技术研究所 | 基于语义保持的跨模态内容检索方法和系统 |
| CN109783657B (zh) * | 2019-01-07 | 2022-12-30 | 北京大学深圳研究生院 | 基于受限文本空间的多步自注意力跨媒体检索方法及系统 |
| CN110245238B (zh) * | 2019-04-18 | 2021-08-17 | 上海交通大学 | 基于规则推理和句法模式的图嵌入方法及系统 |
| CN111125395B (zh) * | 2019-10-29 | 2021-07-20 | 武汉大学 | 一种基于双分支深度学习的cad图纸检索方法及系统 |
| CN111324752B (zh) * | 2020-02-20 | 2023-06-16 | 中国科学技术大学 | 基于图神经网络结构建模的图像与文本检索方法 |
| CN111626058B (zh) * | 2020-04-15 | 2023-05-30 | 井冈山大学 | 基于cr2神经网络的图像-文本双编码实现方法及系统 |
| CN111651661B (zh) * | 2020-06-03 | 2023-02-14 | 拾音智能科技有限公司 | 一种图文跨媒体检索方法 |
| CN111708805B (zh) * | 2020-06-18 | 2025-01-10 | 腾讯科技(深圳)有限公司 | 数据查询方法、装置、电子设备及存储介质 |
| CN111858882B (zh) * | 2020-06-24 | 2022-08-09 | 贵州大学 | 一种基于概念交互和关联语义的文本视觉问答系统及方法 |
| CN111753553B (zh) * | 2020-07-06 | 2022-07-05 | 北京世纪好未来教育科技有限公司 | 语句类型识别方法、装置、电子设备和存储介质 |
| CN111914551B (zh) * | 2020-07-29 | 2022-05-20 | 北京字节跳动网络技术有限公司 | 自然语言处理方法、装置、电子设备及存储介质 |
| CN112052352B (zh) * | 2020-09-07 | 2024-04-30 | 北京达佳互联信息技术有限公司 | 视频排序方法、装置、服务器及存储介质 |
| CN112100457A (zh) * | 2020-09-22 | 2020-12-18 | 国网辽宁省电力有限公司电力科学研究院 | 一种基于元数据的多源异构数据集成方法 |
| CN113191375B (zh) * | 2021-06-09 | 2023-05-09 | 北京理工大学 | 一种基于联合嵌入的文本到多对象图像生成方法 |
| CN113536184B (zh) * | 2021-07-15 | 2022-05-31 | 广东工业大学 | 一种基于多源信息的用户划分方法及系统 |
| CN113742556B (zh) * | 2021-11-03 | 2022-02-08 | 南京理工大学 | 一种基于全局和局部对齐的多模态特征对齐方法 |
| CN114048350B (zh) * | 2021-11-08 | 2024-11-15 | 湖南麓湖数据科技有限公司 | 一种基于细粒度跨模态对齐模型的文本-视频检索方法 |
| CN114048351B (zh) * | 2021-11-08 | 2024-11-05 | 湖南大学 | 一种基于时空关系增强的跨模态文本-视频检索方法 |
| CN114329034B (zh) * | 2021-12-31 | 2024-08-09 | 武汉大学 | 基于细粒度语义特征差异的图像文本匹配判别方法及系统 |
| CN114612749B (zh) * | 2022-04-20 | 2023-04-07 | 北京百度网讯科技有限公司 | 神经网络模型训练方法及装置、电子设备和介质 |
| CN114998607B (zh) * | 2022-05-11 | 2023-01-31 | 北京医准智能科技有限公司 | 超声图像的特征提取方法、装置、电子设备及存储介质 |
| CN114973273B (zh) * | 2022-05-31 | 2025-12-23 | 北京金山数字娱乐科技有限公司 | 文本检测系统及方法 |
| CN115203459B (zh) * | 2022-06-23 | 2025-09-26 | 齐鲁工业大学(山东省科学院) | 基于Bert和自注意机制的图文匹配方法及系统 |
| US20240004913A1 (en) * | 2022-06-29 | 2024-01-04 | International Business Machines Corporation | Long text clustering method based on introducing external label information |
| CN115344736B (zh) * | 2022-08-12 | 2026-03-27 | 电子科技大学 | 一种渐进式的图像文本匹配方法 |
| CN115761036B (zh) * | 2022-09-22 | 2025-05-30 | 西北工业大学 | 基于多视图信息融合的教育领域图像场景图生成方法 |
| CN115757713A (zh) * | 2022-10-25 | 2023-03-07 | 厦门大学 | 面向视频文本检索的端到端多粒度对比学习方法 |
| CN116230003B (zh) * | 2023-03-09 | 2024-04-26 | 北京安捷智合科技有限公司 | 一种基于人工智能的音视频同步方法及系统 |
| CN117520590B (zh) * | 2024-01-04 | 2024-04-26 | 武汉理工大学三亚科教创新园 | 海洋跨模态图文检索方法、系统、设备及存储介质 |
| CN117874262B (zh) * | 2024-03-12 | 2024-06-04 | 北京邮电大学 | 一种基于渐进原型匹配的文本-动态图片跨模态检索方法 |
| CN118069920B (zh) * | 2024-04-19 | 2024-07-09 | 湖北华中电力科技开发有限责任公司 | 一种面向海量多网络协议终端设备接入的数据采集系统 |
| CN118410011B (zh) * | 2024-06-27 | 2025-05-06 | 维飒科技(西安)有限公司 | 工程文件数据自适应匹配方法及系统 |
| CN118917315B (zh) * | 2024-08-16 | 2025-04-04 | 山东省计算中心(国家超级计算济南中心) | 基于Bert与深度学习模型的威胁情报实体检测方法 |
| CN118918516B (zh) * | 2024-10-09 | 2024-12-27 | 山东大学 | 一种基于语义对齐的目标视频片段定位方法、系统及产品 |
| CN119646272B (zh) * | 2025-02-20 | 2025-05-13 | 西南交通大学 | 基于信息增强和多模态全局局部特征对齐的图文检索方法 |
| CN120296186B (zh) * | 2025-03-26 | 2026-01-09 | 国网安徽省电力有限公司超高压分公司 | 一种基于跨模态语义对齐的图像-文本检索方法 |
| CN119903206B (zh) * | 2025-04-02 | 2025-06-03 | 鲁东大学 | 一种基于动态语义增强的层次化文本多粒度食谱检索方法 |
| CN120951988B (zh) * | 2025-07-23 | 2026-03-03 | 四川省文化大数据有限责任公司 | 一种数据集的自动化质量评估方法 |
| CN120953425B (zh) * | 2025-10-20 | 2026-03-24 | 平安创科科技(北京)有限公司 | 基于共享索引的视觉令牌生成方法、装置、设备及介质 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7814040B1 (en) * | 2006-01-31 | 2010-10-12 | The Research Foundation Of State University Of New York | System and method for image annotation and multi-modal image retrieval using probabilistic semantic models |
| CN102138140A (zh) * | 2008-07-01 | 2011-07-27 | 多斯维公司 | 利用综合语义语境的信息处理 |
| CN106202413A (zh) * | 2016-07-11 | 2016-12-07 | 北京大学深圳研究生院 | 一种跨媒体检索方法 |
| CN107330100A (zh) * | 2017-07-06 | 2017-11-07 | 北京大学深圳研究生院 | 基于多视图联合嵌入空间的图像‑文本双向检索方法 |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9904966B2 (en) * | 2013-03-14 | 2018-02-27 | Koninklijke Philips N.V. | Using image references in radiology reports to support report-to-image navigation |
| US20170330059A1 (en) * | 2016-05-11 | 2017-11-16 | Xerox Corporation | Joint object and object part detection using web supervision |
| CN106095829B (zh) * | 2016-06-01 | 2019-08-06 | 华侨大学 | 基于深度学习与一致性表达空间学习的跨媒体检索方法 |
| CN106886601B (zh) * | 2017-03-02 | 2018-09-04 | 大连理工大学 | 一种基于子空间混合超图学习的交叉模态检索方法 |
| CN106777402B (zh) * | 2017-03-10 | 2018-09-11 | 山东师范大学 | 一种基于稀疏神经网络的图像检索文本方法 |
| US20180373955A1 (en) * | 2017-06-27 | 2018-12-27 | Xerox Corporation | Leveraging captions to learn a global visual representation for semantic retrieval |
-
2017
- 2017-07-06 CN CN201710545632.XA patent/CN107330100B/zh not_active Expired - Fee Related
-
2018
- 2018-01-29 US US16/622,570 patent/US11106951B2/en not_active Expired - Fee Related
- 2018-01-29 WO PCT/CN2018/074408 patent/WO2019007041A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7814040B1 (en) * | 2006-01-31 | 2010-10-12 | The Research Foundation Of State University Of New York | System and method for image annotation and multi-modal image retrieval using probabilistic semantic models |
| CN102138140A (zh) * | 2008-07-01 | 2011-07-27 | 多斯维公司 | 利用综合语义语境的信息处理 |
| CN106202413A (zh) * | 2016-07-11 | 2016-12-07 | 北京大学深圳研究生院 | 一种跨媒体检索方法 |
| CN107330100A (zh) * | 2017-07-06 | 2017-11-07 | 北京大学深圳研究生院 | 基于多视图联合嵌入空间的图像‑文本双向检索方法 |
Cited By (21)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109857895B (zh) * | 2019-01-25 | 2020-10-13 | 清华大学 | 基于多环路视图卷积神经网络的立体视觉检索方法与系统 |
| CN109857895A (zh) * | 2019-01-25 | 2019-06-07 | 清华大学 | 基于多环路视图卷积神经网络的立体视觉检索方法与系统 |
| CN110197521A (zh) * | 2019-05-21 | 2019-09-03 | 复旦大学 | 基于语义结构表示的视觉文本嵌入方法 |
| CN110298395B (zh) * | 2019-06-18 | 2023-04-18 | 天津大学 | 一种基于三模态对抗网络的图文匹配方法 |
| CN110298395A (zh) * | 2019-06-18 | 2019-10-01 | 天津大学 | 一种基于三模态对抗网络的图文匹配方法 |
| CN110489551A (zh) * | 2019-07-16 | 2019-11-22 | 哈尔滨工程大学 | 一种基于写作习惯的作者识别方法 |
| CN110489551B (zh) * | 2019-07-16 | 2023-05-30 | 哈尔滨工程大学 | 一种基于写作习惯的作者识别方法 |
| CN111814658A (zh) * | 2020-07-07 | 2020-10-23 | 西安电子科技大学 | 基于语义的场景语义结构图检索方法 |
| CN111814658B (zh) * | 2020-07-07 | 2024-02-09 | 西安电子科技大学 | 基于语义的场景语义结构图检索方法 |
| CN112860847B (zh) * | 2021-01-19 | 2022-08-19 | 中国科学院自动化研究所 | 视频问答的交互方法及系统 |
| CN112860847A (zh) * | 2021-01-19 | 2021-05-28 | 中国科学院自动化研究所 | 视频问答的交互方法及系统 |
| CN115033670A (zh) * | 2022-06-02 | 2022-09-09 | 西安电子科技大学 | 多粒度特征融合的跨模态图文检索方法 |
| CN116842212A (zh) * | 2023-05-26 | 2023-10-03 | 西安电子科技大学 | 基于边界框提取和语义一致性约束的文本-行人检索方法 |
| CN116484878A (zh) * | 2023-06-21 | 2023-07-25 | 国网智能电网研究院有限公司 | 电力异质数据的语义关联方法、装置、设备及存储介质 |
| CN116484878B (zh) * | 2023-06-21 | 2023-09-08 | 国网智能电网研究院有限公司 | 电力异质数据的语义关联方法、装置、设备及存储介质 |
| CN117009569A (zh) * | 2023-07-10 | 2023-11-07 | 浙江工业大学 | 一种基于导向视觉语义对齐的遥感图文检索方法 |
| CN119011969A (zh) * | 2024-08-02 | 2024-11-22 | 浙江大学 | 一种基于检索的文生视频方法 |
| CN119149769A (zh) * | 2024-11-18 | 2024-12-17 | 中国海洋大学 | 基于自适应视角匹配的海洋遥感图文检索方法及系统 |
| CN120216612A (zh) * | 2025-05-28 | 2025-06-27 | 星际空间(天津)科技发展有限公司 | 知识驱动的地下空间信息检索方法、系统及设备 |
| CN120316300A (zh) * | 2025-06-10 | 2025-07-15 | 西北工业大学 | 基于生成式模型的文本信息引导的自进化目标检索方法 |
| CN121093956A (zh) * | 2025-11-11 | 2025-12-09 | 云南电网有限责任公司 | 一种基于图结构节点影响力的多文档关键短语提取方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| US11106951B2 (en) | 2021-08-31 |
| CN107330100B (zh) | 2020-04-03 |
| US20210150255A1 (en) | 2021-05-20 |
| CN107330100A (zh) | 2017-11-07 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2019007041A1 (zh) | 基于多视图联合嵌入空间的图像-文本双向检索方法 | |
| US20200097604A1 (en) | Stacked cross-modal matching | |
| CN110222560B (zh) | 一种嵌入相似性损失函数的文本人员搜索方法 | |
| US8577882B2 (en) | Method and system for searching multilingual documents | |
| WO2021139191A1 (zh) | 数据标注的方法以及数据标注的装置 | |
| WO2020119508A1 (zh) | 视频切割方法、装置、计算机设备和存储介质 | |
| CN111259897A (zh) | 知识感知的文本识别方法和系统 | |
| CN112818889B (zh) | 基于动态注意力的超网络融合视觉问答答案准确性的方法 | |
| WO2023173552A1 (zh) | 目标检测模型的建立方法、应用方法、设备、装置及介质 | |
| CN116737979A (zh) | 基于上下文引导多模态关联的图像文本检索方法及系统 | |
| CN106095829A (zh) | 基于深度学习与一致性表达空间学习的跨媒体检索方法 | |
| WO2023168997A9 (zh) | 一种跨模态搜索方法及相关设备 | |
| CN106228120A (zh) | 查询驱动的大规模人脸数据标注方法 | |
| TWI749441B (zh) | 檢索方法及裝置、儲存介質 | |
| Cao et al. | An improved convolutional neural network algorithm and its application in multilabel image labeling | |
| WO2023109631A1 (zh) | 数据处理方法、装置、设备、存储介质及程序产品 | |
| CN119313966A (zh) | 基于多尺度跨模态提示增强的小样本图像分类方法及系统 | |
| CN105893573A (zh) | 一种基于地点的多模态媒体数据主题提取模型 | |
| CN114238715A (zh) | 基于社会救助的问答系统、构建方法、计算机设备及介质 | |
| Liang et al. | Multi-modal contextual graph neural network for text visual question answering | |
| CN106227836B (zh) | 基于图像与文字的无监督联合视觉概念学习系统及方法 | |
| CN111444313A (zh) | 基于知识图谱的问答方法、装置、计算机设备和存储介质 | |
| CN118643173A (zh) | 一种基于跨模态语义感知的图文检索方法及系统 | |
| CN118427631A (zh) | 一种基于知识增强的跨模态匹配方法及装置 | |
| CN118132788A (zh) | 应用于图文检索的图像文本语义匹配方法及系统 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 18828613 Country of ref document: EP Kind code of ref document: A1 |




