WO2018010365A1 - 一种跨媒体检索方法 - Google Patents
一种跨媒体检索方法 Download PDFInfo
- Publication number
- WO2018010365A1 WO2018010365A1 PCT/CN2016/108196 CN2016108196W WO2018010365A1 WO 2018010365 A1 WO2018010365 A1 WO 2018010365A1 CN 2016108196 W CN2016108196 W CN 2016108196W WO 2018010365 A1 WO2018010365 A1 WO 2018010365A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- text
- features
- vector
- image
- cross
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/33—Querying
- G06F16/3331—Query processing
- G06F16/334—Query execution
- G06F16/3344—Query execution using natural language analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/205—Parsing
- G06F40/216—Parsing using statistical methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/30—Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
- G06F16/33—Querying
- G06F16/3331—Query processing
- G06F16/334—Query execution
- G06F16/3347—Query execution using vector based model
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/50—Information retrieval; Database structures therefor; File system structures therefor of still image data
- G06F16/58—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/50—Information retrieval; Database structures therefor; File system structures therefor of still image data
- G06F16/58—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
- G06F16/583—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
- G06F18/241—Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches
- G06F18/2415—Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches based on parametric or probabilistic models, e.g. based on likelihood ratio or false acceptance rate versus a false rejection rate
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
- G06F18/243—Classification techniques relating to the number of classes
- G06F18/2431—Multiple classes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
- G06F40/284—Lexical analysis, e.g. tokenisation or collocates
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N20/00—Machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/09—Supervised learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N7/00—Computing arrangements based on specific mathematical models
- G06N7/01—Probabilistic graphical models, e.g. probabilistic networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T11/00—Two-dimensional [2D] image generation
- G06T11/60—Creating or editing images; Combining images with text
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
- G06V10/44—Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
- G06V10/443—Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components by matching or filtering
- G06V10/449—Biologically inspired filters, e.g. difference of Gaussians [DoG] or Gabor filters
- G06V10/451—Biologically inspired filters, e.g. difference of Gaussians [DoG] or Gabor filters with interaction between the filter responses, e.g. cortical complex cells
- G06V10/454—Integrating the filters into a hierarchical structure, e.g. convolutional neural networks [CNN]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/764—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F2218/00—Aspects of pattern recognition specially adapted for signal processing
- G06F2218/08—Feature extraction
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F2218/00—Aspects of pattern recognition specially adapted for signal processing
- G06F2218/12—Classification; Matching
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
- Y02D—CLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
- Y02D10/00—Energy efficient computing, e.g. low power processors, power management or thermal management
Definitions
- the invention belongs to the technical field of deep learning and multimedia retrieval, and relates to a cross-media retrieval method, in particular to a cross-media retrieval method for extracting image features and extracting text features by using a convolutional neural network.
- Cross-media retrieval is a medium that does not rely on a single modality, and can realize mutual retrieval between arbitrary modal media. Input information of any type of media, through cross-media retrieval, you can get related other media information in the multi-modal huge amount of data, and retrieve the results that meet the requirements more quickly.
- cross-media metrics mainly involve three key issues: cross-media metrics, cross-media indexing, and cross-media sorting.
- the typical methods for these three key problems are cross-media metrics based on matching models, cross-media indexing methods based on hash learning, and cross-media sorting methods based on sorting learning, as follows:
- the matching model is trained by the training data of the known category to mine the internal relationship between different types of data, and then the similarity between the cross-media data is calculated and returned.
- the most relevant search results There are two matching methods for matching models, one is based on correlation matching, such as the method using Canonical Correlation Analysis (CCA), and the other is based on semantic matching. (SemanticMatching, SM), such as the use of multi-class logistic regression method for semantic classification.
- CCA Canonical Correlation Analysis
- semantic matching SemanticMatching, such as the use of multi-class logistic regression method for semantic classification.
- Hash indexing is an effective way to speed up near-neighbor search.
- the method converts the original feature data into a binary hash code through the learned hash model, while maintaining the neighbor relationship in the original space as much as possible, that is, maintaining the correlation.
- cross-media sorting method based on sorting learning.
- the purpose of cross-media sorting is to learn a semantic similarity-based sorting model between different modalities.
- the specific method is to make a better sorting of the search results after retrieving the semantically related cross-media data, so that the more relevant data is more advanced, and the optimization process is iterated until the convergence is optimally searched.
- the present invention provides a new cross-media retrieval method, which utilizes a convolutional neural network (called VGG net) proposed by the Visual Geometry Group (VGG) to extract image features and utilizes Word2vec-based Fisher Vector extracts text features and performs semantic matching on heterogeneous images and text features through logistic regression to achieve cross-media retrieval.
- VGG net convolutional neural network
- VGG Visual Geometry Group
- Word2vec-based Fisher Vector extracts text features and performs semantic matching on heterogeneous images and text features through logistic regression to achieve cross-media retrieval.
- Existing cross-media retrieval methods are generally based on traditional manual extraction features and artificially defined traditions.
- the feature extraction method of the present invention can effectively represent the deep semantics of images and texts, and can improve the accuracy of cross-media retrieval, thereby greatly improving the cross-media retrieval effect.
- the principle of the present invention is to use the VGG convolutional neural network described in the document [1] (Simonyan K, Zisserman A. Very Deep Convolutional Networks for Large-Scale Image Recognition [J]. Computer Science, 2014) to extract image features, Using the Word2vec-based Fisher Vector (FV) feature as the text feature, and then using the semantic matching based on Semantic Matching (SM) method to find the association between the two heterogeneous features of image and text, thus achieving cross The purpose of media retrieval.
- the features proposed by the present invention can better express images and texts, and can improve the accuracy of cross-media retrieval.
- VGG net convolutional neural network
- Word2vec-based Fisher Vector uses logical regression to treat heterogeneous images and texts.
- the semantic matching is performed to achieve cross-media retrieval; the following steps are included:
- the VGG convolutional neural network method is used to extract the image features.
- step 2) uses the image features and text features in the training data obtained by performing step 2) and step 3) to train the semantic matching model based on logistic regression, and transform the text feature T into a text semantic feature ⁇ T , I ⁇ [1,n],c is the number of categories, and is also the dimension of the semantic features of the text; the image features I i are converted into semantic features composed of posterior probabilities, the posterior probability is K ⁇ [1, C], indicating the probability that the image I i belongs to the category k;
- step 4) trained semantic matching model, using the image features and text features of the test data obtained in steps 2 and 3, testing a picture or text to obtain related text or pictures, that is, cross-media Search Results.
- step 3 extracting text features using the Word2vec-based Fisher Vector method, including the following process:
- the GMM function is defined as Equation 1:
- Equation 2 p(x t
- ⁇ ) the probability value p(x t
- Equation 3 Equation 3
- Equation 4 Equation 4
- Equation 5 The probability that the vector x t is generated by the ith Gaussian function is expressed by ⁇ t (i), expressed as Equation 5:
- the Fisher Vector is a vector combining the derivation results of all the parameters;
- the number of Gaussian functions in the Gaussian mixture model is G, and the vector dimension is dw,
- the dimension of the Fisher Vector is (2 ⁇ dw+1) ⁇ G-1;
- the degree of freedom of the weight ⁇ is N-1;
- step 34 the parameter of the Gaussian model is deductively guided. Specifically, the derivation formula for each parameter is as shown in Equation 6-8:
- the posterior probability is K ⁇ [1,C], indicating the probability that the text T i belongs to the category k, which is calculated by Equation 9.
- Is a parameter of a multi-category logistic regression linear classifier, Express Transpose, Corresponding to the category k, where D T (2 ⁇ dw+1) ⁇ G-1, D T is the dimension of the text feature;
- the posterior probability is K ⁇ [1,C], indicating the probability that the image I i belongs to the category k, where The formula is as follows:
- I a parameter of a multi-category logistic regression linear classifier
- the corresponding category k is a vector of D I dimensions
- D I is the dimension of the image features.
- the step 5 tests a picture or text to obtain a related text or picture;
- the correlation measurement method includes a Kullback-Leibler divergence method, a Normalized Correlation method, and Centered.
- the Correlation method and the L2 paradigm method One or more of the Correlation method and the L2 paradigm method.
- the present invention uses the VGG convolutional neural network to extract image features, uses the Word2vec-based Fisher Vector (FV) feature as a text feature, and both the image and the text use a neural network to extract features.
- VGG convolutional neural network uses the Word2vec-based Fisher Vector (FV) feature as a text feature
- FV Fisher Vector
- both the image and the text use a neural network to extract features.
- neural network features are more complex and more capable of representing the content of images and text. Therefore, the use of neural network features for cross-media retrieval will greatly improve the retrieval effect.
- the present invention has the following advantages: First, the present invention employs a neural network to simulate a biological vision neural network system, representing pixel-level features as higher-level, more abstract features for interpreting image data. Secondly, the technical solution of the present invention benefits from the improvement of computer computing performance, and the neural network features are obtained through more complicated calculations, and can achieve good effects after training through large-scale data.
- FIG. 1 is a flow chart of a cross-media retrieval method provided by the present invention.
- (a) is a pair of images in the wikipedia data set; (b) is the text corresponding to the image, and the text is presented in the form of a long paragraph.
- FIG. 3 is an image and text example of a pascal sentence data set according to an embodiment of the present invention.
- (a) is a pair of images in the pascal sentence data set; (b) is the text corresponding to the image, and the text is five sentences.
- the invention provides a new cross-media retrieval method, which uses a convolutional neural network (called VGG net) proposed by the Visual Geometry Group (VGG) to extract image features, and uses the Word2vec-based Fisher Vector to extract text features through logistic regression.
- VGG net convolutional neural network
- VGG Visual Geometry Group
- the method performs semantic matching on heterogeneous images and text features to achieve cross-media retrieval; existing cross-media retrieval methods are generally based on traditional artificial extraction features, and feature extraction of the present invention compared with manually defined traditional features.
- the method can effectively represent the deep semantics of images and texts, and can improve the accuracy of cross-media retrieval, thereby greatly improving the cross-media retrieval effect.
- FIG. 1 is a flow chart of a cross-media retrieval method provided by the present invention, including the following steps:
- Step 1 collecting a cross-media retrieval data set containing two types of media for image and text, and dividing the image and text into training data and test data respectively;
- Step 2 For all image data in the data set, the image features are extracted using the VGG convolutional neural network method.
- Step 3 For the text features in the dataset, use the Word2vec-based Fisher Vector method to extract the text features.
- Step 4 Train the semantic matching model based on logistic regression using the image and text features in the training data obtained after steps 2 and 3.
- step 5 the trained semantic matching model is used to test the image and text features of the test data obtained in steps 2 and 3 to test the effect of the present invention.
- Each step specifically includes the following process:
- Step 1 Collect a cross-media retrieval data set containing two types of media for image and text, including category labels (such as in the pascal sentence data set, divided into 20 categories, with aircraft, cars, birds, etc.), and divide the data set into Training data and test data.
- category labels such as in the pascal sentence data set, divided into 20 categories, with aircraft, cars, birds, etc.
- L [l 1 , l 2 , ..., l n ], where l i ⁇ [1, 2, ..., C], C is the number of categories, and l i represents the category to which the i-th pair of images and text belong.
- Step 2 For all image data in the data set, the image features are extracted using the VGG convolutional neural network method.
- the VGG convolutional neural network has five configurations of A to E, and the number of convolution layers is increased from 8 to 16. In the embodiment of the present invention, preferably, the number of convolution layers used is 16 layers, plus 3 fully connected layers, and a total of 19 layers of VGG networks.
- a 4096-dimensional vector is obtained at the seventh-layer fully-connected layer (fc7).
- this vector is used as the image feature.
- the original image data D I Enter the VGG network and extract image features.
- Step 3 For the text features in the dataset, use the Word2vec-based Fisher Vector method to extract the text features.
- f word2vec (w) W i , i ⁇ [1, n]. which is Where w i,j ⁇ R dw ,j ⁇ [1,b i ],w i,j is Contains the word vector corresponding to the word, dw is the dimension of the word vector, and b i is The number of words contained in it.
- GMM Gaussian Mixture Model
- the GMM function is defined as follows:
- Equation 2 p(x t
- ⁇ ) the probability value p(x t
- the weight ⁇ i has the following constraint, the sum is 1, which is expressed as Equation 3:
- Equation 4 Equation 4
- Equation 5 The probability that the vector x t is generated by the ith Gaussian function is expressed by ⁇ t (i), expressed as Equation 5:
- the Fisher Vector is obtained by deriving the parameters of the Gaussian model, and the derivation formula for each parameter is as shown in Equation 6 to Equation 8, where the superscript d represents the dth dimension of the vector:
- the Fisher Vector is a vector that combines the derivation results of all the above parameters. Because the number of Gaussian functions in the Gaussian mixture model is G and the vector dimension is dw, the dimension of the Fisher Vector is (2 ⁇ dw+1) ⁇ G-1; for the weight ⁇ , the constraint with a sum of 1 is free. The degree is G-1; G is the number of Gaussian functions in the Gaussian model.
- step 4 the semantic regression-based semantic matching model is trained using the image and text features in the training data obtained after performing steps 2 and 3.
- L [l 1 , l 2 , ..., l n ], where l i ⁇ [1, 2, ..., C].
- Equation 9 We convert the text feature T i into a semantic feature consisting of posterior probabilities, and the posterior probability is K ⁇ [1, C], indicating the probability that the text T i belongs to the category k, where the calculation is obtained by Equation 9.
- Is a multi-class logistic regression parameter, Express Transpose, Corresponding to the category k, where D T (2 ⁇ dw + 1) ⁇ G - 1, D T is the dimension of the text feature.
- the posterior probability is K ⁇ [1,C], indicating the probability that the image I i belongs to the category k, where The formula is as follows:
- I a parameter of multi-class logistic regression
- the corresponding category k is a vector of D I dimensions
- D I is the dimension of the image features.
- Equation 12 the dth dimension in the vector is expressed as Equation 12:
- the above calculates the semantic features of images and texts, and trains to obtain a semantic matching model.
- Step 5 using the semantic matching model trained in step 4, using the image and text features of the test data obtained in steps 2 and 3, testing a picture (or text) to obtain related text (or picture); The effects of the present invention were examined.
- the correlation between the image semantic feature ⁇ I and the text semantic feature ⁇ T is calculated, and the text semantic features ⁇ T are sorted according to the correlation from large to small, and the more relevant the image ⁇ I is, the more the text before.
- the correlation between the text semantic feature ⁇ T and the image semantic feature ⁇ I is calculated, and the image semantic feature ⁇ I is sorted according to the correlation from large to small, and the image related to the text ⁇ T is more relevant. The more forward.
- Correlation measures include Kullback–Leibler divergence (KL), Normalized Correlation (NC), Centered Correlation (CC), and L2 paradigm (L2).
- the MAP value (Mean Average Precision) is calculated to measure the search result.
- the first example uses wikipedia dataset, including 2866 pairs of images and their texts, there are 10 categories, namely: Art&architecture (Art & Architecture), Biology (Bio), Geography & places (Geography & Location) , History, Literature & theatre, Media, Music, Royality & Nobility, Sport & Regulation, Warfare. 2173 data were divided into training data and 693 data were test data.
- An example of the image and text of the data set is shown in Figure 2, each image corresponding to a long piece of text. Image features and text features are obtained by steps 2 and 3.
- the text data first extracts the first two topic sentences of each text with textteaser (an open source text automatic summary tool), extracts the Fisher Vector features for each topic sentence, and then connects the Fisher Vector features of the two sentences together. Form a higher dimensional feature as the final feature.
- the Fisher vector feature is d-dimensional, and the final feature after the two sentences are connected is the 2d dimension.
- the semantic matching model is obtained, and the test sample is processed according to step 5 to obtain the retrieval result.
- the experimental results show that compared with the existing methods, the method of the present invention achieves superior results in both the Img2Text and Text2Img tasks.
- Methods for extracting traditional artificial features for cross-media retrieval for comparison include CCA [2], LCFS [3], CDLFA [4], HSNN [5].
- the text features they use are 10-dimensional Latent Dirichlet Allocation (LDA) features, and the image features are 128-dimensional SIFT features.
- LDA Latent Dirichlet Allocation
- the present invention is compared with the results of the latest paper CVF [6] using deep learning for cross-media retrieval.
- the CVF[6] Chinese text feature uses 100-dimensional LDA features, and the image features use the 4096-dimensional DeCAF deep network CNN. feature.
- the second embodiment uses the Pascal Sentence dataset, which contains 1000 pairs of image-text data, divided into 20 categories (corresponding to category labels), including aircraft, cars, birds, etc., as shown in Table 2; each category contains 50 pairs of images and text.
- FIG. 3 An example of image and text data is shown in Figure 3, each image corresponding to 5 sentences. From each class, 30 pairs of images and texts were randomly selected, a total of 600 pairs were used as training data, and the remaining 400 pairs were used as test data. Extracted by step 2 and step 3 Corresponding image features and text features, in which the text data in Pascal Sentence is already a sentence, no need to do text summary processing, you can directly extract the Fisher Vector feature, the Fisher vector feature of a sentence is d dimension, and then, according to step 4 The training obtains a semantic matching model, and the test sample is processed according to step 5 to obtain a search result.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Artificial Intelligence (AREA)
- General Engineering & Computer Science (AREA)
- Evolutionary Computation (AREA)
- Data Mining & Analysis (AREA)
- General Health & Medical Sciences (AREA)
- Health & Medical Sciences (AREA)
- Software Systems (AREA)
- Computational Linguistics (AREA)
- Computing Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Databases & Information Systems (AREA)
- Mathematical Physics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Biomedical Technology (AREA)
- Molecular Biology (AREA)
- Medical Informatics (AREA)
- Multimedia (AREA)
- Probability & Statistics with Applications (AREA)
- Biophysics (AREA)
- Library & Information Science (AREA)
- Evolutionary Biology (AREA)
- Bioinformatics & Computational Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Algebra (AREA)
- Computational Mathematics (AREA)
- Mathematical Analysis (AREA)
- Mathematical Optimization (AREA)
- Pure & Applied Mathematics (AREA)
- Biodiversity & Conservation Biology (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
一种跨媒体检索方法,利用VGG提出的卷积神经网络VGG net提取图像特征,将VGG卷积神经网络中的第七层全连接层fc7通过ReLU激活函数之后的4096维特征作为图像特征;利用基于Word2vec的Fisher Vector提取文本特征,通过逻辑回归的方法对异构图像、文本特征进行语义匹配,通过基于逻辑回归的语义匹配方法找到图像、文本这两种异构特征之间的关联,从而实现跨媒体检索;上述特征提取方法能有效地表示图像和文本的深层语义,可提高跨媒体检索的准确度,从而大幅度提升跨媒体检索效果。
Description
本发明属于深度学习和多媒体检索技术领域,涉及跨媒体检索方法,尤其涉及一种利用卷积神经网络提取图像特征和Fisher Vector提取文本特征的跨媒体检索方法。
随着互联网的高速发展,图像、文本、视频、音频等不同类型的多媒体数据呈现出爆炸性的增长。这些多媒体数据经常会同时出现,用来描述一个相同的事物。不同模态的信息反映了事物的不同属性,人们需要获取不同模态的信息来满足对事物不同形式的描述的需求。比如,对于一副图像,我们想要找到与其相关的文字描述;或者对于一段文本,找到符合这段文本语义的图像或是视频。要满足上述需求,就需要实现跨媒体检索的相关技术。
现有检索系统大都是建立在单一模态文本信息的基础上,例如谷歌、百度等搜索引擎。通过查询请求检索图像、音频、视频的功能本质上是对一个由文字信息组成的元数据库上的内容匹配,这种检索仍然属于传统的基于关键字的检索技术。虽然关键字能够准确地描述概念的细节信息,但是它很难完整、生动地呈现一幅图片或一段视频的内容,并可能带有标注人的主观意愿。其固有缺陷使得大批学者开始转向研究基于内容的检索技术,通过充分挖掘多媒体数据的语义关联,使计算机能够更准确地理解多媒体信息表达的内容。然而,基于内容的检索一般只关注媒体底层特征,且通常针对单一模态媒体对象,使得查询和检索结果必须为相同的模态,无法实现跨越各种媒体类型的综合检索。因此,跨媒体检索的概念被提出。跨媒体检索是不依托于某个单一模态的媒体,可以实现任意模态媒体之间的相互检索。输入任意类型媒体的信息,通过跨媒体检索即可得到相关的其他媒体信息在多模态的巨量数据中,更快地检索出符合要求的结果。
现有的跨媒体检索方法主要涉及三个关键问题:跨媒体度量、跨媒体索引、跨媒体排序。针对这三个关键问题的典型方法分别是基于匹配模型的跨媒体度量方法、基于哈希学习的跨媒体索引方法和基于排序学习的跨媒体排序方法,具体如下:
第一,基于匹配模型的跨媒体度量方法,通过已知类别的训练数据对匹配模型进行训练,来挖掘不同类型数据之间的内在联系,进而对跨媒体数据之间的相似度进行计算,返回相关性最高的检索结果。匹配模型有两种匹配方法,一种是基于相关性的匹配,如利用典型相关性分析(Canonical Correlation Analysis,CCA)的方法;另一种是基于语义的匹配
(SemanticMatching,SM),如利用多类逻辑回归的方法进行语义分类。
第二,基于哈希学习的跨媒体索引方法。由于互联网中海量大数据的出现,使得人们对检索速度提出了更高的要求。哈希索引是加快近似近邻检索的一种有效方法。该方法通过学习到的哈希模型将原始特征数据转化为二进制哈希码,同时尽可能地保持原空间中的近邻关系,即保持相关性。
第三,基于排序学习的跨媒体排序方法。跨媒体排序的目的是学习不同模态之间的基于语义相似度的排序模型。具体做法是在检索出语义相关的跨媒体数据之后,对检索结果做一个更优的排序,使得相关性更高的数据更加靠前,不断迭代优化过程,直到收敛得到最优检索。
上述这些方法中,所用的图像和文本特征几乎都是使用人工定义的传统特征,如SIFT特征。随着计算机处理性能和计算能力的不断提高,传统的人工特征极大地阻碍了跨媒体检索性能的提升,近一年,人们开始关注深度学习相关技术与跨媒体检索的结合。事实证明,深度学习的有效应用往往能对检索效果带来突破性的进展。
发明内容
为了克服上述现有技术的不足,本发明提供一种新的跨媒体检索方法,利用Visual Geometry Group团队(简称VGG)提出的卷积神经网络(称作VGG net)提取图像特征,利用基于Word2vec的Fisher Vector提取文本特征,通过逻辑回归的方法对异构图像、文本特征进行语义匹配,从而实现跨媒体检索;现有跨媒体检索方法普遍都是基于传统的人工提取的特征,与人工定义的传统特征相比,本发明的特征提取方法能有效地表示图像和文本的深层语义,可提高跨媒体检索的准确度,从而大幅度提升跨媒体检索效果。
本发明的原理是:将文献[1](Simonyan K,Zisserman A.Very Deep Convolutional Networks for Large-Scale Image Recognition[J].Computer Science,2014)记载的VGG卷积神经网络用来提取图像特征,使用基于Word2vec的Fisher Vector(简称,FV)特征作为文本特征,再通过基于逻辑回归的语义匹配(Semantic Matching,SM)方法找到图像、文本这两种异构特征之间的关联,由此达到跨媒体检索的目的。本发明所提出的特征能更好的对图像和文本进行表达,可提高跨媒体检索的准确度。
本发明提供的技术方案是:
一种跨媒体检索方法,利用VGG提出的卷积神经网络(称作VGG net)提取图像特征,利用基于Word2vec的Fisher Vector提取文本特征,通过逻辑回归的方法对异构图像、文本特
征进行语义匹配,从而实现跨媒体检索;包括如下步骤:
1)收集含有类别标签的跨媒体检索数据集,设为D={D1,D2,…,Dn},n表示数据集的大小;所述跨媒体检索数据集中数据的类型包括图像和文本两种媒体类型,表示为图像-文本对Di(Di∈D),其中表示图像的原始数据,表示文本的原始数据;类别标签设为L,L=[l1,l2,…,ln],其中li∈[1,2,…,C],C为类别的数目,li表示第i对图像和文本所属的类别;将所述跨媒体检索数据集划分为训练数据和测试数据;
2)对于数据集D中的所有图像数据DI,其中使用VGG卷积神经网络方法提取得到图像特征,将VGG卷积神经网络中的第七层全连接层fc7通过ReLU激活函数之后的4096维特征,记作I={I1,I2,…,In},其中Ij∈R4096,j∈[1,n],作为图像特征;
3)对于数据集中的文本特征数据DT,其中使用基于Word2vec的Fisher Vector方法提取文本特征;具体将DT转换成词向量集合W={W1,W2,…,Wn},W为DT包含的单词的词向量集合;将W={W1,W2,…,Wn}中的每个文本词向量集合Wi代入式1中的X,求得每个文本的Fisher Vector,记作T={T1,T2,…,Tn},Ti∈R(2×dw+1)×G-1,i∈[1,n],其中,Ti表示由第i个文本计算出来的Fisher Vector;由此提取得到文本特征;
4)使用执行步骤2)和步骤3)得到的训练数据中的图像特征和文本特征对基于逻辑回归的语义匹配模型进行训练,将文本特征T转换成了文本语义特征ПT,i∈[1,n],c是类别的个数,也是文本语义特征的维数;将图像特征Ii转换成后验概率组成的语义特征,后验概率为k∈[1,C],表示图像Ii属于类别k的概率;
5)利用步骤4)训练好的语义匹配模型,使用步骤2和步骤3得到的测试数据的图像特征和文本特征,针对一幅图片或文本进行测试,得到相关的文本或图片,即为跨媒体检索结果。
针对上述跨媒体检索方法,进一步地,步骤3)使用基于Word2vec的Fisher Vector方法提取文本特征,具体包括如下过程:
32)将单词记作w,单词w所对应的词向量为fword2vec(w);对于有fword2vec(w)∈Wi,i∈[1,n],即其中wi,j∈Rdw,j∈[1,bi],wi,j为包含单词所对应的词向量,dw为词向量的维度,bi为中包含的单词个数;
33)用X={x1,x2,…,xnw}表示一个文本的词向量集合,nw为词向量个数;令混合高斯模型GMM的参数为λ,λ={ωi,μi,∑i,i=1..G},其中ωi,μi,∑i分别表示GMM中每个高斯函数的权重、均值向量和协方差矩阵,G表示模型中高斯函数的个数;
GMM函数定义为式1:
其中,p(xt|λ)表示对于向量xt(t∈[1,nw]),由GMM产生的概率值p(xt|λ),表示为式2:
对权重ωi设置总和为1约束,表示为式3:
其中,pi(x|λ)表示GMM中的第i个高斯函数,由式4给出:
其中,dw是向量的维度,|∑i|表示求∑i的行列式;
用γt(i)来表示向量xt由第i个高斯函数产生的概率,表示为式5:
34)对高斯模型的参数求偏导即得到Fisher Vector;所述Fisher Vector是将所有参数的求导结果连接组成的向量;所述高斯混合模型中高斯函数个数为G,向量维度为dw,所述Fisher Vector的维度为(2×dw+1)×G-1;权重ω的自由度为N-1;
35)将W={W1,W2,…,Wn}中的每个文本词向量集合Wi代入式1中的文本的词向量集合X,求得每个文本的Fisher Vector,记作T={T1,T2,…,Tn},Ti∈R(2×dw+1)×G-1,i∈[1,n],其中,Ti表示由第i个文本计算出来的Fisher Vector。
更进一步地,步骤34)所述对高斯模型的参数求偏导,具体地,对各个参数的求导公式如式6~式8:
其中,上标d表示向量的第d个维度。
针对上述跨媒体检索方法,进一步地,步骤4)所述使用训练数据中的图像特征和文本特征对基于逻辑回归的语义匹配模型进行训练,所述图像特征为I={I1,I2,…,In},Ij∈R4096;所述文本特征为T={T1,T2,…,Tn},Ti∈R(2×dw+1)×G-1;相应的图像特征和文本特征具有共同的标签为L=[l1,l2,…,ln],其中li∈[1,2,…,C];所述训练具体包括:
针对上述跨媒体检索方法,进一步地,步骤5)所述针对一幅图片或文本进行测试,得到相关的文本或图片;所述相关性的度量方法包括Kullback–Leibler divergence方法、Normalized Correlation方法、Centered Correlation方法和L2范式方法中的一种或多种。
与现有技术相比,本发明的有益效果是:
本发明使用VGG卷积神经网络提取图像特征,使用基于Word2vec的Fisher Vector(FV)特征作为文本特征,图像和文本都使用了神经网络提取特征的方法。与传统的人工特征相比,神经网络特征更加复杂,更能表现出图像和文本的内容。所以,使用神经网络特征来进行跨媒体检索,在检索效果上会有较大提升。
具体地,本发明具有如下优点:第一,本发明采用神经网络模拟生物视觉神经网络系统,将像素级别的特征表示成高层的更加抽象的特征,用来解释图像数据。第二,本发明技术方案得益于计算机计算性能的提升,神经网络特征经过更加复杂的计算得到,能够在通过大规模数据的训练后取得很好的效果。
图1是本发明提供的跨媒体检索方法的流程框图。
图2是本发明实施例采用wikipedia数据集中的图像和文本实例;
其中,(a)是wikipedia数据集中的一副图像;(b)是该图像所对应的文本,文本呈现形式为长段落。
图3是本发明实施例采用pascal sentence数据集的图像和文本实例;
其中,(a)是pascal sentence数据集中的一副图像;(b)是该图像所对应的文本,文本为五个句子。
下面结合附图,通过实施例进一步描述本发明,但不以任何方式限制本发明的范围。
本发明提供一种新的跨媒体检索方法,利用Visual Geometry Group团队(简称VGG)提出的卷积神经网络(称作VGG net)提取图像特征,利用基于Word2vec的Fisher Vector提取文本特征,通过逻辑回归的方法对异构图像、文本特征进行语义匹配,从而实现跨媒体检索;现有跨媒体检索方法普遍都是基于传统的人工提取的特征,与人工定义的传统特征相比,本发明的特征提取方法能有效地表示图像和文本的深层语义,可提高跨媒体检索的准确度,从而大幅度提升跨媒体检索效果。
图1是本发明提供的跨媒体检索方法的流程框图,包括如下步骤:
步骤1,收集含有类别标签的针对图像和文本两种媒体类型的跨媒体检索数据集,分别将图像和文本划分为训练数据和测试数据;
步骤2,对于数据集中的所有图像数据,使用VGG卷积神经网络的方法提取图像特征。
步骤3,对于数据集中的文本特征,使用基于Word2vec的Fisher Vector方法提取文本特征。
步骤4,使用步骤2,3后得到的训练数据中的图像和文本特征对基于逻辑回归的语义匹配模型进行训练。
步骤5,利用训练好的语义匹配模型,使用步骤2,3得到的测试数据的图像和文本特征进行测试,检验本发明的效果。
各步骤具体包括如下过程:
步骤1,收集含有类别标签(如在pascal sentence数据集中,分为20类,有飞机,汽车,鸟等类别)的针对图像和文本两种媒体类型的跨媒体检索数据集,将数据集划分为训练数据和测试数据。
将数据集定义为D={D1,D2,…,Dn},其中n表示数据集的大小,对数据集中的任一
图像-文本对Di(Di∈D),可表示为其中表示图像的原始数据,表示文本的原始数据。L=[l1,l2,…,ln],其中li∈[1,2,…,C],C为类别的数目,li表示第i对图像和文本所属的类别。
步骤2,对于数据集中的所有图像数据,使用VGG卷积神经网络的方法提取图像特征。
VGG卷积神经网络有A~E五种配置,卷积层数从8到16递增。本发明实施例中,优选地,使用的卷积层数为16层,再加上3个全连接层,一共是19层的VGG网络。
每幅图像输入VGG网络后,在第七层全连接层(fc7)得到一个4096维的向量,通过ReLU(Rectified Linear Units)激活函数后,用这个向量作为图像特征。具体地,将原始图像数据DI,其中输入VGG网络中并提取图像特征。图像特征是第七层全连接层(fc7)通过ReLU(Rectified Linear Units)激活函数之后的4096维特征,记作I={I1,I2,…,In},其中Ij∈R4096,j∈[1,n]。
步骤3,对于数据集中的文本特征,使用基于Word2vec的Fisher Vector方法提取文本特征。
进一步地,将单词记作w,单词w所对应的词向量为fword2vec(w),则对于
有fword2vec(w)∈Wi,i∈[1,n]。即其中wi,j∈Rdw,j∈[1,bi],wi,j为包含单词所对应的词向量,dw为词向量的维度,bi为中包含的单词个数。
这里先假设用X={x1,x2,…,xnw}表示一个文本的词向量集合,nw为词向量个数。令混合高斯模型(Gaussion Mixture Model,GMM)参数为λ,则λ={ωi,μi,∑i,i=1..G},其中ωi,μi,∑i分别表示GMM中每个高斯函数的权重、均值向量和协方差矩阵,G表示模型中高斯函数的个数。
对GMM函数定义如下:
其中,p(xt|λ)表示对于向量xt(t∈[1,nw]),由GMM产生的概率值p(xt|λ),表示为式2:
对权重ωi有如下约束,总和为1,表示为式3:
其中,pi(x|λ)表示GMM中的第i个高斯函数,由式4给出:
其中,dw是向量的维度,|∑i|表示求∑i的行列式
用γt(i)来表示向量xt由第i个高斯函数产生的概率,表示为式5:
对高斯模型的参数求偏导即得到Fisher Vector,对各个参数的求导公式如式6~式8,其中,上标d表示向量的第d个维度:
Fisher Vector就是将上述所有参数的求导结果连接组成的向量。因为高斯混合模型中高斯函数个数为G,向量维度为dw,所以,Fisher Vector的维度为(2×dw+1)×G-1;对于权重ω,含有总和为1的约束条件,其自由度为G-1;G为高斯模型中高斯函数的个数。
最后,将W={W1,W2,…,Wn}中的每个文本词向量集合Wi代入式1中的X,求得每个文本的Fisher Vector,记作T={T1,T2,…,Tn},Ti∈R(2×dw+1)×G-1,i∈[1,n],其中,Ti表示由第i个文本计算出来的Fisher Vector。
步骤4,使用执行步骤2、3之后得到的训练数据中的图像和文本特征对基于逻辑回归的语义匹配模型进行训练。
得到的图像特征为I={I1,I2,…,In},Ij∈R4096。
得到的文本特征为T={T1,T2,…,Tn},Ti∈R(2×dw+1)×G-1。
对于相应的图像和文本特征,有着共同的标签,L=[l1,l2,…,ln],其中li∈[1,2,…,C]。
以上对图像和文本语义特征进行计算,训练得到语义匹配模型。
步骤5,利用步骤4训练好的语义匹配模型,使用步骤2和步骤3得到的测试数据的图像和文本特征,针对一幅图片(或文本)进行测试,得到相关的文本(或图片);并检验本发明的效果。
对于图像检索文本(Img2Text),计算图像语义特征ПI和文本语义特征ПT的相关性,将文本语义特征ПT按相关性从大到小排序,则和图像ПI越相关的文本越靠前。
同理,对于文本检索图像(Text2Img)计算文本语义特征ПT和图像语义特征ПI的相关性,将图像语义特征ПI按相关性从大到小排序,则和文本ПT越相关的图像越靠前。
其中相关性的度量方法包括Kullback–Leibler divergence(KL)、Normalized Correlation(NC)、Centered Correlation(CC)以及L2范式(L2)。
对于图像检索文本(Img2Text)和文本检索图像(Text2Img)的结果,计算其MAP值(Mean Average Precision),衡量检索结果。
在具体实施实验中,实施例一使用wikipedia的数据集,共包括2866对图像及其文本,有10个类别,分别为:Art&architecture(艺术&建筑)、Biology(生物)、Geography&places(地理&地点)、History(历史)、Literature&theatre(文学&戏剧)、Media(媒体)、Music(音乐)、Royalty&nobility(皇室&贵族)、Sport&recreation(运动&娱乐)、Warfare(战争)。划分其中的2173个数据为训练数据,693个数据为测试数据。数据集的图像和文本实例如图2所示,每个图像对应一段长文本。通过步骤2和步骤3得到图像特征和文本特征。其中,文本数据先用textteaser(一种开源文本自动摘要工具)提取出每个文本的前两个主题句,对于每个主题句提取Fisher Vector特征,然后将这两句的Fisher Vector特征连接在一起形成更高维度的特征,作为最终的特征。如一句话的Fisher vector特征是d维,两句话连接后的最终特征是2d维。之后,按照步骤4训练得到语义匹配模型,按照步骤5对待测试样本得到检索结果。
实验结果表明,与现有方法相比,本发明方法在Img2Text和Text2Img两个任务中,都取得了较优的结果。用于对比的提取传统人工特征进行跨媒体检索的方法包括CCA[2],LCFS[3],CDLFA[4],HSNN[5]。他们使用的文本特征为10维的隐狄利克雷分布(Latent Dirichlet Allocation,LDA)特征,图像特征为128维的SIFT特征。
同时本发明与最新的利用深度学习进行跨媒体检索的论文CVF[6]中的结果进行比较。CVF[6]中文本特征使用100维的LDA特征,图像特征使用4096维的DeCAF深度网络的CNN
特征.
下表给出了实验结果,Proposed表示的是本发明的结果,通过对比可知,本发明较CCA[2],LCFS[3],CDLFA[4],HSNN[5]中的方法效果有很大提升,和最新的CVF[6]中的方法效果相近,使用CC相关性度量的方法较CVF[6]效果有一定的提升。
表1 Wikipedia数据集实验结果
第二个实施例使用Pascal Sentence数据集,该数据集包含1000对图像-文本数据,分为20类(对应类别标签),包括飞机、汽车、鸟等类别,如表2所示;每类包含50对图像和文本。
表2 Pascal Sentence数据集的20个类别
| aeroplane | 飞机 | diningtable | 饭桌 |
| bicycle | 自行车 | dog | 狗 |
| boat | 船 | house | 房子 |
| bird | 鸟 | motorbike | 摩托车 |
| bottle | 瓶子 | person | 人 |
| bus | 公交车 | pottedplant | 盆栽 |
| car | 汽车 | sheep | 羊 |
| cat | 猫 | sofa | 沙发 |
| chair | 椅子 | train | 火车 |
| cow | 牛 | tvmonitor | 电视 |
图像和文本数据实例如图3所示,每个图像对应5个句子。从每类中随机抽取30对图像和文本,共600对作为训练数据,其余的400对作为测试数据。通过步骤2和步骤3提取出
相应的图像特征和文本特征,其中,由于Pascal Sentence中的文本数据已经是句子,不需要做文本摘要处理,可直接提取Fisher Vector特征,一句话的Fisher vector特征是d维,然后,按照步骤4训练得到语义匹配模型,按照步骤5对待测试样本得到检索结果。
由于文献[2]~[5]中记载的方法没有使用本数据集做评测,我们直接与CVF[6]的结果进行比较,结果如表3:
表3 PascalSentence数据集实验结果
从实验结果可以看出,我们的方法对于Pascal Sentence数据集的检索正确率有较大提升。
需要注意的是,公布实施例的目的在于帮助进一步理解本发明,但是本领域的技术人员可以理解:在不脱离本发明及所附权利要求的精神和范围内,各种替换和修改都是可能的。因此,本发明不应局限于实施例所公开的内容,本发明要求保护的范围以权利要求书界定的范围为准。
Claims (5)
- 一种跨媒体检索方法,利用VGG提出的卷积神经网络提取图像特征,利用基于Word2vec的Fisher Vector提取文本特征,通过逻辑回归的方法对异构图像特征和文本特征进行语义匹配,从而实现跨媒体检索;包括如下步骤:1)收集含有类别标签的跨媒体检索数据集,设为D={D1,D2,…,Dn},n表示数据集的大小;所述跨媒体检索数据集中数据的类型包括图像和文本两种媒体类型,表示为图像-文本对Di(Di∈D),其中表示图像的原始数据,表示文本的原始数据;类别标签设为L,L=[l1,l2,…,ln],其中li∈[1,2,…,C],C为类别的数目,li表示第i对图像和文本所属的类别;将所述跨媒体检索数据集划分为训练数据和测试数据;2)对于数据集D中的所有图像数据DI,其中使用VGG卷积神经网络方法提取得到图像特征,将VGG卷积神经网络中的第七层全连接层fc7通过ReLU激活函数之后的4096维特征,记作I={I1,I2,…,In},其中Ij∈R4096,j∈[1,n],作为图像特征;3)对于数据集中的文本特征数据DT,其中使用基于Word2vec的Fisher Vector方法提取文本特征;具体将DT转换成词向量集合W={W1,W2,…,Wn},W为DT包含的单词的词向量集合;将W={W1,W2,…,Wn}中的每个文本词向量集合Wi代入式1中的X,求得每个文本的Fisher Vector,记作T={T1,T2,…,Tn},Ti∈R(2×dw+1)×G-1,i∈[1,n],其中,Ti表示由第i个文本计算出来的Fisher Vector;由此提取得到文本特征;4)使用执行步骤2)和步骤3)得到的训练数据中的图像特征和文本特征对基于逻辑回归的语义匹配模型进行训练,将文本特征T转换成了文本语义特征ПT,c是类别的个数,也是文本语义特征的维数;将图像特征Ii转换成后验概率组成的语义特征,后验概率为k∈[1,C],表示图像Ii属于类别k的概率;5)利用步骤4)训练好的语义匹配模型,使用步骤2和步骤3得到的测试数据的图像特征和文本特征,针对一幅图片或文本进行测试,得到相关的文本或图片,即为跨媒体检索结果。
- 如权利要求1所述跨媒体检索方法,其特征是,步骤3)使用基于Word2vec的FisherVector方法提取文本特征,具体包括如下过程:32)将单词记作w,单词w所对应的词向量为fword2vec(w);对于有fword2vec(w)∈Wi,i∈[1,n],即其中wi,j∈Rdw,j∈[1,bi],wi,j为包含单词所对应的词向量,dw为词向量的维度,bi为中包含的单词个数;33)用X={x1,x2,…,xnw}表示一个文本的词向量集合,nw为词向量个数;令混合高斯模型GMM的参数为λ,λ={ωi,μi,∑i,i=1..G},其中ωi,μi,∑i分别表示GMM中每个高斯函数的权重、均值向量和协方差矩阵,G表示模型中高斯函数的个数;GMM函数定义为式1:其中,p(xt|λ)表示对于向量xt(t∈[1,nw]),由GMM产生的概率值p(xt|λ),表示为式2:对权重ωi设置总和为1约束,表示为式3:其中,pi(x|λ)表示GMM中的第i个高斯函数,由式4给出:其中,dw是向量的维度,|∑i|表示求∑i的行列式;用γt(i)来表示向量xt由第i个高斯函数产生的概率,表示为式5:34)对高斯模型的参数求偏导即得到Fisher Vector;所述Fisher Vector是将所有参数的求导结果连接组成的向量;所述高斯混合模型中高斯函数个数为G,向量维度为dw,所述Fisher Vector的维度为(2×dw+1)×G-1;权重ω的自由度为G-1;35)将W={W1,W2,…,Wn}中的每个文本词向量集合Wi代入式1中的文本的词向量集合X,求得每个文本的Fisher Vector,记作T={T1,T2,…,Tn},Ti∈R(2×dw+1)×G-1,i∈[1,n],其中,Ti表示由第i个文本计算出来的Fisher Vector。
- 如权利要求1所述跨媒体检索方法,其特征是,步骤4)所述使用训练数据中的图像特征和文本特征对基于逻辑回归的语义匹配模型进行训练,所述图像特征为I={I1,I2,…,In},Ij∈R4096;所述文本特征为T={T1,T2,…,Tn},Ti∈R(2×dw+1)×G-1;相应的图像特征和文本特征具有共同的标签为L=[l1,l2,…,ln],其中li∈[1,2,…,C];所述训练具体包括:
- 如权利要求1所述跨媒体检索方法,其特征是,步骤5)所述针对一幅图片或文本进行测试,得到相关的文本或图片;所述相关性的度量方法包括Kullback–Leibler divergence方法、Normalized Correlation方法、Centered Correlation方法和L2范式方法中的一种或多种。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US16/314,673 US10719664B2 (en) | 2016-07-11 | 2016-12-01 | Cross-media search method |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201610544156.5 | 2016-07-11 | ||
| CN201610544156.5A CN106202413B (zh) | 2016-07-11 | 2016-07-11 | 一种跨媒体检索方法 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2018010365A1 true WO2018010365A1 (zh) | 2018-01-18 |
Family
ID=57476922
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2016/108196 Ceased WO2018010365A1 (zh) | 2016-07-11 | 2016-12-01 | 一种跨媒体检索方法 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US10719664B2 (zh) |
| CN (1) | CN106202413B (zh) |
| WO (1) | WO2018010365A1 (zh) |
Cited By (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108595688A (zh) * | 2018-05-08 | 2018-09-28 | 鲁东大学 | 基于在线学习的潜在语义跨媒体哈希检索方法 |
| TWI695277B (zh) * | 2018-06-29 | 2020-06-01 | 國立臺灣師範大學 | 自動化網站資料蒐集方法 |
| CN111782853A (zh) * | 2020-06-23 | 2020-10-16 | 西安电子科技大学 | 基于注意力机制的语义图像检索方法 |
| CN112364192A (zh) * | 2020-10-13 | 2021-02-12 | 中山大学 | 一种基于集成学习的零样本哈希检索方法 |
| CN112613451A (zh) * | 2020-12-29 | 2021-04-06 | 民生科技有限责任公司 | 一种跨模态文本图片检索模型的建模方法 |
| EP3761187A4 (en) * | 2018-04-13 | 2021-12-15 | Tencent Technology (Shenzhen) Company Limited | METHOD AND DEVICE FOR ALIGNING MULTIMEDIA RESOURCES AND STORAGE MEDIUM AND ELECTRONIC DEVICE |
| CN116245103A (zh) * | 2022-09-07 | 2023-06-09 | 京东科技信息技术有限公司 | 模型训练方法、实体确定方法、装置、电子设备和介质 |
| CN117648579A (zh) * | 2022-08-15 | 2024-03-05 | 腾讯科技(深圳)有限公司 | 数据信息与事件信息的匹配方法、装置和计算机设备 |
| CN120747125A (zh) * | 2024-05-29 | 2025-10-03 | 荣耀终端股份有限公司 | 图像处理方法及电子设备 |
Families Citing this family (43)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106777040A (zh) * | 2016-12-09 | 2017-05-31 | 厦门大学 | 一种基于情感极性感知算法的跨媒体微博舆情分析方法 |
| CN106649715B (zh) * | 2016-12-21 | 2019-08-09 | 中国人民解放军国防科学技术大学 | 一种基于局部敏感哈希算法和神经网络的跨媒体检索方法 |
| CN107256221B (zh) * | 2017-04-26 | 2020-11-03 | 苏州大学 | 基于多特征融合的视频描述方法 |
| CN107016439A (zh) * | 2017-05-09 | 2017-08-04 | 重庆大学 | 基于cr2神经网络的图像‑文本双编码机理实现模型 |
| CN107273517B (zh) * | 2017-06-21 | 2021-07-23 | 复旦大学 | 基于图嵌入学习的图文跨模态检索方法 |
| CN107256271B (zh) * | 2017-06-27 | 2020-04-03 | 鲁东大学 | 基于映射字典学习的跨模态哈希检索方法 |
| CN107330100B (zh) * | 2017-07-06 | 2020-04-03 | 北京大学深圳研究生院 | 基于多视图联合嵌入空间的图像-文本双向检索方法 |
| CN107832351A (zh) * | 2017-10-21 | 2018-03-23 | 桂林电子科技大学 | 基于深度关联网络的跨模态检索方法 |
| CN110020078B (zh) * | 2017-12-01 | 2021-08-20 | 北京搜狗科技发展有限公司 | 一种生成相关性映射字典及其验证相关性的方法和相关装置 |
| CN108319686B (zh) * | 2018-02-01 | 2021-07-30 | 北京大学深圳研究生院 | 基于受限文本空间的对抗性跨媒体检索方法 |
| JP6765392B2 (ja) * | 2018-02-05 | 2020-10-07 | 株式会社デンソーテン | 電源制御装置および電源制御方法 |
| CN108647350A (zh) * | 2018-05-16 | 2018-10-12 | 中国人民解放军陆军工程大学 | 一种基于双通道网络的图文关联检索方法 |
| CN108876643A (zh) * | 2018-05-24 | 2018-11-23 | 北京工业大学 | 一种社交策展网络上采集(Pin)的多模态表示方法 |
| CN108920648B (zh) * | 2018-07-03 | 2021-06-22 | 四川大学 | 一种基于音乐-图像语义关系的跨模态匹配方法 |
| CN109145974B (zh) * | 2018-08-13 | 2022-06-24 | 广东工业大学 | 一种基于图文匹配的多层次图像特征融合方法 |
| CN109344279B (zh) * | 2018-12-12 | 2021-08-10 | 山东山大鸥玛软件股份有限公司 | 基于哈希检索的手写英文单词智能识别方法 |
| US11704487B2 (en) * | 2019-04-04 | 2023-07-18 | Beijing Jingdong Shangke Information Technology Co., Ltd. | System and method for fashion attributes extraction |
| CN110059201A (zh) * | 2019-04-19 | 2019-07-26 | 杭州联汇科技股份有限公司 | 一种基于深度学习的跨媒体节目特征提取方法 |
| US11403700B2 (en) * | 2019-04-23 | 2022-08-02 | Target Brands, Inc. | Link prediction using Hebbian graph embeddings |
| CN112182281B (zh) * | 2019-07-05 | 2023-09-19 | 腾讯科技(深圳)有限公司 | 一种音频推荐方法、装置及存储介质 |
| CN110472079B (zh) * | 2019-07-08 | 2022-04-05 | 杭州未名信科科技有限公司 | 目标图像的检索方法、装置、设备及存储介质 |
| CN110516026A (zh) * | 2019-07-15 | 2019-11-29 | 西安电子科技大学 | 基于图正则化非负矩阵分解的在线单模态哈希检索方法 |
| US11520993B2 (en) * | 2019-07-24 | 2022-12-06 | Nec Corporation | Word-overlap-based clustering cross-modal retrieval |
| US20210027157A1 (en) * | 2019-07-24 | 2021-01-28 | Nec Laboratories America, Inc. | Unsupervised concept discovery and cross-modal retrieval in time series and text comments based on canonical correlation analysis |
| CN110647632B (zh) * | 2019-08-06 | 2020-09-04 | 上海孚典智能科技有限公司 | 基于机器学习的图像与文本映射技术 |
| CN110598739B (zh) * | 2019-08-07 | 2023-06-23 | 广州视源电子科技股份有限公司 | 图文转换方法、设备、智能交互方法、设备及系统、客户端、服务器、机器、介质 |
| CN112528624B (zh) * | 2019-09-03 | 2024-05-14 | 阿里巴巴集团控股有限公司 | 文本处理方法、装置、搜索方法以及处理器 |
| CN110705283A (zh) * | 2019-09-06 | 2020-01-17 | 上海交通大学 | 基于文本法律法规与司法解释匹配的深度学习方法和系统 |
| CN110580281A (zh) * | 2019-09-11 | 2019-12-17 | 江苏鸿信系统集成有限公司 | 一种基于语义相似度的相似案件匹配方法 |
| CN111782921B (zh) * | 2020-03-25 | 2025-01-14 | 北京沃东天骏信息技术有限公司 | 检索目标的方法和装置 |
| CN111753190B (zh) * | 2020-05-29 | 2024-07-05 | 中山大学 | 一种基于元学习的无监督跨模态哈希检索方法 |
| WO2022041940A1 (en) * | 2020-08-31 | 2022-03-03 | Guangdong Oppo Mobile Telecommunications Corp., Ltd. | Cross-modal retrieval method, training method for cross-modal retrieval model, and related device |
| CN112037215B (zh) * | 2020-09-09 | 2024-05-28 | 华北电力大学(保定) | 一种基于零样本学习的绝缘子缺陷检测方法及系统 |
| CN112559820B (zh) * | 2020-12-17 | 2022-08-30 | 中国科学院空天信息创新研究院 | 基于深度学习的样本数据集智能出题方法、装置及设备 |
| CN112818113A (zh) * | 2021-01-26 | 2021-05-18 | 山西三友和智慧信息技术股份有限公司 | 一种基于异构图网络的文本自动摘要方法 |
| CN113254678B (zh) * | 2021-07-14 | 2021-10-01 | 北京邮电大学 | 跨媒体检索模型的训练方法、跨媒体检索方法及其设备 |
| CN114048282A (zh) * | 2021-11-16 | 2022-02-15 | 中山大学 | 一种基于文本树局部匹配的图文跨模态检索方法及系统 |
| CN114328506B (zh) * | 2021-11-19 | 2024-07-02 | 集美大学 | 一种智能船舶自动控制系统 |
| CN116775980B (zh) | 2022-03-07 | 2024-06-07 | 腾讯科技(深圳)有限公司 | 一种跨模态搜索方法及相关设备 |
| CN115062208B (zh) * | 2022-05-30 | 2024-01-23 | 苏州浪潮智能科技有限公司 | 数据处理方法、系统及计算机设备 |
| CN116127171A (zh) * | 2022-11-29 | 2023-05-16 | 北京海卓飞网络科技有限公司 | 企业信息采集及展示方法、装置、电子设备及存储介质 |
| CN116881482A (zh) * | 2023-06-27 | 2023-10-13 | 四川九洲视讯科技有限责任公司 | 一种公共安全数据的跨媒体智能感知与分析处理方法 |
| CN117812381B (zh) * | 2023-12-05 | 2024-06-04 | 世优(北京)科技有限公司 | 基于人工智能的视频内容制作方法 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104166684A (zh) * | 2014-07-24 | 2014-11-26 | 北京大学 | 一种基于统一稀疏表示的跨媒体检索方法 |
| CN104317834A (zh) * | 2014-10-10 | 2015-01-28 | 浙江大学 | 一种基于深度神经网络的跨媒体排序方法 |
| CN104346440A (zh) * | 2014-10-10 | 2015-02-11 | 浙江大学 | 一种基于神经网络的跨媒体哈希索引方法 |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5848378A (en) * | 1996-02-07 | 1998-12-08 | The International Weather Network | System for collecting and presenting real-time weather information on multiple media |
| US7814040B1 (en) * | 2006-01-31 | 2010-10-12 | The Research Foundation Of State University Of New York | System and method for image annotation and multi-modal image retrieval using probabilistic semantic models |
| CN101901249A (zh) * | 2009-05-26 | 2010-12-01 | 复旦大学 | 一种图像检索中基于文本的查询扩展与排序方法 |
| CN102629275B (zh) * | 2012-03-21 | 2014-04-02 | 复旦大学 | 面向跨媒体新闻检索的人脸-人名对齐方法及系统 |
-
2016
- 2016-07-11 CN CN201610544156.5A patent/CN106202413B/zh not_active Expired - Fee Related
- 2016-12-01 US US16/314,673 patent/US10719664B2/en not_active Expired - Fee Related
- 2016-12-01 WO PCT/CN2016/108196 patent/WO2018010365A1/zh not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104166684A (zh) * | 2014-07-24 | 2014-11-26 | 北京大学 | 一种基于统一稀疏表示的跨媒体检索方法 |
| CN104317834A (zh) * | 2014-10-10 | 2015-01-28 | 浙江大学 | 一种基于深度神经网络的跨媒体排序方法 |
| CN104346440A (zh) * | 2014-10-10 | 2015-02-11 | 浙江大学 | 一种基于神经网络的跨媒体哈希索引方法 |
Cited By (11)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP3761187A4 (en) * | 2018-04-13 | 2021-12-15 | Tencent Technology (Shenzhen) Company Limited | METHOD AND DEVICE FOR ALIGNING MULTIMEDIA RESOURCES AND STORAGE MEDIUM AND ELECTRONIC DEVICE |
| US11914639B2 (en) | 2018-04-13 | 2024-02-27 | Tencent Technology (Shenzhen) Company Limited | Multimedia resource matching method and apparatus, storage medium, and electronic apparatus |
| CN108595688A (zh) * | 2018-05-08 | 2018-09-28 | 鲁东大学 | 基于在线学习的潜在语义跨媒体哈希检索方法 |
| TWI695277B (zh) * | 2018-06-29 | 2020-06-01 | 國立臺灣師範大學 | 自動化網站資料蒐集方法 |
| CN111782853A (zh) * | 2020-06-23 | 2020-10-16 | 西安电子科技大学 | 基于注意力机制的语义图像检索方法 |
| CN111782853B (zh) * | 2020-06-23 | 2022-12-02 | 西安电子科技大学 | 基于注意力机制的语义图像检索方法 |
| CN112364192A (zh) * | 2020-10-13 | 2021-02-12 | 中山大学 | 一种基于集成学习的零样本哈希检索方法 |
| CN112613451A (zh) * | 2020-12-29 | 2021-04-06 | 民生科技有限责任公司 | 一种跨模态文本图片检索模型的建模方法 |
| CN117648579A (zh) * | 2022-08-15 | 2024-03-05 | 腾讯科技(深圳)有限公司 | 数据信息与事件信息的匹配方法、装置和计算机设备 |
| CN116245103A (zh) * | 2022-09-07 | 2023-06-09 | 京东科技信息技术有限公司 | 模型训练方法、实体确定方法、装置、电子设备和介质 |
| CN120747125A (zh) * | 2024-05-29 | 2025-10-03 | 荣耀终端股份有限公司 | 图像处理方法及电子设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN106202413B (zh) | 2018-11-20 |
| US20190205393A1 (en) | 2019-07-04 |
| CN106202413A (zh) | 2016-12-07 |
| US10719664B2 (en) | 2020-07-21 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2018010365A1 (zh) | 一种跨媒体检索方法 | |
| Pang et al. | Text matching as image recognition | |
| CN106997382B (zh) | 基于大数据的创新创意标签自动标注方法及系统 | |
| US10430689B2 (en) | Training a classifier algorithm used for automatically generating tags to be applied to images | |
| Nowak et al. | The CLEF 2011 Photo Annotation and Concept-based Retrieval Tasks. | |
| CN109960763B (zh) | 基于用户细粒度摄影偏好的摄影社区个性化好友推荐方法 | |
| CN107315772B (zh) | 基于深度学习的问题匹配方法以及装置 | |
| CN106095829A (zh) | 基于深度学习与一致性表达空间学习的跨媒体检索方法 | |
| CN113934835B (zh) | 结合关键词和语义理解表征的检索式回复对话方法及系统 | |
| CN108804595B (zh) | 一种基于word2vec的短文本表示方法 | |
| CN108038099B (zh) | 基于词聚类的低频关键词识别方法 | |
| CN108763348A (zh) | 一种扩展短文本词特征向量的分类改进方法 | |
| CN110489548A (zh) | 一种基于语义、时间和社交关系的中文微博话题检测方法及系统 | |
| CN112800249A (zh) | 基于生成对抗网络的细粒度跨媒体检索方法 | |
| Yao et al. | A new web-supervised method for image dataset constructions | |
| CN103559193A (zh) | 一种基于选择单元的主题建模方法 | |
| CN109033478A (zh) | 一种用于搜索引擎的文本信息规律分析方法与系统 | |
| CN110442736A (zh) | 一种基于二次判别分析的语义增强子空间跨媒体检索方法 | |
| CN104077419B (zh) | 结合语义与视觉信息的长查询图像检索重排序方法 | |
| Peng et al. | The effect of pets on happiness: A large-scale multi-factor analysis using social multimedia | |
| CN105989094A (zh) | 基于隐层语义中层表达的图像检索方法 | |
| TW202004519A (zh) | 影像自動分類的方法 | |
| CN113868424B (zh) | 文本主题的确定方法、装置、计算机设备及存储介质 | |
| CN103198117B (zh) | 基于内容的图像伪相关重排序方法 | |
| CN113742520A (zh) | 基于半监督学习的密集视频描述算法的视频查询检索方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 16908688 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 16908688 Country of ref document: EP Kind code of ref document: A1 |










