WO2020194792A1 - 検索装置、学習装置、検索方法、学習方法及びプログラム - Google Patents
検索装置、学習装置、検索方法、学習方法及びプログラム Download PDFInfo
- Publication number
- WO2020194792A1 WO2020194792A1 PCT/JP2019/035526 JP2019035526W WO2020194792A1 WO 2020194792 A1 WO2020194792 A1 WO 2020194792A1 JP 2019035526 W JP2019035526 W JP 2019035526W WO 2020194792 A1 WO2020194792 A1 WO 2020194792A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- region
- neural network
- learning
- feature vector
- media data
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/40—Information retrieval; Database structures therefor; File system structures therefor of multimedia data, e.g. slideshows comprising image and additional audio data
- G06F16/43—Querying
- G06F16/432—Query formulation
- G06F16/434—Query formulation using image data, e.g. images, photos, pictures taken by a user
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/40—Information retrieval; Database structures therefor; File system structures therefor of multimedia data, e.g. slideshows comprising image and additional audio data
- G06F16/43—Querying
- G06F16/432—Query formulation
- G06F16/433—Query formulation using audio data
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/50—Information retrieval; Database structures therefor; File system structures therefor of still image data
- G06F16/53—Querying
- G06F16/532—Query formulation, e.g. graphical querying
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/73—Querying
- G06F16/732—Query formulation
- G06F16/7335—Graphical querying, e.g. query-by-region, query-by-sketch, query-by-trajectory, GUIs for designating a person/face/object as a query predicate
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/21—Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
- G06F18/214—Generating training patterns; Bootstrap methods, e.g. bagging or boosting
- G06F18/2148—Generating training patterns; Bootstrap methods, e.g. bagging or boosting characterised by the process organisation or structure, e.g. boosting cascade
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/044—Recurrent networks, e.g. Hopfield networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/044—Recurrent networks, e.g. Hopfield networks
- G06N3/0442—Recurrent networks, e.g. Hopfield networks characterised by memory or gating, e.g. long short-term memory [LSTM] or gated recurrent units [GRU]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/084—Backpropagation, e.g. using gradient descent
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/092—Reinforcement learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N7/00—Computing arrangements based on specific mathematical models
- G06N7/01—Probabilistic graphical models, e.g. probabilistic networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
Definitions
- the present invention relates to a search device, a learning device, a search method, a learning method, and a program that can be used to search media data for a target area that matches the query data.
- the preferred way to solve this problem is robust enough to find exact matches in situations that occur in real-world scenarios such as background clutter, occlusion and geometric transformations. It is desirable to be a method. In addition, it is desired that the method is sufficiently fast so that the query image can be identified from a large number of candidate regions in the reference image within a reasonable time.
- Non-Patent Document 1 reduces the execution time by adaptively sliding the window more than one pixel. Judgment about the amount of slides is based on the rank defined for each feature in the pattern.
- Non-Patent Document 2 proposes a method for ensuring robustness and high speed against a complex background by using a subset of pixel pairs (usually a small subset).
- Non-Patent Document 3 speeds up the search process by combining geometric transformation with random sampling of a distance approximation method and a branch-and-bound search.
- Non-Patent Document 4 reduces the calculation cost by skipping the positions that do not match based on the difference features in the principal component direction (principal orientation differences features).
- Non-Patent Documents 1 to 4 above still requires evaluation of a large number of windows or images in order to identify a target area that matches the query image, which is the whole.
- the search process is inefficient.
- the above method does not consider optimizing the search path (that is, the order of windows or pixels to be evaluated) for specifying the target area while specifying the target area. Therefore, the above method is not robust for complex backgrounds, occlusions, geometric transformations, etc., for example.
- Reducing the number of candidate regions such as windows or pixels to be evaluated is desired not only for image search applications but also for other applications that handle media data such as moving images and acoustic signals.
- the present invention has been made in view of the above problems, and the present invention reduces the number of candidate regions to be evaluated in the media data and more efficiently finds a target region matching the query data from the media data. It is an object of the present invention to provide a search device, a learning device, a search method, a learning method, and a program capable of the present invention.
- One embodiment of the present invention provides a search device that searches media data for a target area that matches the query data.
- a first feature extraction unit that extracts a first feature vector from the query data using the first trained neural network
- a second feature extraction unit that acquires a first region from the media data and extracts a second feature vector from the first region using a second trained neural network.
- Candidates for the target region based on the first feature vector, the second feature vector, and the position of the first region or the first region using a third trained neural network.
- the position identification part that determines By using the determined target region candidate as the first region used by the second feature extraction unit, the second feature extraction unit and the position identification unit are specified until a predetermined condition is satisfied.
- a control unit that repeats the operation of the unit and Have.
- Another embodiment of the present invention provides a learning device that learns a neural network used to search media data for a target area that matches the query data.
- a first feature extraction unit that extracts a first feature vector from learning query data using a first neural network
- a second feature extraction unit that acquires a first region from the learning media data and extracts a second feature vector from the first region using a second neural network.
- a candidate for the target region is determined based on the first feature vector, the second feature vector, and the position of the first region or the first region. Positioning part to do and By using the determined target region candidate as the first region used by the second feature extraction unit, the second feature extraction unit and the position identification unit are specified until a predetermined condition is satisfied.
- a control unit that repeats the operation of the unit and It is determined whether or not the candidate of the target area determined by the position specifying unit has captured the target area, and based on the determination result, the first neural network, the second neural network, and the first neural network are used.
- a learning unit that updates the parameters of the neural network of 3 and Have.
- Another embodiment of the present invention provides a search method used by a search device that searches media data for a target area that matches the query data.
- the first step of extracting the first feature vector from the query data using the first trained neural network and A second step of acquiring a first region from the media data and extracting a second feature vector from the first region using a second trained neural network.
- Candidates for the target region based on the first feature vector, the second feature vector, and the position of the first region or the first region using a third trained neural network.
- the third step to determine and By using the determined target region candidate as the first region used in the second step, the second step and the third step are repeated until a predetermined condition is satisfied. 4th step and Have.
- Another embodiment of the present invention provides a learning method used by a learning device that learns a neural network used to search media data for a target area that matches the query data.
- the first step of extracting the first feature vector from the training query data using the first neural network A second step of acquiring a first region from the learning media data and extracting a second feature vector from the first region using a second neural network, Using a third neural network, a candidate for the target region is determined based on the first feature vector, the second feature vector, and the position of the first region or the first region.
- the third step to do and By using the determined target region candidate as the first region used in the second step, the second step and the third step are repeated until a predetermined condition is satisfied. 4th step and It is determined whether or not the candidate of the target area has captured the target area, and the parameters of the first neural network, the second neural network, and the third neural network are updated based on the determination result.
- Another aspect of the present invention provides a program that causes a computer to function as the above-mentioned search device or learning device.
- the present invention it is possible to reduce the number of candidate regions to be evaluated in the media data and more efficiently find a target region that matches the query data from the media data.
- a search device that searches media data for a target area (hereinafter, also referred to as a “target window”) that matches the query data
- media data and query data are still images, moving images, acoustic signals or other data, respectively.
- the search device is a model trained by the learning device in order to find a target area matching the query data from a plurality of target area candidates (hereinafter, also referred to as “candidate area” or “candidate window”) in the media data.
- a neural network is used.
- the search device finds the target area by repeatedly comparing the query data and the candidate area by using the trained model and determining whether or not the candidate area matches the query data. Since the model is trained to reduce the candidate areas to be evaluated, the search device can find the target area in a smaller number of candidate areas.
- the learning device for learning the model used by the search device will be further described.
- the learning device may be different from the search device and may be the same as the search device. In the following description, it is assumed that the learning device is different from the search device.
- FIG. 1 is a diagram showing a functional configuration of the search device 100 according to the first embodiment of the present invention.
- the purpose of the search device 100 is to identify a target area represented by the query data Q, which exists at a specific position l g in the media data R.
- the position l g may be the center of the area, the corners of the area, or any other position that can be used to define the area.
- the position l g can be represented by xy image coordinates.
- the position l g can be represented by a time stamp (or frame index). In this case, the position l g may include the xy image coordinates of the time frame.
- the position l g can be represented by a time stamp.
- the search device 100 may determine other types of information used to define the target area, such as window size, orientation, or a combination of any of the above.
- the search device 100 includes a first feature extraction unit 120, a second feature extraction unit 130, a position identification unit 140, and a control unit 150.
- the search device 100 may further include an initial region prediction unit 110.
- the initial region prediction unit 110 is a neural network that inputs media data R and outputs the initial region R (l 0 ) to be evaluated or its position l 0 .
- the neural network may be a convolutional neural network (CNN) or another neural network.
- CNN convolutional neural network
- RNN recurrent neural network
- LSTM long-short term memory
- Initial region R (l 0) is the region that is extracted from the media data at the initial position l 0, the initial position l 0 is acquired by linear projection feature amount of media data R coarse after the down-sampling the position vector You may.
- the initial region R (l 0 ) is the first candidate region evaluated by the second feature extraction unit 130.
- the initial area prediction unit 110 may not be included in the search device 100, and the initial position l 0 may be arbitrarily determined.
- the first feature extraction unit 120 is a neural network that inputs query data Q and outputs feature vector f (Q) of query data Q.
- the neural network may be CNN or another neural network.
- RNN or LSTM may be used.
- the first feature extraction unit 120 extracts the feature vector f (Q) from the query data Q.
- This is a neural network that outputs the feature vector f (R (l t )) of the candidate region of the media data R.
- the second feature extraction unit 130 extracts the candidate region R (l t ) at the position l t from the media data R.
- the neural network of the second feature extraction unit 130 is the same as the neural network of the first feature extraction unit 120, and shares the same parameters as the neural network of the first feature extraction unit 120.
- Second feature extraction unit 130 acquires the candidate region R (l t), extracts a candidate region R (l t) from the feature vector f (R (l t)) .
- the neural network of the position specifying unit 140 may be an LSTM or another neural network.
- the position specifying unit 140 is based on the feature vector f (Q), the feature vector f (R (l t )), the candidate region R (l t ), or its position l t, and the next candidate region R (l). t + 1 ) is determined. More specifically, the position specifying unit 140 combines the feature vector f (Q) and the feature vector f (R (l t )) into a single vector, and then the combined vector and the current candidate. The next candidate position R (l t + 1 ) or its position l based on the region R (l t ) or its position l t and the current internal state (also called the "hidden state") of the position identifying part 140. Determine t + 1 .
- the position specifying unit 140 outputs the candidate area R (l T ) or its position l T. Further, when the candidate area R (l t ) matches the query data Q, the position specifying unit 140 determines that the candidate area R (l T ) is the target area matching the query data Q, and determines the target area or its target area. Output the position.
- the control unit 150 is a processing unit that receives the next candidate area R (l t + 1 ) or its position l t + 1 or the final result as an input and determines whether or not to end the repetition.
- the control unit 150 inputs the next candidate region R (l t + 1 ) or its position l t + 1 to the second feature extraction unit 130, and the second feature extraction unit 130 until a predetermined condition is satisfied.
- the operation of the position specifying unit 140 is repeated. For example, the control unit 150 increases the number of repetitions by 1 for each repetition, and ends the repetition when the number of repetitions t reaches a predetermined limit value T. Further, for example, the control unit 150, the position specifying unit 140, if the candidate region R (l t)) is determined to be the target area that matches the query data, may end the repetition.
- the neural network of the first feature extraction unit 120, the second feature extraction unit 130, and the position identification unit 140 (and the initial region prediction unit 110 if the initial region prediction unit 110 is included) , Learning query data and learning media data are used for learning.
- the learning query data is input to the first feature extraction unit 120
- the learning media data is input to the second feature extraction unit 130 (or the initial area prediction unit 110 if the initial area prediction unit 110 is included).
- the parameters of the neural network are learned based on the result of determining whether or not the candidate region R (l t + 1 ) determined by the position specifying unit 140 has captured the target region.
- the neural network has a degree of similarity between the feature vector f (Q) extracted by the first feature extraction unit 120 and the feature vector f (R (l t )) extracted by the second feature extraction unit 130. Is learned to grow.
- FIG. 2 is a diagram showing a functional configuration of the learning device 200 according to the first embodiment of the present invention.
- the learning device 200 includes a first feature extraction unit 220, a second feature extraction unit 230, a position identification unit 240, a control unit 250, and a learning unit 260.
- the learning device 200 may further include an initial region prediction unit 210.
- the initial region prediction unit 210, the first feature extraction unit 220, the second feature extraction unit 230, the position identification unit 240, and the control unit 250 are the initial region prediction unit 110, the first feature extraction unit 120, in the search device 100. It is the same as the second feature extraction unit 130, the position identification unit 140, and the control unit 150, respectively.
- the learning device 200 uses learning data (also referred to as a "query reference pair") including learning media data R and learning query data Q as inputs.
- the learning query data Q may be a part of the learning media data R, or may be data similar to a part of the learning media data R.
- the exact position of the learning query data Q in the learning media data R does not necessarily have to be given.
- the feature vector f (Q) can be acquired by the first feature extraction unit 220, and the feature vector f ((R (l t ))) is the second. Can be acquired by the feature extraction unit 230, and the next candidate region R (l t + 1 ) or its position l t + 1 , or the final result can be acquired by the position identification unit 240.
- the learning unit 260 inputs the feature vector f (Q), the feature vector f ((R (l t )), the next candidate region R (l t + 1 ) or its position l t + 1 , or the final result.
- a processing unit that outputs neural network parameters of the first feature extraction unit 220, the second feature extraction unit 230, and the position identification unit 240 (and the initial region prediction unit 210 if the initial region prediction unit 210 is included).
- the learning unit 260 updates the parameters of the neural network based on the result of determining whether or not the candidate area R (l t + 1 ) determined by the position specifying unit 240 has captured the target area.
- the learning unit 260 calculates the similarity between the feature vector f (Q) and the feature vector f ((R (l t )), and sets the feature vector f (Q) and the feature vector f ((R (l t ))).
- the media data is the reference image R and the query data is the query image Q.
- FIG. 3 is a conceptual diagram of the search device 100 according to the second embodiment.
- the reference image R is downsampled to the low-resolution reference image R coarse , the feature vector f (R coarse ) is extracted from the low-resolution reference image R coarse by CNN, and the feature vector f (R) is extracted. coarse ) is linearly projected at position l 0 .
- the feature vector f (Q) representing the image feature amount is extracted from the query image Q by CNN.
- the feature vector f (R (l 0 )) representing the image feature amount is extracted from the reference image R at the position l 0 by the CNN.
- the current position l 0 is linearly projected onto the position vector having the same dimension as the feature vector f (R coarse ) or f (R (l 0 )), and the next hidden state is the feature vector f.
- the RSTM based on the combination of (R coarse ) and f (R (l 0 )), the current hidden state of the position identification part 140, and the position vector, the next hidden state is at the next position l 1 . It is linearly projected.
- the control unit 150 (not shown in FIG. 3) of the second feature extraction unit 130 and the position identification unit 140 until the number of repetitions reaches a predetermined limit value T. Repeat the operation. When the limit value is reached, the image search process ends and the image search result is output.
- Steps S101 to S103 are related to the initialization step, and in the initialization step, the reference image R is input and the initial position l 0 is output.
- step S101 the initial region prediction unit 110 downsamples the reference image R to a low resolution reference image R coarse . Downsampling of the reference image R may be performed using a scaling factor of 3.
- Step 102 will be described in detail with reference to FIG.
- FIG. 5 is a detailed view of the CNN in the initial region prediction unit 110.
- the initial region prediction unit 110 uses three convolution layers that map the downsampled reference image R coarse to the feature vector f (R coarse ). Since the feature vector f (R coarse ) is obtained from the downsampled reference image R coarse , it effectively gives an indication of where the potentially interesting region is in a given reference image R. More specifically, the first convolution layer takes R coarse as an input and applies a maximum pooling layer of stride 2 following 32 2D convolution filters of size 7x7. Each of the second and third convolution layers consists of 32 2D convolution filters of the same size 3x3 followed by a maximum pooling layer of the same stride 2. Finally, there is a fully-connected (FC) layer that receives the output of the third convolution layer and produces a fixed-length feature vector of length 256.
- FC fully-connected
- step S103 linear projection is applied to the obtained feature vector f (R coarse) in step S102, the conversion length 256 feature vectors f a (R coarse) to the position vector of length 2.
- the position vector is normalized to the range -1 to 1.
- Steps S104 to S106 relate to the feature extraction step.
- step S104 the image region is extracted from the reference image R at position l 0 represented by the position vector using the position vector of step S103.
- FIG. 6 is a detailed view of the CNN in the first feature extraction unit 120 and the second feature extraction unit 130. Since the position l t is repeatedly determined by the control of the control unit 150, the number of repetitions is generally indicated as t in the following description. When the number of repetitions is 0, the initial position l 0 is used.
- the first feature extraction unit 120 and the second feature extraction unit 130 map the query image Q and the extracted image region R (l t ) to the feature vectors f (Q) and f (R (l t )), respectively.
- Conv-ReLU convolutional-rectified linear unit
- GAP global average pooling
- each convolution layer The specifications of each convolution layer are as follows: 1st layer: 32 2D convolution filters (filter: 7x7 and stride: 1x1), 2nd layer: 64 2D convolution filters (filter: 5x5 and stride: 1x1), 3rd layer: 128 2D convolution filters (filter: 3x3 and stride: 1x1), 4th layer: 256 2D convolution filters (filter: 1x1 and stride: 1x) 1), 5th layer: 128 2D convolution filters (filter: 1 ⁇ 1 and stride 1 ⁇ 1).
- step S107 the first feature extraction unit 120 and the second feature acquired from the extraction unit 130 fixed length feature vector f (Q) and f (R (l t)) is of length 256 of a single Combined into a vector.
- Step S108 ⁇ S110 relates localization step, the position specifying unit 140 includes the image feature vector f (Q) and f (R (l t)) , and the current position l t, the current state h t of LSTM Based on the three inputs, the LSTM predicts the next position l t + 1 in sequence.
- step S108 the current position l t is first encoded by linear projection into a position vector having the same dimensions as the feature vector, and then processed in combination with the feature vector.
- the combined vector, the current position l t, and the current hidden state h t are combined to form a single vector that is input to the LSTM.
- the output of the positioning unit 140 is a fixed-length vector (256) of the next hidden state h t + 1 .
- step S110 the next hidden state h t + 1 as a result of the LSTM is the expected value of the predicted position of the next region by linear projection.
- step S111 the control unit 150 increases the number of repetitions t by 1, and repeats steps S104 and S106 to 110 until the maximum number of repetitions T is reached.
- the position l t is output.
- the maximum number of iterations T may be fixed at 6. Also, T may be determined adaptively.
- the decision process is modeled as a partially observable Markov decision process (POMDP), and the learning can be performed by the reinforcement learning method.
- POMDP partially observable Markov decision process
- the policy gradient method is used to learn the neural network.
- the policy is to decide how to select the next position to match.
- the next position is the average value
- the learning process starts with inputting learning data (query reference pair) into the initial area prediction unit 210.
- the initial region prediction unit 210 outputs the initial position l 0 as described in steps S101 to S103.
- steps S204 to S206 the first feature extraction unit 220 and the second feature extraction unit 230 extract the feature vector as described in steps S204 to S206.
- steps S207-S210 a random stochastic process (Gaussian distribution) is applied to the output of the positioning unit 240 to produce a result (ie, the next predicted position l t ).
- a random stochastic process Gaussian distribution
- step S211 the learning unit 260 calculates the reward based on the accuracy of the predicted position l t . Further, in step S212, the learning unit 260 calculates the degree of similarity between the feature vector f (Q) and the feature vector f (R (l t )).
- step S213 the learning unit 260 starts backpropagation to update the parameters of the neural network using the calculated reward and similarity.
- the backpropagation bypasses the stochastic process of step S210 and updates the neural network parameters of the position specifying unit 240, the second feature extraction unit 230, the first feature extraction unit 220, and the initial region prediction unit 210.
- the backpropagation of the neural network is considered to update the parameters of the neural network so that more rewards will be given in the future.
- this learning method can learn together the feature amount of the image and the search path (that is, the order of the positions of the search targets) in a unified framework.
- the learning method of the second embodiment is customized for an image search task and is designed to effectively learn image features for a similar search.
- the entire model is made to be unsupervised, i.e., unlike the model above, no class label for learning is required.
- the learning method will be explained in more detail below.
- ⁇ ⁇ f , ⁇ l ⁇ is a set parameter of the entire model
- ⁇ f is a set parameter of the first feature extraction unit 220 and the second feature extraction unit 230
- ⁇ l is the position identification. It is assumed that it is a set of parameters of part 240.
- ⁇ ⁇ f , ⁇ l , ⁇ i ⁇ may be defined
- ⁇ i is a set of parameters of the initial region prediction unit 210. Reinforcement learning is used to adjust ⁇ .
- l t is determined on condition of all past positions.
- the strategy of this model can be expressed as a conditional distribution ⁇ (l t
- r t is a reward function.
- the reward function r t can be based on the success or failure of the search at the number of iterations t.
- Whether or not the window accurately captures the position l g may be determined based on the IoU (intersection over union) between the window and the region R (l g ).
- the expected value of the overall reward is given by the following equation (1).
- the gradient of the expected reward at the number of repetitions t can be defined by the following equation (3).
- equation (3) By exchanging the addition and the gradient in the equation (3) and multiplying and dividing by the policy, the equation (3) can be rewritten as the following equation (4).
- the left side of equation (6) means that the parameter should be updated in the direction of the gradient of the reward function at the current position l t in order to increase the reward in the future. , It is the same as the direction to maximize the likelihood log ⁇ (l t
- ⁇ can be repeatedly updated in the direction of increasing gradient.
- the feature vectors f (Q) and f (R (l t )) of the images extracted from the first feature extraction unit 220 and the second feature extraction unit 230 are f (Q) and f, respectively. It is used to measure the similarity with (R (l t )). For matching pairs, the distance between the two feature vectors should be small, and for unmatched pairs, the distance between the two feature vectors should be large.
- the true label can be inferred directly from the reward given to each pair during the training of the neural network. If the reward is 1, the pair is treated as a match, otherwise it is treated as a mismatched pair. Therefore, the following loss functions (8) from the first feature extraction unit 220 and the second feature extraction unit 230, similar to the contrasting loss function of the widely used Siamese network (Siamese network). Is incorporated.
- the first dataset is called "Translated MNIST" and each reference image is by placing a 28x28 number image (28x28 pixel image) at random positions on a 100x100 blank image. Generated. More specifically, the position coordinates are random numbers in the range from 0 to the size difference between the blank image and the numerical image. This range is set in order to avoid arranging a numerical image at the boundary position of the blank image.
- the second dataset is called "Cluttered MNIST" and is used to evaluate the robustness of image search methods for complex backgrounds.
- each reference image was generated by adding a random 9 ⁇ 9 sub-patch from another random number image to a random position in the Translated MNIST reference image. Specifically, first, an image of 28 ⁇ 28 numbers is randomly selected, then a partial image of 9 ⁇ 9 pixels of another number image is cut out at a random position, and finally, 9 ⁇ A partial image of 9 pixels was embedded at a randomly selected position in the 100 x 100 Translated MNIST reference image. The partial image was embedded so that it did not overlap the existing numeric image. The clutter was controlled by fixing the total number of partial images to be inserted.
- the third dataset is called "Mixed MNIST”. Randomly selected 28x28 number images different from the target number image were placed at random positions in each Cluttered MNIST reference image. Similar to Translated MNIST, the position coordinates were chosen to avoid boundary positions and not overlap with existing numeric images.
- each 100x100 image as a reference image, select a 28x28 query image with the same number as the target number image from the clean image master set centered on 10 numbers 0-9.
- all query reference pairs were prepared for the above three types of MNIST datasets. According to MNIST standard intent, 10,000 query reference pairs were used for testing and 60,000 query reference pairs were used for learning.
- Each pair was generated by considering the logo image in the dataset as a reference image and the just-cut logo with the same brand name as the query image as the query image. A total of 32 query images were generated for each logo. All reference images were resized to half their original size, and each query image was resized to the same size as the reference image logo.
- the image search method was evaluated from the viewpoint of accuracy and speed.
- a query image and a reference image were given, a prediction window corresponding to the area in the reference image was output, and this prediction window was used to evaluate the accuracy.
- the image search is considered to be successful when the IOU (intersection over union) between the prediction window and the true value window is greater than 0.5 according to the same criteria as the object detection method.
- the success rate is the ratio of the number of exact matching image pairs to all pairs.
- the efficiency aspect was evaluated from the viewpoint of two indicators.
- One is the number of windows evaluated and the other is the execution time required to process each query reference pair.
- the execution time was determined by averaging the total time required to match the query reference pairs in the test set.
- the image search method according to the example was evaluated in comparison with two existing image search methods, BBS (Non-Patent Document 2) and MTM (Non-Patent Document 8).
- BBS Non-Patent Document 2
- MTM Non-Patent Document 8
- the model for the example was trained from the beginning using Adam with a batch size of 64 for MNIST and 1 for FlickrLogos-32.
- the learning rate was kept in the range [10 -4 , 10 -3 ] by exponential decay.
- the results of the MNIST dataset were repeated by 3 epochs and the results of FlickrLogos-32 were repeated by 45 epochs.
- the hyperparameter ⁇ of the Gaussian distribution used to sample the next position was fixed at 0.22.
- the margin of the hyperparameter of the contrast loss function was fixed at 0.2. If you set the margin too high, the network will only consider “mismatch”, and if you set the margin too low, the network will learn nothing about "mismatch", so there is a trade-off in margin selection. Exists. Appropriate margins were manually determined in this evaluation.
- the neural network is approximated only by using the gradient stop function for passing the gradient directly to the neural network.
- a loss function was used for training.
- the parameters of the neural network have been updated so that the desired average value is output from the neural network.
- FIGS. 8-10 For all datasets, the success rate, the number of windows evaluated and the execution time are shown in FIGS. 8-10, respectively.
- “Joint-Training” is the result of an embodiment of the present invention.
- a search path for identifying a query image according to an embodiment of the present invention is shown in FIG.
- the success rate of the method of the example is the highest among all the image search methods.
- the maximum gain of the embodiment is as high as 0.25 for BBS and 0.27 for MTM. This result clearly shows that the examples can learn the search path very accurately in the Translated MNIST dataset.
- FIG. 9 the examples are clearly superior to BBS and MTM in terms of the total number of candidate windows evaluated to identify the query image.
- the method of the embodiment evaluates only six candidate windows, while the other method evaluates thousands of windows.
- the advantage of processing only a few windows is reflected in the execution time.
- the execution time is not directly proportional to the number of windows processed, as each method has different computational requirements for each pixel. Nevertheless, the method of the example is as fast as or better than the other two methods, while having very good matching accuracy as shown in FIG.
- the success rate of the method of the example is the highest among all the image search methods. This indicates that the method of the embodiment succeeds in learning the search path even when there is a clutter with a wide range of complex backgrounds.
- the success rate of BBS is the lowest of the three methods. This is because the match between the query image and the candidate window in the BBS is evaluated according to the consistency of the pixel distribution, and when the pixel distribution in the (x, y, R, G, B) space is similar. This is because it is determined that the two windows are a matching pair.
- This method is not effective in Cluttered MNIST and Mixed MNIST where noise can have the same distribution as the target image.
- the method of the embodiment can accurately identify the query image for both Cluttered MNIST and Mixed MNIST with only eight candidate windows. Further, the method of the embodiment is as good as or better than that in terms of execution time as shown in FIG.
- FIG. 8 shows that the method of the embodiment is superior to all other methods in terms of accuracy.
- execution time as shown in FIG. 10
- the gain of the method of the embodiment is larger than that of the MNIST dataset. This is because the size of the reference image is larger than that of MNIST, and the execution time of BBS and MTM is almost linear in the size of the reference image.
- the execution time of the method of the embodiment depends only on the number of windows to be evaluated and is considerably smaller than the two reference methods, as shown in FIG. This shows that the method of the embodiment is more efficient when applied to more realistic and larger size images.
- FIG. 11 shows a search path for identifying a query image. It can be seen that the method of the embodiment has an excellent ability to learn the search path.
- the results in the MNIST dataset show that query images can be successfully identified with approximately the same number of evaluated candidate windows, even if the search level is difficult due to clutter. This is because the method of the embodiment learns the search path for matching and the effective features together.
- the embodiments of the present invention address the problem of matching query images to regions within the reference image.
- the method is based on a neural network (eg, a combination of CNN and LSTM) that sequentially outputs the next position towards the target region at each iteration count.
- the embodiments of the present invention use the positioning section to determine where in the reference image the next region is to be extracted.
- the embodiments of the present invention incorporate a technique based on reinforcement learning for predicting the next position. Therefore, it is possible to focus on the relevant area of the reference image, significantly reducing the number of windows (candidate areas) required to identify the query image, and faster, especially for large reference images. Brings image search.
- the number of candidate windows processed to identify the query image can be determined because the search path and valid features can be learned together based on the similarity between the query image and the reference image. It can be significantly smaller than existing methods, resulting in faster image retrieval, especially for large reference images. Second, as can be seen from the evaluation results, the query image can be accurately identified even for the reference image having a clutter with a considerably complicated background.
- FIG. 12 shows a hardware configuration example of each device (search device 100 or learning device 200) according to the embodiment of the present invention.
- Each device may be a computer composed of a processor such as a CPU (Central Processing Unit) 151, a memory device 152 such as a RAM (Random Access Memory) or a ROM (Read Only Memory), and a storage device 153 such as a hard disk.
- a processor such as a CPU (Central Processing Unit) 151
- a memory device 152 such as a RAM (Random Access Memory) or a ROM (Read Only Memory)
- a storage device 153 such as a hard disk.
- the functions and processes of each device are realized by the CPU 151 executing data or a program stored in the storage device 153 or the memory device 152.
- the information required for each device may be input from the input / output interface device 154, and the result obtained by each device may be output from the input / output interface device 154.
- each device search device 100 or learning device 200
- each device is described using a functional block diagram, but each device is described by hardware, software, or a combination thereof. It may be realized.
- the examples of the present invention include a program for causing a computer to realize the functions of each device according to the embodiment of the present invention, a program for causing the computer to execute each procedure of the method according to the embodiment of the present invention, and the like. , May be realized.
- each functional part may be used in combination as necessary.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Data Mining & Analysis (AREA)
- General Engineering & Computer Science (AREA)
- Mathematical Physics (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- Software Systems (AREA)
- Computing Systems (AREA)
- Life Sciences & Earth Sciences (AREA)
- Multimedia (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Biophysics (AREA)
- Biomedical Technology (AREA)
- Molecular Biology (AREA)
- Databases & Information Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Probability & Statistics with Applications (AREA)
- Algebra (AREA)
- Computational Mathematics (AREA)
- Mathematical Analysis (AREA)
- Mathematical Optimization (AREA)
- Pure & Applied Mathematics (AREA)
- Medical Informatics (AREA)
- Human Computer Interaction (AREA)
- Image Analysis (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
Description
第1の学習済みニューラルネットワークを用いて、前記クエリデータから第1の特徴ベクトルを抽出する第1の特徴抽出部と、
前記メディアデータから第1の領域を取得し、第2の学習済みニューラルネットワークを用いて、前記第1の領域から第2の特徴ベクトルを抽出する第2の特徴抽出部と、
第3の学習済みニューラルネットワークを用いて、前記第1の特徴ベクトルと、前記第2の特徴ベクトルと、前記第1の領域又は前記第1の領域の位置とに基づいて、前記目標領域の候補を決定する位置特定部と、
前記決定された目標領域の候補を、前記第2の特徴抽出部により用いられる前記第1の領域として使用することにより、所定の条件が満たされるまで、前記第2の特徴抽出部及び前記位置特定部の動作を繰り返す制御部と、
を有する。
第1のニューラルネットワークを用いて、学習用クエリデータから第1の特徴ベクトルを抽出する第1の特徴抽出部と、
学習用メディアデータから第1の領域を取得し、第2のニューラルネットワークを用いて、前記第1の領域から第2の特徴ベクトルを抽出する第2の特徴抽出部と、
第3のニューラルネットワークを用いて、前記第1の特徴ベクトルと、前記第2の特徴ベクトルと、前記第1の領域又は前記第1の領域の位置とに基づいて、前記目標領域の候補を決定する位置特定部と、
前記決定された目標領域の候補を、前記第2の特徴抽出部により用いられる前記第1の領域として使用することにより、所定の条件が満たされるまで、前記第2の特徴抽出部及び前記位置特定部の動作を繰り返す制御部と、
前記位置特定部により決定された前記目標領域の候補が、前記目標領域を捉えたか否かを判定し、当該判定結果に基づいて、前記第1のニューラルネットワーク、前記第2のニューラルネットワーク及び前記第3のニューラルネットワークのパラメータを更新する学習部と、
を有する。
第1の学習済みニューラルネットワークを用いて、前記クエリデータから第1の特徴ベクトルを抽出する第1のステップと、
前記メディアデータから第1の領域を取得し、第2の学習済みニューラルネットワークを用いて、前記第1の領域から第2の特徴ベクトルを抽出する第2のステップと、
第3の学習済みニューラルネットワークを用いて、前記第1の特徴ベクトルと、前記第2の特徴ベクトルと、前記第1の領域又は前記第1の領域の位置とに基づいて、前記目標領域の候補を決定する第3のステップと、
前記決定された目標領域の候補を、前記第2のステップにより用いられる前記第1の領域として使用することにより、所定の条件が満たされるまで、前記第2のステップ及び前記第3のステップを繰り返す第4のステップと、
を有する。
第1のニューラルネットワークを用いて、学習用クエリデータから第1の特徴ベクトルを抽出する第1のステップと、
学習用メディアデータから第1の領域を取得し、第2のニューラルネットワークを用いて、前記第1の領域から第2の特徴ベクトルを抽出する第2のステップと、
第3のニューラルネットワークを用いて、前記第1の特徴ベクトルと、前記第2の特徴ベクトルと、前記第1の領域又は前記第1の領域の位置とに基づいて、前記目標領域の候補を決定する第3のステップと、
前記決定された目標領域の候補を、前記第2のステップにより用いられる前記第1の領域として使用することにより、所定の条件が満たされるまで、前記第2のステップ及び前記第3のステップを繰り返す第4のステップと、
前記目標領域の候補が、前記目標領域を捉えたか否かを判定し、当該判定結果に基づいて、前記第1のニューラルネットワーク、前記第2のニューラルネットワーク及び前記第3のニューラルネットワークのパラメータを更新する第5のステップと、
を有する。
第1の実施例では、クエリデータに一致する目標領域(以下、「目標ウィンドウ」とも呼ばれる)を求めてメディアデータを検索する検索装置の全体構成について説明する。例えば、メディアデータ及びクエリデータは、それぞれ、静止画像、動画像、音響信号又は他のデータである。検索装置は、メディアデータの中の複数の目標領域の候補(以下、「候補領域」又は「候補ウィンドウ」とも呼ばれる)からクエリデータに一致する目標領域を見つけるために、学習装置により学習されたモデル(具体的には、ニューラルネットワーク)を用いる。具体的には、検索装置は、学習済みモデルを用いることにより、クエリデータと候補領域とを繰り返し比較し、候補領域がクエリデータに一致するか否かを判断することによって、目標領域を見つける。モデルは、評価対象の候補領域を減少させるように学習されているため、検索装置は、より少ない数の候補領域で目標領域を見つけることができる。
図1は、本発明の第1の実施例に係る検索装置100の機能構成を示す図である。検索装置100の目的は、メディアデータRの中の特定の位置lgに存在する、クエリデータQによって表される目標領域を特定することである。位置lgは領域の中心でもよく、領域の角でもよく、領域を定義するために用いることができる他の位置でもよい。メディアデータRが静止画像である場合、位置lgはxy画像座標により表すことができる。メディアデータRが動画像である場合、位置lgはタイムスタンプ(又はフレームインデックス)により表すことができる。この場合、位置lgは時間フレームのxy画像座標を含んでもよい。メディアデータRが音響信号である場合、位置lgはタイムスタンプにより表すことができる。検索装置100は、ウィンドウサイズ、方向又は上記のいずれかの組み合わせ等のように、目標領域を定義するために用いられる他の種類の情報を決定してもよい。
図2は、本発明の第1の実施例に係る学習装置200の機能構成を示す図である。学習装置200の目的は、目標領域を特定しつつ、より小さい繰り返し回数Tで検索パス{lt}t=0 Tを決定することである。
第2の実施例では、第1の実施例の概念を用いた画像検索方法について説明する。第2の実施例では、メディアデータは参照画像Rであり、クエリデータはクエリ画像Qである。
図3は、第2の実施例に係る検索装置100の概念図である。
以下、学習装置200により実行される学習方法の各ステップについて、図7を参照して詳細に説明する。
本発明の実施例に記載の画像検索方法を評価するために、MNIST(http://yann.lecun.com/exdb/mnist/)及びFlickrLogos-32(http://www.multimedia-computing.de/flickrlogos/)という2つのベンチマーク用データセットを用いた。
MNISTに関して、3つのデータセットをMNISTデータセットに基づいて生成した。第1のデータセットは「Translated MNIST」と呼ばれ、各参照画像は、28×28の数字の画像(28×28画素の画像)を100×100のブランク画像のランダムな位置に配置することにより生成された。より具体的には、位置座標は、0からブランク画像と数字の画像とのサイズ差までの範囲内の乱数である。この範囲は、ブランク画像の境界位置に数字の画像を配置することを回避するために設定されたものである。
性能指標に関して、画像検索方法を精度及び速度の観点で評価した。クエリ画像及び参照画像を与えて、参照画像の中の領域に対応する予測ウィンドウを出力し、精度を評価するためにこの予測ウィンドウを用いた。特に、画像検索は、物体検出手法と同じ基準に従って、予測ウィンドウと真値ウィンドウとのIoU(intersection over union)が0.5より大きい場合に成功であると考えられるものとした。成功率は、全てのペアに対する正確に一致した画像ペアの数の比である。
全てのデータセットについて、成功率、評価されたウィンドウの数及び実行時間を、それぞれ図8~図10に示す。図8~10において、「Joint-Training」は本発明の実施例の結果である。本発明の実施例に従ってクエリ画像を特定するための検索パスは図11に示されている。
上記のように、本発明の実施例は、参照画像の中の領域に対してクエリ画像をマッチングする問題に対処する。当該方法は、各繰り返し回数において目標領域に向かって次の位置を順に出力するニューラルネットワーク(例えば、CNN及びLSTMの組み合わせ)に基づいている。より具体的には、本発明の実施例は、参照画像のどこで次の領域を抽出するかを決定するために位置特定部を用いる。位置特定部の性能を最大化するために、本発明の実施例は、次の位置を予測するための強化学習に基づく技術を取り入れる。したがって、参照画像の関係する領域に着目することが可能になり、クエリ画像を特定するために必要なウィンドウ(候補領域)の数がかなり減少し、特に大きい参照画像の場合には、より高速な画像検索をもたらす。
図12に、本発明の実施例における各装置(検索装置100又は学習装置200)のハードウェア構成例を示す。各装置は、CPU(Central Processing Unit)151等のプロセッサ、RAM(Random Access Memory)やROM(Read Only Memory)等のメモリ装置152、ハードディスク等の記憶装置153等から構成されたコンピュータでもよい。例えば、各装置の機能及び処理は、記憶装置153又はメモリ装置152に格納されているデータやプログラムをCPU151が実行することによって実現される。また、各装置に必要な情報は、入出力インタフェース装置154から入力され、各装置において求められた結果は、入出力インタフェース装置154から出力されてもよい。
説明の便宜上、本発明の実施例に係る各装置(検索装置100又は学習装置200)は機能的なブロック図を用いて説明しているが、各装置は、ハードウェア、ソフトウェア又はそれらの組み合わせで実現されてもよい。例えば、本発明の実施例は、コンピュータに対して本発明の実施例に係る各装置の機能を実現させるプログラム、コンピュータに対して本発明の実施例に係る方法の各手順を実行させるプログラム等により、実現されてもよい。また、各機能部が必要に応じて組み合わせて使用されてもよい。
110 初期領域予測部
120 第1の特徴抽出部
130 第2の特徴抽出部
140 位置特定部
150 制御部
200 学習装置
210 初期領域予測部
220 第1の特徴抽出部
230 第2の特徴抽出部
240 位置特定部
250 制御部
260 学習部
Claims (9)
- クエリデータに一致する目標領域を求めてメディアデータを検索する検索装置であって、
第1の学習済みニューラルネットワークを用いて、前記クエリデータから第1の特徴ベクトルを抽出する第1の特徴抽出部と、
前記メディアデータから第1の領域を取得し、第2の学習済みニューラルネットワークを用いて、前記第1の領域から第2の特徴ベクトルを抽出する第2の特徴抽出部と、
第3の学習済みニューラルネットワークを用いて、前記第1の特徴ベクトルと、前記第2の特徴ベクトルと、前記第1の領域又は前記第1の領域の位置とに基づいて、前記目標領域の候補を決定する位置特定部と、
前記決定された目標領域の候補を、前記第2の特徴抽出部により用いられる前記第1の領域として使用することにより、所定の条件が満たされるまで、前記第2の特徴抽出部及び前記位置特定部の動作を繰り返す制御部と、
を有する検索装置。 - 学習用クエリデータが前記第1の特徴抽出部に入力され、学習用メディアデータが前記第2の特徴抽出部に入力された場合、前記第1の学習済みニューラルネットワーク、前記第2の学習済みニューラルネットワーク及び前記第3の学習済みニューラルネットワークは、前記位置特定部により決定された前記目標領域の候補が、前記目標領域を捉えたか否かを判定した結果に基づいて、学習用クエリデータ及び学習用メディアデータを用いて学習されている、請求項1に記載の検索装置。
- 前記第1の学習済みニューラルネットワークのパラメータは、前記第2の学習済みニューラルネットワークのパラメータと同じである、請求項1又は2に記載の検索装置。
- 前記メディアデータをダウンサンプリング後のメディアデータにダウンサンプリングし、第4の学習済みニューラルネットワークを用いて、前記ダウンサンプリング後のメディアデータに基づいて前記第1の領域を取得する初期領域予測部を更に有する、請求項1乃至3のうちいずれか1項に記載の検索装置。
- クエリデータに一致する目標領域を求めてメディアデータを検索するために用いられるニューラルネットワークを学習する学習装置であって、
第1のニューラルネットワークを用いて、学習用クエリデータから第1の特徴ベクトルを抽出する第1の特徴抽出部と、
学習用メディアデータから第1の領域を取得し、第2のニューラルネットワークを用いて、前記第1の領域から第2の特徴ベクトルを抽出する第2の特徴抽出部と、
第3のニューラルネットワークを用いて、前記第1の特徴ベクトルと、前記第2の特徴ベクトルと、前記第1の領域又は前記第1の領域の位置とに基づいて、前記目標領域の候補を決定する位置特定部と、
前記決定された目標領域の候補を、前記第2の特徴抽出部により用いられる前記第1の領域として使用することにより、所定の条件が満たされるまで、前記第2の特徴抽出部及び前記位置特定部の動作を繰り返す制御部と、
前記位置特定部により決定された前記目標領域の候補が、前記目標領域を捉えたか否かを判定し、当該判定結果に基づいて、前記第1のニューラルネットワーク、前記第2のニューラルネットワーク及び前記第3のニューラルネットワークのパラメータを更新する学習部と、
を有する学習装置。 - 前記学習用メディアデータをダウンサンプリング後の学習用メディアデータにダウンサンプリングし、第4のニューラルネットワークを用いて、前記ダウンサンプリング後の学習用メディアデータに基づいて前記第1の領域を取得する初期領域予測部を更に有し、
前記第4のニューラルネットワークのパラメータは、前記位置特定部により決定された前記目標領域の候補が、前記目標領域を捉えたか否かを判定し、当該判定結果に基づいて、前記第1のニューラルネットワーク、前記第2のニューラルネットワーク及び前記第3のニューラルネットワークのパラメータと共に更新される、請求項5に記載の学習装置。 - クエリデータに一致する目標領域を求めてメディアデータを検索する検索装置により使用される検索方法であって、
第1の学習済みニューラルネットワークを用いて、前記クエリデータから第1の特徴ベクトルを抽出する第1のステップと、
前記メディアデータから第1の領域を取得し、第2の学習済みニューラルネットワークを用いて、前記第1の領域から第2の特徴ベクトルを抽出する第2のステップと、
第3の学習済みニューラルネットワークを用いて、前記第1の特徴ベクトルと、前記第2の特徴ベクトルと、前記第1の領域又は前記第1の領域の位置とに基づいて、前記目標領域の候補を決定する第3のステップと、
前記決定された目標領域の候補を、前記第2のステップにより用いられる前記第1の領域として使用することにより、所定の条件が満たされるまで、前記第2のステップ及び前記第3のステップを繰り返す第4のステップと、
を有する検索方法。 - クエリデータに一致する目標領域を求めてメディアデータを検索するために用いられるニューラルネットワークを学習する学習装置により使用される学習方法であって、
第1のニューラルネットワークを用いて、学習用クエリデータから第1の特徴ベクトルを抽出する第1のステップと、
学習用メディアデータから第1の領域を取得し、第2のニューラルネットワークを用いて、前記第1の領域から第2の特徴ベクトルを抽出する第2のステップと、
第3のニューラルネットワークを用いて、前記第1の特徴ベクトルと、前記第2の特徴ベクトルと、前記第1の領域又は前記第1の領域の位置とに基づいて、前記目標領域の候補を決定する第3のステップと、
前記決定された目標領域の候補を、前記第2のステップにより用いられる前記第1の領域として使用することにより、所定の条件が満たされるまで、前記第2のステップ及び前記第3のステップを繰り返す第4のステップと、
前記目標領域の候補が、前記目標領域を捉えたか否かを判定し、当該判定結果に基づいて、前記第1のニューラルネットワーク、前記第2のニューラルネットワーク及び前記第3のニューラルネットワークのパラメータを更新する第5のステップと、
を有する学習方法。 - 請求項1乃至6のうちいずれか1項に記載の装置としてコンピュータを機能させるプログラム。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/440,166 US12277160B2 (en) | 2019-03-26 | 2019-09-10 | Search apparatus, training apparatus, search method, training method, and program |
| JP2021508692A JP7192966B2 (ja) | 2019-03-26 | 2019-09-10 | 検索装置、学習装置、検索方法、学習方法及びプログラム |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2019-059437 | 2019-03-26 | ||
| JP2019059437 | 2019-03-26 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020194792A1 true WO2020194792A1 (ja) | 2020-10-01 |
Family
ID=72610464
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2019/035526 Ceased WO2020194792A1 (ja) | 2019-03-26 | 2019-09-10 | 検索装置、学習装置、検索方法、学習方法及びプログラム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US12277160B2 (ja) |
| JP (1) | JP7192966B2 (ja) |
| WO (1) | WO2020194792A1 (ja) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2023509105A (ja) * | 2020-10-26 | 2023-03-07 | 3アイ インコーポレイテッド | ディープラーニングを利用した屋内位置測位方法 |
| JP2025520071A (ja) * | 2022-05-23 | 2025-07-01 | セールスフォース インコーポレイテッド | プログラム合成のためのシステムおよび方法 |
| JP7855727B2 (ja) | 2022-05-23 | 2026-05-08 | セールスフォース インコーポレイテッド | プログラム合成のためのシステムおよび方法 |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11587345B2 (en) * | 2020-07-22 | 2023-02-21 | Honda Motor Co., Ltd. | Image identification device, method for performing semantic segmentation, and storage medium |
| CN112200198B (zh) * | 2020-07-31 | 2023-11-24 | 星宸科技股份有限公司 | 目标数据特征提取方法、装置及存储介质 |
| CN112288003B (zh) * | 2020-10-28 | 2023-07-25 | 北京奇艺世纪科技有限公司 | 一种神经网络训练、及目标检测方法和装置 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2006221525A (ja) * | 2005-02-14 | 2006-08-24 | Chuden Gijutsu Consultant Kk | オブジェクト検索システムおよび方法 |
| JP2006338313A (ja) * | 2005-06-01 | 2006-12-14 | Nippon Telegr & Teleph Corp <Ntt> | 類似画像検索方法,類似画像検索システム,類似画像検索プログラム及び記録媒体 |
| JP2009251667A (ja) * | 2008-04-01 | 2009-10-29 | Toyota Motor Corp | 画像検索装置 |
| JP2018022390A (ja) * | 2016-08-04 | 2018-02-08 | 日本電信電話株式会社 | 検証装置、方法、及びプログラム |
| JP2019028700A (ja) * | 2017-07-28 | 2019-02-21 | 日本電信電話株式会社 | 検証装置、方法、及びプログラム |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP6754619B2 (ja) * | 2015-06-24 | 2020-09-16 | 三星電子株式会社Samsung Electronics Co.,Ltd. | 顔認識方法及び装置 |
| US10860898B2 (en) * | 2016-10-16 | 2020-12-08 | Ebay Inc. | Image analysis and prediction based visual search |
-
2019
- 2019-09-10 WO PCT/JP2019/035526 patent/WO2020194792A1/ja not_active Ceased
- 2019-09-10 US US17/440,166 patent/US12277160B2/en active Active
- 2019-09-10 JP JP2021508692A patent/JP7192966B2/ja active Active
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2006221525A (ja) * | 2005-02-14 | 2006-08-24 | Chuden Gijutsu Consultant Kk | オブジェクト検索システムおよび方法 |
| JP2006338313A (ja) * | 2005-06-01 | 2006-12-14 | Nippon Telegr & Teleph Corp <Ntt> | 類似画像検索方法,類似画像検索システム,類似画像検索プログラム及び記録媒体 |
| JP2009251667A (ja) * | 2008-04-01 | 2009-10-29 | Toyota Motor Corp | 画像検索装置 |
| JP2018022390A (ja) * | 2016-08-04 | 2018-02-08 | 日本電信電話株式会社 | 検証装置、方法、及びプログラム |
| JP2019028700A (ja) * | 2017-07-28 | 2019-02-21 | 日本電信電話株式会社 | 検証装置、方法、及びプログラム |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2023509105A (ja) * | 2020-10-26 | 2023-03-07 | 3アイ インコーポレイテッド | ディープラーニングを利用した屋内位置測位方法 |
| JP7336653B2 (ja) | 2020-10-26 | 2023-09-01 | 3アイ インコーポレイテッド | ディープラーニングを利用した屋内位置測位方法 |
| US11961256B2 (en) | 2020-10-26 | 2024-04-16 | 3I Inc. | Method for indoor localization using deep learning |
| JP2025520071A (ja) * | 2022-05-23 | 2025-07-01 | セールスフォース インコーポレイテッド | プログラム合成のためのシステムおよび方法 |
| JP7855727B2 (ja) | 2022-05-23 | 2026-05-08 | セールスフォース インコーポレイテッド | プログラム合成のためのシステムおよび方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| JPWO2020194792A1 (ja) | 2020-10-01 |
| US12277160B2 (en) | 2025-04-15 |
| JP7192966B2 (ja) | 2022-12-20 |
| US20220188345A1 (en) | 2022-06-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN112561027B (zh) | 神经网络架构搜索方法、图像处理方法、装置和存储介质 | |
| CN110852349B (zh) | 一种图像处理方法、检测方法、相关设备及存储介质 | |
| CN112288770A (zh) | 基于深度学习的视频实时多目标检测与跟踪方法和装置 | |
| WO2020228446A1 (zh) | 模型训练方法、装置、终端及存储介质 | |
| JP7192966B2 (ja) | 検索装置、学習装置、検索方法、学習方法及びプログラム | |
| CN112215332A (zh) | 神经网络结构的搜索方法、图像处理方法和装置 | |
| CN111144425B (zh) | 检测拍屏图片的方法、装置、电子设备及存储介质 | |
| CN116563285A (zh) | 一种基于全神经网络的病灶特征识别与分割方法及系统 | |
| CN113095185B (zh) | 人脸表情识别方法、装置、设备及存储介质 | |
| CN119296143B (zh) | 基于方向场引导和空间注意力技术的指纹特征识别分析方法 | |
| CN115063831A (zh) | 一种高性能行人检索与重识别方法及装置 | |
| CN110033012A (zh) | 一种基于通道特征加权卷积神经网络的生成式目标跟踪方法 | |
| WO2024078112A1 (zh) | 一种舾装件智能识别方法、计算机设备 | |
| CN118786440A (zh) | 使用自监督学习训练对象发现神经网络和特征表示神经网络 | |
| CN116824330A (zh) | 一种基于深度学习的小样本跨域目标检测方法 | |
| CN116258877A (zh) | 土地利用场景相似度变化检测方法、装置、介质及设备 | |
| CN119206853A (zh) | 行人检测方法、装置、设备、存储介质及产品 | |
| CN116823734B (zh) | 用于跟踪医学图像中的对象组的系统和方法 | |
| CN111179270A (zh) | 基于注意力机制的图像共分割方法和装置 | |
| JP5430243B2 (ja) | 画像検索装置及びその制御方法並びにプログラム | |
| CN120375207A (zh) | 一种遥感图像变化检测方法、装置、设备及介质 | |
| Wang et al. | A single-stream adaptive scene layout modeling method for scene recognition | |
| CN116860998B (zh) | 一种基于全局上下文特征融合知识图谱的目标检测方法 | |
| CN114549591B (zh) | 时空域行为的检测和跟踪方法、装置、存储介质及设备 | |
| Antar et al. | Robust Object Recognition with Deep Learning on a Variety of Datasets. |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19920792 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2021508692 Country of ref document: JP Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19920792 Country of ref document: EP Kind code of ref document: A1 |
|
| WWG | Wipo information: grant in national office |
Ref document number: 17440166 Country of ref document: US |













