WO2024158073A1 - 학습 모델을 이용한 영상들 간의 유사도 판단 - Google Patents
학습 모델을 이용한 영상들 간의 유사도 판단 Download PDFInfo
- Publication number
- WO2024158073A1 WO2024158073A1 PCT/KR2023/001254 KR2023001254W WO2024158073A1 WO 2024158073 A1 WO2024158073 A1 WO 2024158073A1 KR 2023001254 W KR2023001254 W KR 2023001254W WO 2024158073 A1 WO2024158073 A1 WO 2024158073A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- image
- images
- learning model
- processor
- labels
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T3/00—Geometric image transformations in the plane of the image
- G06T3/40—Scaling of whole images or parts thereof, e.g. expanding or contracting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
- G06V10/54—Extraction of image or video features relating to texture
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
- G06V10/56—Extraction of image or video features relating to colour
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/74—Image or video pattern matching; Proximity measures in feature spaces
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/774—Generating sets of training patterns; Bootstrap methods, e.g. bagging or boosting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/70—Labelling scene content, e.g. deriving syntactic or semantic representations
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/20—Movements or behaviour, e.g. gesture recognition
Definitions
- This disclosure relates to technology for determining similarity between images using a learning model.
- Measures to protect copyright for content can be divided into proactive measures to make it difficult to copy, distribute, and distribute copyrighted works, and reactive measures to detect and crack down on illegally copied, distributed, and distributed works.
- proactive measures much progress has been made in terms of technology, such as watermarking technology to prevent duplication or limit the number of duplications.
- watermarking technology to prevent duplication or limit the number of duplications.
- the method based on prior measures has been largely neutralized by the development of technology to lift restrictions, and its application is often inappropriate in reality due to the effect of prohibiting copies that do not constitute direct infringement of copyrighted works without distinction. . Therefore, as a follow-up measure, detection and detection of acts that infringe copyright must be continuously carried out in parallel.
- the technical problem to be solved through an embodiment of the present disclosure is to provide a technology for training a learning model capable of determining similarity between images through learning images.
- the technical problem to be solved through an embodiment of the present disclosure is to provide a technology for determining the similarity between images by determining the similarity of objects appearing in an image using a learning model.
- An electronic device may be proposed.
- An electronic device may include one or more processors and one or more memories.
- the processor samples a training image composed of a sequence of image frames into a plurality of frames, detects one or more objects appearing in each of the plurality of frames, generates a plurality of object images corresponding to the one or more objects, and generates a plurality of object images corresponding to the one or more objects. It may be configured to label object images and train a learning model to classify each of the one or more objects.
- An electronic device may be proposed.
- An electronic device may include one or more processors, a network interface, and one or more memories.
- the memory may store a learning model, and the learning model may be a model in which the correlation between objects and labels included in the learning image is learned.
- the processor loads the learning model from the memory, inputs a plurality of images obtained from the input target image into the learning model, obtains similarity for each of the one or more labels, and calculates the similarity to each of the one or more labels. It may be configured to calculate the total similarity of the target image to the training image by multiplying the weight of each of the one or more labels.
- a method may be proposed.
- a method includes sampling an input learning image into a plurality of frames, detecting one or more objects appearing in each of the plurality of frames, and generating a plurality of object images corresponding to the one or more objects.
- a learning model can be trained to classify images through training images.
- the accuracy of the learning model can be improved by modifying the images of the learning video in various ways and training the learning model.
- the similarity of an input image to an image previously learned in a learning model can be determined. Therefore, it can be determined whether the input image is similar to a specific image.
- copyright infringement can be prevented by preventing the upload of images that may pose a copyright infringement problem.
- FIG. 1 is a diagram illustrating an exemplary sharing platform system for determining whether images are similar according to an embodiment of the present disclosure.
- Figure 2 is a block diagram of an electronic device according to an embodiment of the present disclosure.
- Figure 3 is a diagram illustrating sampling of a learning image according to an embodiment of the present disclosure.
- FIGS. 4A and 4B are diagrams illustrating exemplary training of a learning model according to an embodiment of the present disclosure.
- Figure 5 is a diagram illustrating a method for generating an image group according to an embodiment of the present disclosure.
- Figure 6 is a diagram illustrating a similarity comparison method according to an embodiment of the present disclosure.
- FIGS. 7A and 7B are diagrams illustrating a method for determining similarity according to an embodiment of the present disclosure.
- Figure 8 is a flowchart of a method for determining whether an input image is similar to a training image according to an embodiment of the present disclosure.
- the expression "based on” is used to describe one or more factors that influence the act or action of decision, judgment, or action described in the phrase or sentence containing the expression, and this expression It does not exclude additional factors that may influence the decision, act or action of judgment.
- a component when referred to as being “connected” or “connected” to another component, it means that the component can be directly connected or connected to the other component, or as a new component. It should be understood that it can be connected or connected through other components.
- a processor configured to perform a specific operation refers to a general purpose processor capable of performing that specific operation by executing software. It may mean, or it may mean a special purpose computer structured through programming to perform a specific operation.
- FIG. 1 is a diagram illustrating a sharing platform system for determining whether images are similar according to an embodiment of the present disclosure.
- the shared platform system 100 may include an electronic device 103, a learning model 104, and a server 105.
- user 101 may create content using a user device (not shown).
- Content may include videos, images, music, etc.
- a video file is a general term for moving video material and may mean, for example, a file with an avi, mp4, or mkv extension.
- user 101 may upload video 102 to sharing platform system 100 to share with other users.
- the first user may upload the video 102 to the sharing platform system 100
- the second user may view the video 102 input by the first user on the sharing platform system 100. there is.
- the electronic device 103 may upload the input image 102 to the server 105 and may perform various processing before uploading it to the server 105.
- a plurality of components may be implemented in an integrated form in an actual physical environment, and some components may be added or deleted as needed.
- the electronic device 103 of the present disclosure may be of various types.
- electronic device 103 may be a computing device or a portable communication device.
- the electronic device 103 is connected to the user 101 and can transmit and receive various data.
- the electronic device 103 and the user 101 use a wired network such as a Local Area Network (LAN), a Wide Area Network (WAN), or a Value Added Network (VAN), or a mobile communication network. It can be run through all types of wireless networks such as (mobile radio communication network), satellite communication network, Bluetooth, Wibro (Wireless Broadband Internet), HSDPA (High Speed Downlink Packet Access), etc.
- the electronic device 103 may prohibit the upload of the video 102.
- the electronic device 103 may receive the image 102 and check whether it is similar to another image through the learning model 104 before uploading it to the server 105 or another storage device (not shown).
- the input image may be distinguished from other images being compared and may be referred to as an input image or a target image.
- the learning model 104 may be integrated and implemented within the electronic device 103, or may be communicated separately.
- the electronic device 103 may train (learn) the learning model 104 in advance before the image 102 is input.
- the learning model 104 is a model for classifying images and can be learned using multiple images. Hereinafter, for convenience of explanation, it is assumed that the learning model 104 is implemented separately from the electronic device 103.
- the electronic device 103 may receive the learning image 111, which will be specified later by description in the specification, from the operator 110. Alternatively, the electronic device 103 may download the learning video 111 from a network, server, etc.
- the training image 111 may include a number of consecutive frames in time series.
- the electronic device 103 can train the learning model 104 based on the received training image 111.
- the electronic device 103 may sample the learning image 111 into a plurality of frames and generate a plurality of object images by detecting one or more objects appearing in the plurality of sampled frames.
- the electronic device 103 can train the learning model 104 by labeling the plurality of object images it has created.
- the learning model 104 may be trained for correlation by associating the image of the object with the label of the object. After training, the learning model 104 can view objects included in a specific image and classify them into one of the designated labels.
- the learned learning model 104 can determine the degree of similarity between images.
- the electronic device 103 can determine whether copyright infringement exists based on the similarity between images determined by the learning model 104. For example, the learning model 104 determines how similar the images of the object in the image 102 uploaded by the user 101 are to the images of the object in the learning image 111 and compares the image 102 and the learning image. The total similarity of (111) can be calculated. The electronic device 103 may determine whether to upload the image 102 based on the calculated total similarity.
- FIG. 2 is a block diagram of an electronic device according to an embodiment of the present disclosure.
- the reference number of the electronic device is written as '200' for convenience of explanation, but the electronic device 103 and the electronic device 200 do not have different device configurations and may be the same or corresponding devices.
- the electronic device 200 monitors whether the image 102 input by the user is similar to the learning image 111 and provides a service that can determine copyright infringement. can be provided.
- the electronic device 200 may include at least one processor 210, at least one memory 220, and at least one network interface 230.
- the processor 210 controls the memory 220 and/or the network interface 230 and executes instructions stored in the memory 220 to implement the description, function, procedure, proposal, method, and/or operation flowchart of the present disclosure. It can be configured to do so.
- the processor 210 may receive signals and/or data through the network interface 230 and store information included in the signals and/or data in the memory 220.
- the memory 220 may be connected to the processor 210 and may store various information related to the operation of the processor 210. For example, memory 220 may perform some or all of the processes controlled by processor 210, or may include instructions for performing the descriptions, functions, procedures, suggestions, methods, and/or operational flowcharts of the present disclosure. Software code containing (instructions) can be stored. Memory 220 may include non-transitory computer-readable media, such as high-speed random access memory and/or non-volatile computer-readable storage media (e.g., one or more disk storage devices, flash memory devices, or other non-volatile solid-state memory devices). It can be included. When the learning model 104 is implemented in the electronic device 200, the learning model 104 may be stored and implemented in the memory 220.
- non-transitory computer-readable media such as high-speed random access memory and/or non-volatile computer-readable storage media (e.g., one or more disk storage devices, flash memory devices, or other non-volatile solid-state memory devices).
- the network interface 230 may be connected to the processor 210 and may transmit and/or receive wired/wireless signals or data.
- the network interface 230 may be connected to a user terminal and/or a server and a database through a wired/wireless communication network.
- the wireless communication network may include a mobile communication network, wireless LAN, and short-range wireless communication network.
- wireless communication networks include LTE, LTE Advance (LTE-A), code division multiple access (CDMA), wideband CDMA (WCDMA), universal mobile telecommunications system (UMTS), Wireless Broadband (WiBro), and Global System (GSM). for Mobile Communications), etc. may include cellular communications using at least one of the following.
- the wireless communication network may include at least one of wireless fidelity (WiFi), Bluetooth, Bluetooth low energy (BLE), Zigbee, near field communication (NFC), and radio frequency (RF).
- WiFi wireless fidelity
- BLE Bluetooth low energy
- NFC near field communication
- RF radio frequency
- the wired communication network may include at least one of USB (Universal Serial Bus), USART (Universal Synchronous/Asynchronous Receiver Transmitter), and Ethernet.
- USB Universal Serial Bus
- USART Universal Synchronous/Asynchronous Receiver Transmitter
- Ethernet Ethernet
- Each of the at least one network interface 230 may correspond to the wired/wireless communication network described above.
- the functional units included in the electronic device 200 disclosed below are hardware including the processor 210, memory 220, and network interface 230, software for implementing instructions, or a combination of hardware and software. It can be implemented.
- Figure 3 is a diagram showing a processor sampling a learning image according to an embodiment.
- the processor 210 may sample the received image 102 at a predetermined sampling period. Sampling may mean acquiring a frame from an image.
- the sampling cycle may be stored in advance in the memory 220 and may be determined by receiving input from the operator 110.
- the sampling time point may be determined by the input of the operator 110, the sampling period may be performed periodically, and the frame at a specific time point may be sampled aperiodically according to the input.
- Computing load can be reduced by processing only sampled frames rather than entire frames.
- the processor 210 may obtain a motion vector for the movement of an object in the image 102 and measure the magnitude of the motion vector for a predetermined period of time. If the size of the motion vector is greater than a certain value during a certain period of time, it may mean that the object's movement is large.
- a motion vector can be obtained by using an adaptive difference image on a background model generated through a Gaussian mixture model as a method to separate an object with noticeable movement from the background. If the amount of change in the motion vector during a given time is greater than or equal to the first value, this may mean that the importance of the object in the image is high.
- the processor 210 may track the object and sample a frame in which the object appears.
- the processor 210 may use point tracking, kernel tracking, silhouette tracking, etc. to track the object. Accordingly, the processor 210 may reduce computing load by sampling only meaningful frames.
- the processor 210 may obtain a plurality of frames 310 corresponding to different times by sampling the image 300.
- the interval between sampled times T1, T2, T3, and T4 may be constant or may be an arbitrary time interval that is not constant.
- each frame 311, 312, 313, and 314 represents a frame of an image at the sampled time.
- FIGS. 4A and 4B are diagrams illustrating exemplary training of a learning model according to an embodiment of the present disclosure.
- the reference number of the learning model is indicated as '414' for convenience of explanation, but the learning model 104 and the learning model 414 are not different models and may be the same or corresponding learning models.
- the processor 210 may obtain a plurality of frames 311, 312, 313, and 314 from the received image 102 through sampling.
- one of the frames 410 may include an object 411.
- the object 411 may be a character, person, specific shape, clothing, building, or logo.
- the processor 210 may identify a region (hereinafter referred to as Region of Interest (ROI)) where the object is located in order to detect the object 411 in the frame.
- ROI Region of Interest
- the processor 210 may find similar areas in a frame based on at least one of color, texture, and pattern, and set an ROI by grouping similar areas based on at least one of color, texture, and pattern. For example, there may be a sharp difference in RGB values based on the boundary of the object 411. That is, the processor 210 can set areas with similar colors as similar areas, zone them, and identify them as ROI. The processor 210 may take the differential of the color value in the frame to determine similar areas, view the portion where the differential value exceeds a certain threshold as a boundary, and group similar areas. In this case, ROI can be easily set based on similar areas.
- the processor 210 can use a Sobel filter to detect the outline of an object, using a change in the brightness value of a pixel or a change in color (RGB) value by a differential operator to find the outline.
- the Sobel filter is a non-linear operator that calculates the difference in the RGB sum between pixels at both ends of the mask window area and then calculates the average size in the horizontal and vertical directions to emphasize the contour area. do. Since the function of the Sobel filter is known, detailed description is omitted in this specification.
- the processor 210 may classify the frame into multiple regions, histogram the texture features for each region, and group portions with histogram values within a certain size range into one region to identify the ROI.
- the processor 210 may receive input from the operator 110 and set the ROI.
- an ROI can be set to that area.
- the processor 210 may separate the object from the background based on a set ROI.
- the processor 210 may generate a box 411 based on the ROI and crop the image based on the box 411. The size of the cropped image may be different.
- the processor 210 may additionally process the size of the cropped images to adjust them to a certain size. For example, images can be adjusted to a certain size for training data.
- the processor 210 may pre-classify the cropped images before training depending on whether they are images corresponding to a specific object. For example, the processor 210 may find feature points between images and create a group of similar images. The processor 210 may create an image group indicating a specific object and train the learning model 104 based on the image group. Group creation of images is explained in detail in FIG. 5 below.
- processor 210 may identify an ROI from a frame and acquire the image as well as process the image or generate additional images.
- a plurality of objects may overlap in the cropped image.
- a first object and a second object may overlap with a plurality of objects.
- the processor 210 may acquire an image of the first object in a frame different from the corresponding frame.
- the processor 210 may separate the image of the frame based on the image of the first object in the frame from another viewpoint.
- the other time point may be before the point of the frame, or it may be a point in time after some time has elapsed after the point of the frame.
- the processor 210 may separate the image of the first object in the previous frame from the corresponding frame and generate the remaining portion as an image of the second object. Separating an image may mean removing the color of the image or separating the layers of the image. In this case, even when objects overlap, each object can be separated through object images from different viewpoints to create a variety of images for training.
- the processor 210 may generate training data by merging images. Partial images of the object obtained in the first frame and partial images of the object obtained in the second frame may be merged. For example, if the processor 210 determines that the images have continuity or if it determines that the images have a common shape, the processor 210 may merge the images to generate learning data.
- the processor 210 may train the learning model 414 based on a plurality of images sampled and cropped from the learning image 111 and additionally processed images.
- the operator 110 can sufficiently train the model by inputting one learning image 111 without the need to separately input a large amount of learning images.
- the processor 210 may train the learning model 414 based on images of various objects.
- the processor 210 may perform supervised learning on the learning model 414.
- Supervised learning means training a model with certain images according to a given label so that the model can classify a specific image with a specific label.
- the processor 210 may receive a label for the image from the operator 110 each time the image is trained. For example, processor 210 may allow an operator to check one of the cropped images through a display (not shown). The operator 110 can view the image and input the corresponding label by typing, clicking, etc.
- the learning model 414 can be supervised by receiving a label for each image.
- the processor 210 may train the learning model 414 by classifying similarly classified images with arbitrary labels. That is, a label can be determined for each image group, and the learning model 414 can be trained with the same label for images belonging to the same image group. For example, the processor 210 classifies the plurality of images 412, 413, and 414 as corresponding to the same label and labels the images 412, 413, and 414 as “Label 1” to create the learning model 414. can be trained. For example, the learning model 414 may be trained to classify an image similar to the plurality of images 412, 413, and 414 as “Label 1.”
- the processor 210 may load labels pre-stored in the memory 220 and images corresponding to each label and use them for training.
- the processor 210 can find images corresponding to the loaded labels, initially classify them according to the labels, and then use the classified images as training data for the learning model 414.
- the processor 210 may calculate the probability of extracting an image of a specific object and determine a learning target based on this. For example, the processor 210 may calculate the ratio that the image 422 for object A occupies in all object images. Alternatively, the processor 210 may calculate the number of images corresponding to the image 422 for object A among the acquired images. For example, when a label is input from an operator, the processor 210 can calculate the number of images corresponding to a specific label. For example, if the processor 210 creates an image group by classifying similar images, the processor 210 may calculate the number of images belonging to the corresponding image group.
- the processor 210 may select the specific object as a learning target if the number of images corresponding to the specific object is greater than a predetermined number. Selecting it as a learning target means using it as an input for training the learning model 414.
- the processor 210 may select the specific object as a learning target if the image corresponding to the specific object occupies a certain percentage or more of the total image. For example, the ratio of the image for object A among all images cropped from the entire sampled frame may be more than 70%.
- the processor 210 can select images with a ratio exceeding 30% as learning targets, and since the ratio occupied by the images for A exceeds the corresponding threshold, the images for A can be selected as learning targets. We can then train images for A.
- the processor 210 may measure the total number of cropped images to calculate the probability that it corresponds to an image of a specific object, and among them, an image labeled A or an image similar to A may be selected. The number of grouped images can be measured. The processor 210 may calculate the probability by dividing images labeled or grouped with the same type by the total number of images.
- the processor 210 may determine that the probability is lower than the probability defined above if the image of an object does not appear frequently. Alternatively, if images of an object do not appear frequently, the processor 210 may determine that the absolute number of images for that object is less than a predetermined value and exclude the images of the corresponding object from the learning target. The processor 210 stores all cropped images in the memory 220 and can load and use them when necessary. If excluded from the learning target, the corresponding image can be removed from the memory 220.
- the processor 210 Before learning the learning model 414, the processor 210 selects learning targets through pre-filtering and selects only objects with a certain frequency or more as learning targets, thereby preventing unnecessary waste of resources. In addition, the processor 210 may prevent the need for excessive labels by limiting the learning target to a portion, thereby saving the time required for later classification by the learned model.
- Figure 5 is a diagram illustrating a method for generating an image group according to an embodiment of the present disclosure.
- the reference number of the learning model is indicated as '530', but the learning model 104, learning model 414, and learning model 530 are not different models and are the same or corresponding learning models. You can.
- the processor 210 may classify similar objects into one group.
- feature points of each object may be extracted from the image. Extracting feature points of an image may include extracting edges, corners, size of the object, color of the object, rigidity of the object, illuminance, etc. For example, to extract a boundary line where pixel values suddenly change within an image, the size of the gradient vector obtained by differentiating the image can be used. If the extracted feature points are similar between images, they can be classified as similar images.
- group 1 510 may include different images 511 and 512 for object A.
- the different images 511 and 512 may be images that have similar feature points but have overall different sizes, expressions, actions, poses, etc., when group classification is based on feature points.
- Group 2 520 may include different images 521 and 522 for object B.
- the processor 210 may first classify the images into groups and then check the similarity once more using other feature points to further process the images so that images of different objects are not included in the group. . For example, if the processor 210 has primarily classified images based on corners, it can secondarily determine whether or not images of objects are similar based on edges.
- the processor 210 may adjust the classification of images by adjusting a reference value for determining similarity between images.
- the processor 210 broadly views the similarity between images, the number of labels decreases as the total group decreases, and the output complexity of the learning model 530 can be reduced.
- the processor 210 may set up groups (i.e., classify) and input the grouped images as training data for the learning model 530.
- the processor 210 has the effect of efficiently training the learning model 530 by classifying images and organizing learning data. Accordingly, when the processor 210 classifies images, the user does not need to classify them one by one, thereby improving user convenience.
- the processor 210 may assign a label to each group.
- a label may represent a meaningless sequence, but in other cases it may represent a property of an object.
- the operator 110 can set the corresponding character names for group 1 (510) and group 2 (520) as labels.
- the label can be set as a characteristic of the object of the group (for example, the name of a character, the name of a person, the name of a logo, a specific letter, etc.), or it can be set simply as an ordinal number or alphabet.
- the number of images for learning may be insufficient.
- the number of images for group 1 may be less than the number required for learning.
- the number of images required for learning may be predetermined in the processor 210.
- the minimum number of images required for learning may be 50, and the number of images for group 1 may be less than 50.
- the processor 210 may additionally generate images for the group of the corresponding label. For example, if the number of images for an object is less than a predetermined number, the processor 210 may generate additional images through an image processing model and use the additional images as training data for the learning model 530.
- the image processing model may be stored in advance in the memory 220 and loaded by the processor 210.
- image processing models can resize an image, change its color, perform a color transformation, adjust saturation, crop part of an image, rotate an image, or merge an image with another image. Alternatively, processing can be done to adjust the aspect ratio of the image.
- the processor 210 may generate additional images for specific objects through an image processing model. Additional images increase the number of images for an object and can be used to train the learning model 530 using additional data even when there is insufficient data for training.
- processor 210 may use a predefined image generation algorithm to generate images for further learning.
- the image creation algorithm can perform not only the image conversion described above, but also the operation of compositing other images and layers into the image.
- the learning model 530 may additionally request an image corresponding to a required label from the operator 110.
- the learning model 530 may be configured to be learned by at least one of linear discriminant analysis, K Nearest Neighbor (KNN), decision tree, neural network, and Support Vector Machine (SVM) methods.
- KNN K Nearest Neighbor
- SVM Support Vector Machine
- the learning model 530 may correspond to a deep neural network (DNN) including a plurality of layers, and may be simply referred to as a 'neural network'.
- the plurality of layers may include an input layer, a hidden layer, and an output layer.
- Neural networks may include fully connected networks (FCN), convolutional neural networks (CNN), and recurrent neural networks (RNN).
- FCN fully connected networks
- some of the plurality of layers in the neural network may correspond to a convolutional neural network (CNN), and other parts may correspond to a fully connected network (FCN).
- the convolutional neural network (CNN) may be referred to as a convolutional layer
- the fully connected network (FCN) may be referred to as a fully connected layer.
- the data input to each layer may be referred to as an input feature map, and the data output from each layer may be referred to as an output feature map. You can.
- the input feature map and output feature map may be referred to as activation data.
- Deep learning is a machine learning technique to solve problems such as image or voice recognition from big data sets.
- Figure 6 is a diagram illustrating a similarity comparison method according to an embodiment of the present disclosure. It is assumed that the learning model 104 has been trained for classification in advance through the learning image 111 prior to inputting the image 102 for comparing similarity.
- the image 102 may be transmitted to the electronic device 103 and input into the learning model 104.
- the electronic device 103 samples the image 102 into a plurality of frames and inputs the plurality of frames into the learning model 104, so that the learning model 104 can perform classification and determine the degree of similarity based on the classification result. there is.
- the electronic device 103 is not limited to sampling and inputting an image for a frame or an image for an object into the learning model 104, and the learning model 104 receives the image 102 itself and creates a frame. It may be configured to perform sampling and then determine similarity.
- an image of the object 612 may be extracted from one frame 611 among a plurality of sampled frames.
- Object 612 can be identified in the frame, cropped, and extracted into image 613 in the same way as creating data for training described in FIG. 4A.
- the image 613 may be input to the learning model 104 and compared with a plurality of images 621, 622, and 623 for each one or more labels 620.
- the learning model 104 may compare the input image with all images of all labels input in labeling for learning.
- the image 612 is not cropped separately, but the image for the entire frame 611 is input to the learning model 104 and compared with a plurality of images 621, 622, and 623 corresponding to the label 620. It can be. That is, the learning model 104 can determine whether a plurality of images 621, 622, and 623 corresponding to the label 620 are included in the image for the input frame. For example, if it is determined that an image corresponding to a label is included in an input frame, the frame may be classified with the determined label. If a frame is not extracted as one or more object images, but is input to the learning model 104 as a whole frame, there may be images with multiple labels in the frame. If a frame includes an image corresponding to one or more labels, the frame may be classified as the image corresponding to the label with the highest weight, which will be explained later in FIG. 7B.
- the learning model 104 may change the input image to determine similarity.
- the image 613 may be an inverted shape of a specific object.
- Learning model 104 may change image 613 and compare it to images for labels.
- the learning model 104 may change the image 613 by changing its size, color, cropping the image, or rotating the image and compare the changed image with the images for each label.
- the learning model 104 may rotate the image 613 by 180 degrees to create the image 614, and compare the rotated image with a plurality of images 621, 622, and 623 for the label 620. .
- similarity can be accurately calculated even for transformed images.
- the learning model 104 is not limited to transforming the image, and the processor 210 can transform the image in advance (e.g., change size, change color, crop image, rotate image, etc.) before inputting the image. It may be possible.
- the learning model 104 may obtain feature points of an image and compare them to the image corresponding to the label. For example, the learning model 104 may extract corners from images as feature points and compare them with images for labels.
- FIG. 7A and 7B are diagrams showing a method for determining similarity according to an embodiment of the present disclosure.
- the learning model 104 may determine a degree of similarity regarding how similar the image 102 input by the user is to the learning image 111 used as training data for the learning model 104. To this end, the learning model 104 may receive a first image of an object among the frames of the image 102 and then extract feature points of the image. Feature points may be corners, edges, etc. The first image 701 of the extracted object or the extracted feature points of the first image 701 may be compared with the feature points of a plurality of images of labels stored in the learning model 104. For example, similarity to the label may be calculated based on the number of matching feature points between the extracted feature points of the first image 701 and images corresponding to the label.
- the similarity with the corresponding label when there are multiple images for one label, the similarity with the corresponding label may be determined based on the image with the highest similarity. In other cases, when there are multiple images in one label, the similarity with the corresponding label may be determined by calculating the average similarity of the multiple images.
- the learning model 104 may generate a representative image for a label based on learned images corresponding to each label. In this case, when determining similarity, the image of the object extracted from the frames of the image 102 is not all images corresponding to the label, but rather the image between the representative image for each label of the learning model 104 and the first image 701. Similarity can be judged. For example, the learning model 104 may extract feature points of the input first image 701 and a representative image of the label and determine the similarity by calculating the number of matching feature points. By using representative images, the comparison speed of the learning model 104 can be improved.
- the similarity may be calculated using a neural network model through a softmax function for each of one or more labels such that the sum of the similarities for all labels is 1.
- the first image 701 may have similarities of 0.75, 0.24, and 0.01 for labels 1, 2, and 3, respectively. This may indicate that the first image 701 is similar to the image of label 1 by 0.75, that is, by 75%.
- the second image 711 may have similarities of 0.11, 0.37, and 0.52 for labels 1, 2, and 3, respectively. This may mean that the second image 711 is 11% similar to the image of label 1, 37% similar to the image of label 2, and 52% similar to the image of label 3.
- learning model 104 can classify which label a particular image belongs to. In one embodiment, the learning model 104 may classify the image according to the label with the highest similarity among the labels. For example, since label 1 has the highest similarity to the first image 701, the first image 701 may be classified as label 1. For example, since the learning model 104 has the highest similarity to label 3 for the second image 711, the second image 711 may be classified as label 3.
- the learning model 104 may classify a label as having the highest similarity to a specific label and if the similarity exceeds a certain threshold. On the other hand, even if the learning model 104 has the highest similarity with a specific label, if the similarity is below the threshold, it may not classify it as any label. For example, the learning model 104 can classify with the label if the similarity with the label exceeds 0.5 (50%). For example, the first image 701 has a similarity with label 1 of 0.75, so it can be classified as label 1 702. That is, the first image 701 may be determined to correspond to label 1. For example, a particular image (not shown) may have a similarity of no more than 0.5 for all labels. In this case, the corresponding image may not be classified with any label.
- the image 102 is similar to the learning image 111 even though it is not similar to it. Problems that may be considered problematic can be resolved.
- FIG. 7B shows the total value of the training image 111 of the input image 102 based on the similarity to the labels trained in the learning model of each of one or more images in the image 102, according to an embodiment of the present disclosure. This diagram shows how to determine similarity.
- the learning model 104 compares the images of each object extracted from the input image 102 with a plurality of images for each of one or more labels, as shown in FIG. 7A, or selects a representative image for each label. After comparing with , the similarity of each label of the object image can be determined. In one embodiment, the processor 210 may determine the total similarity of all extracted objects in the frames of the image 102 output from the learning model 104 based on the similarity determined for each label. The total similarity refers to the similarity of the input image 102 containing one or more objects to the training image 111.
- the total similarity may be a function of the weight of the label.
- the processor 210 may set different weights for each label and calculate the total similarity of the input image 102 to the training image 111 by multiplying the similarity for each of one or more labels by the weight of one or more labels. there is. Through this, the degree of similarity between the input image 102 and the training image 111 can be meaningfully and conveniently calculated.
- the processor 210 may calculate the frequency of how many object images for the corresponding label appear in the training image 111. A reasonable similarity judgment can be made by calculating weights based on frequency. Based on the calculated frequency, the processor 210 may assign a higher weight to the label of an object that appears frequently, and may assign a lower weight to the label of an object that appears less frequently. The processor 210 may receive a weight for the label from the operator 110. Information about weights may be stored in advance in the processor 103.
- the processor 210 may set the weight of the matching label for objects that frequently appear in the input image 102 to be high.
- the label evaluation model can set the weight of the label.
- the label evaluation model may be learned as a function based on the frequency of objects corresponding to the label in the image 102.
- the processor 210 may not set the weight for the label but may directly set the weight for the objects 721, 722, and 723 from the input image 102.
- the processor 210 measures the motion vectors of the objects 721, 722, and 723 for a predetermined period of time, and, if the amount of the motion vector is greater than a predetermined value, determines that the amount of motion vector is high and sets the weight high.
- the total similarity of the input image 102 to the training image 111 may be calculated based on weights determined for the labels or objects 721, 722, and 723.
- the learning model 104 may multiply the similarity of each of one or more labels by the weight of each label and calculate the total similarity based on this. For example, the learning model 104 may calculate the total similarity 724 that the input image 102 is 73% similar to the training image 111 based on the respective similarities of the objects 721, 722, and 723. .
- the electronic device 103 may restrict uploading of the input image 102 if the calculated total similarity 724 is greater than or equal to a threshold.
- the electronic device 103 may display a warning message about video infringement to the user 101.
- the electronic device 103 may prohibit uploading the video 102 to the server 105.
- the electronic device 103 may prohibit storing the image 102 in the memory of the electronic device 103.
- the electronic device 103 may prevent copyright infringement by prohibiting uploading to the server 105 a video that is similar to a video in which another person's copyright is at issue, which may be an example of the learning video 111. .
- FIG 8 is a flowchart of a method for determining whether an input image is similar to a training image, according to an embodiment of the present disclosure.
- each step may be performed sequentially, but is not necessarily performed sequentially.
- the order of each step may be changed, and at least two steps may be performed in parallel.
- step S810 the electronic device 103 samples a plurality of frames from the uploaded video 102. For example, when a sampling period is set, the electronic device 103 can sample and acquire a frame at each corresponding period. If the sampling period is not determined, the electronic device 103 can receive a desired frame time from the operator 110 and sample it.
- the electronic device 103 may generate an image corresponding to the object by cropping the area where the object is located in a plurality of sampled frames.
- the area where the object is located can be identified as a ROI and can be designated by the operator 110 by dragging the mouse. If the area where the object is located is identified by ROI designation rather than the input of the operator 110, the area where the object is located may be identified by an algorithm for finding the ROI.
- the electronic device 103 may supervised training a learning model by labeling images corresponding to objects.
- the electronic device 103 may input an image corresponding to an object and a label for the corresponding image into a learning model as learning data.
- the learning model 104 can find a function that maps input variables and output variables. Training may be performed before the image 102 for which similarity is to be determined is input to the electronic device 103.
- step S840 similarity for each label of objects in the input image 102 may be obtained through the learning model 104.
- the input image may be divided into a plurality of frames and input to the learning model 104.
- the input image may be divided into a plurality of frames, and then images of the object may be cropped and input to the learning model 104.
- the learning model 104 may output the most similar label among labels for the corresponding image as an output. For example, if the learning model 104 is supervised with labels A, B, and C, the image of the object input to the learning model 104 may be classified as one of A, B, and C.
- the learning model 104 may output the similarity that the image of the object has with each label as an additional output. For example, the similarity with labels A, B, and C can be output.
- the electronic device 103 may calculate the total similarity of the target image to the training image by setting different weights for each label. For example, the total similarity can be obtained by multiplying the similarity for each of one or more labels by the weight of each of one or more labels.
- the electronic device 103 may use the weight according to the label by loading a value previously stored in the memory 220 of the electronic device 103. In one embodiment, the electronic device 103 may set the weight based on the frequency of the object image according to the label. In one embodiment, the electronic device 103 may set the weight based on the input of the operator 110. For example, even if the image of an object corresponding to a specific label appears less frequently than other images in the input image 102 according to the weight, it can contribute to increasing the total similarity between the input image and the training image 111.
- the electronic device 103 may calculate the total similarity based on the number of object images matching the corresponding label only when the weight for a specific label is greater than or equal to a predetermined value. In this case, rather than calculating the total similarity by considering the similarity of all labels, only meaningful labels are calculated, which has the effect of speeding up the calculation speed and reducing the computing load.
- the embodiments described above may be implemented with hardware components, software components, and/or a combination of hardware components and software components.
- the devices, methods, and components described in the embodiments may include, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, and a field programmable gate (FPGA).
- ALU arithmetic logic unit
- FPGA field programmable gate
- It may be implemented using a general-purpose computer or a special-purpose computer, such as an array, programmable logic unit (PLU), microprocessor, or any other device capable of executing and responding to instructions.
- the processing device may execute an operating system (OS) and software applications running on the operating system. Additionally, a processing device may access, store, manipulate, process, and generate data in response to the execution of software.
- OS operating system
- a processing device may access, store, manipulate, process, and generate data in response to the execution of software.
- a single processing device may be described as being used; however, those skilled in the art will understand that a processing device includes multiple processing elements and/or multiple types of processing elements. It can be seen that it may include.
- a processing device may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are possible.
- Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing unit to operate as desired, or may be processed independently or collectively. You can command the device.
- Software and/or data may be used on any type of machine, component, physical device, virtual equipment, computer storage medium or device to be interpreted by or to provide instructions or data to a processing device. , or may be permanently or temporarily embodied in a transmitted signal wave.
- Software may be distributed over networked computer systems and stored or executed in a distributed manner.
- Software and data may be stored on a computer-readable recording medium.
- the method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium.
- a computer-readable medium can store program instructions, data files, data structures, etc., singly or in combination, and the program instructions recorded on the medium may be specially designed and constructed for the embodiment or may be known and available to those skilled in the art of computer software. It may be possible.
- Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, and magnetic media such as floptical disks.
- Examples of program instructions include machine language code, such as that produced by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.
- the hardware devices described above may be configured to operate as one or multiple software modules to perform the operations of the embodiments, and vice versa.
- the described techniques are performed in a different order than the described method, and/or components of the described system, structure, device, circuit, etc. are combined or combined in a different form than the described method, or other components are used.
- appropriate results may be achieved even if substituted or substituted by an equivalent. Therefore, other implementations, other embodiments and equivalents of the claims also fall within the scope of the claims described below.
- One or more processors are One or more processors.
- Comprising one or more memories storing instructions that cause the processor to perform an operation when executed by the processor
- the processor When the instructions are executed by the processor, the processor:
- a training video consisting of a sequence of image frames is sampled into multiple frames
- An electronic device configured to label the plurality of object images and train a learning model to classify each of the one or more objects.
- the processor executes
- the electronic device is further configured to train the learning model based on the image group.
- Detecting the one or more objects and generating a plurality of object images corresponding to the one or more objects includes grouping regions with similar values within a frame based on at least one of color, texture, or pattern into regions of interest, and An electronic device comprising generating an image of an object by separating an area from a background.
- the object is an electronic device corresponding to at least one of a character, person, specific shape, costume, building, or logo.
- the processor selects the first object at a second viewpoint different from the first viewpoint.
- the electronic device is configured to generate an image of the second object from the first frame using the first frame and a second frame in which appears.
- the processor executes
- An electronic device configured to train the learning model based on images of one or more objects corresponding to objects in which the number of images is a predetermined number or more.
- the processor executes
- the electronic device further configured to sample a frame in which the specific object appears.
- the processor is configured to sample the learning image according to a predetermined cycle or to sample a frame corresponding to a specific point in time based on an input.
- the processor generates additional images through an image processing model when the number of images for a specific object for training the learning model is less than a predetermined number, and creates a learning model for the specific object with the additional images for the specific object.
- An electronic device further configured to train.
- the electronic device wherein the image processing model performs at least one of resizing, changing color, cropping, or rotating an image for an object.
- the electronic device is further configured to train the learning model by at least one of linear discriminant analysis, K Nearest Neighbor (KNN), decision tree, neural network, and Support Vector Machine (SVM) methods.
- KNN K Nearest Neighbor
- SVM Support Vector Machine
- One or more processors are One or more processors;
- the learning model is a model in which the correlation between objects and labels included in the learning image is learned
- the processor executes
- Input a plurality of images obtained from the input target image into the learning model to obtain similarity for each of one or more labels
- the electronic device is configured to calculate the total similarity of the target image to the training image by multiplying the similarity of each of the one or more labels by a weight of each of the one or more labels.
- the processor executes
- the learning model is configured to compare the plurality of object images with images corresponding to each of the one or more labels.
- the learning model is further configured to perform at least one of size change, color change, cropping, or rotation on the plurality of object images and then compare them with images corresponding to each of the one or more labels.
- the learning model is,
- An electronic device configured to calculate similarity by comparing images corresponding to each of the one or more labels using the feature points.
- the similarity is calculated based on the number of matching feature points between the feature point and feature points of images corresponding to each of the one or more labels.
- the similarity is calculated using a neural network model through a softmax function for each of the one or more labels so that the sum of the similarities for the one or more labels is 1.
- the weight for a specific label among the one or more labels is determined according to the frequency of objects corresponding to the specific label or determined by receiving an input.
- the processor executes
- the electronic device is further configured to restrict uploading of the target image if the total similarity is greater than or equal to a threshold.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Physics & Mathematics (AREA)
- Multimedia (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Computing Systems (AREA)
- Software Systems (AREA)
- Medical Informatics (AREA)
- Evolutionary Computation (AREA)
- Databases & Information Systems (AREA)
- Artificial Intelligence (AREA)
- Psychiatry (AREA)
- Social Psychology (AREA)
- Human Computer Interaction (AREA)
- Computational Linguistics (AREA)
- Image Analysis (AREA)
Abstract
본 개시의 다양한 실시예에 따른 전자 장치는 하나 이상의 프로세서; 및 프로세서에 의한 실행 시 프로세서가 연산을 수행하도록 하는 명령어들이 저장된 하나 이상의 메모리를 포함하고, 프로세서에 의해 상기 명령어들이 실행될 시, 프로세서는, 이미지 프레임의 시퀀스로 구성된 학습 영상을 복수의 프레임으로 샘플링하고, 복수의 프레임 각각에 나타난 하나 이상의 객체를 검출하여 상기 하나 이상의 객체에 대응하는 객체 이미지들을 생성하고, 객체 이미지들을 레이블링(labeling)하여 하나 이상의 객체 각각으로 분류되도록 학습 모델을 훈련하도록 구성될 수 있다.
Description
본 개시는 학습 모델을 이용하여 영상들 간의 유사도를 판단하기 위한 기술에 관한 것이다.
웹상에서 다양한 컨텐츠가 자유롭게 배포 및 유통됨에 따라 컨텐츠에 대한 저작물 보호가 중요한 문제가 되었다. 컨텐츠는 복제, 유통 및 배포가 매우 용이하며 이와 같이 배포된 컨텐츠는 진본과 실질적으로 동일하므로 컨텐츠에 의한 저작권 침해는 저작권자의 권익을 심각하게 훼손하게 된다.
컨텐츠에 대한 저작권 보호를 위한 방안은 저작물의 복제, 유통, 배포가 어렵게 하는 사전적 조치와, 불법으로 복제, 유통, 배포된 저작물에 대해 검출하고 단속하는 사후적 조치로 나눌 수 있다. 사전적 조치는 예컨대 복제가 불가능하거나 복제 회수를 제한하기 위한 워터마킹 기술 등과 같이 기술적 측면에서 개발되는 방식으로서 많은 발전이 이루어져 왔다. 그러나 사전적 조치에 의한 방식은 제한을 해제하는 기술의 개발에 의해 대부분 무력화되고 있으며, 또한 저작물의 직접적 침해에 해당되지 않는 복제 등에 대해서도 구분을 두지 않고 금지하는 효과로 인해 현실적으로 적용이 부적절한 경우가 많다. 따라서 사후적 조치로서 저작권의 침해를 하고 있는 행위에 대해 검출 및 적발이 지속적으로 병행되어야 한다.
그런데 현재의 컨텐츠 저작물에 대한 침해 검출은 저작권자 스스로 또는 저작권을 위탁 받아 관리하는 위탁기관이 개별적으로 웹사이트들을 접속하여 검출하는 수작업에 의존하고 있다. 이러한 방식은 무수히 많은 웹사이트들에 대한 검출 및 적발을 매우 어렵게 하며, 기 검출된 웹사이트에서도 새로이 저작권 침해 사례가 추가되는 경우에 재접속 및 재검출을 하지 않는 한 지속적인 감시가 어렵게 된다. 나아가, 저작권자가 소자본 개인인 경우 위탁 기관에 자신의 저작물에 대한 권리 보호를 위탁시키는 것조차도 쉽지 않은 경우가 많다.
본 개시의 일 실시예를 통해 해결하고자 하는 기술적 과제는 학습 영상을 통해 영상들 간의 유사도를 판단할 수 있는 학습 모델을 훈련시키는 기술을 제공하는 것이다.
본 개시의 일 실시예를 통해 해결하고자 하는 기술적 과제는 학습 모델을 이용하여 영상에 나타난 객체들의 유사도를 판단하여 영상들 간의 유사도를 판단하는 기술을 제공하는 것이다.
본 개시의 한 측면으로서, 전자 장치가 제안될 수 있다. 본 개시의 한 측면에 따른 전자 장치는, 하나 이상의 프로세서, 및 하나 이상의 메모리를 포함할 수 있다. 상기 프로세서는 이미지 프레임의 시퀀스로 구성된 학습 영상을 복수의 프레임으로 샘플링하고, 상기 복수의 프레임 각각에 나타난 하나 이상의 객체를 검출하여 상기 하나 이상의 객체에 대응하는 복수의 객체 이미지를 생성하고, 상기 복수의 객체 이미지를 레이블링하여 상기 하나 이상의 객체 각각으로 분류되도록 학습 모델을 훈련하도록 구성될 수 있다.
본 개시의 한 측면으로서, 전자 장치가 제안될 수 있다. 본 개시의 한 측면에 따른 전자 장치는, 하나 이상의 프로세서, 네트워크 인터페이스, 및 하나 이상의 메모리를 포함할 수 있다. 상기 메모리는 학습 모델을 저장할 수 있고, 학습 모델은 학습 영상에 포함된 객체와 레이블의 상관 관계가 학습된 모델일 수 있다. 상기 프로세서는 상기 메모리로부터 상기 학습 모델을 로드하고, 입력된 타겟 영상으로부터 획득한 복수의 이미지를 상기 학습 모델에 입력하여 하나 이상의 레이블 각각에 대한 유사도를 획득하고, 상기 하나 이상의 레이블 각각에 대한 유사도에 상기 하나 이상의 레이블 각각의 가중치를 곱하여 상기 타겟 영상의 상기 학습 영상에 대한 총 유사도를 산출하도록 구성될 수 있다.
본 개시의 한 측면으로서, 방법이 제안될 수 있다. 본 개시의 한 측면에 따른 방법은, 입력된 학습 영상을 복수의 프레임으로 샘플링하는 단계, 상기 복수의 프레임 각각에 나타난 하나 이상의 객체를 검출하여 상기 하나 이상의 객체에 대응하는 복수의 객체 이미지를 생성하는 단계, 상기 복수의 객체 이미지를 레이블링하여 상기 하나 이상의 객체 각각으로 분류되도록 학습 모델을 훈련시키는 단계, 입력된 타겟 영상으로부터 획득한 복수의 이미지를 상기 학습 모델에 입력하여 하나 이상의 레이블 각각에 대한 유사도를 획득하는 단계, 및 상기 하나 이상의 레이블 각각에 대한 유사도에 상기 하나 이상의 레이블 각각의 가중치를 곱하여, 상기 타겟 영상의 상기 학습 영상에 대한 총 유사도를 산출하는 단계를 포함할 수 있다.
본 개시의 다양한 실시예에 따르면, 학습 영상을 통해 학습 모델이 이미지를 분류하도록 훈련시킬 수 있다.
본 개시의 다양한 실시예에 따르면, 학습 영상의 이미지를 다양하게 변형시켜서 학습 모델에 훈련시킴으로써 학습 모델의 정확도를 향상시킬 수 있다.
본 개시의 다양한 실시예에 따르면, 학습 모델에 미리 학습된 영상에 대한 입력된 영상의 유사도를 판단할 수 있다. 따라서, 입력된 영상이 특정 영상과 유사한지 여부를 판단할 수 있다.
본 개시의 다양한 실시예에 따르면, 학습 모델에 저작권이 있는 영상을 미리 입력하여 지도 훈련시킴으로써, 특정한 영상을 업로드하기 전에 해당 영상이 저작권을 침해하는 지 여부를 미리 판단할 수 있다.
본 개시의 다양한 실시예에 따르면, 저작권 침해가 문제될 수 있는 영상에 대해서는 업로드를 방지함으로써 저작권 침해를 방지할 수 있다.
본 개시의 다양한 실시예에 따르면, 영상의 유사 여부를 판단하여 이를 기초로 저작권 침해 여부를 판단함으로써 저작권 침해 사례 검출을 간편하게 할 수 있다.
도 1은 본 개시의 일 실시예에 따른 영상의 유사 여부 판단을 위한 공유 플랫폼 시스템을 예시적으로 도시한 도면이다.
도 2는 본 개시의 일 실시예에 따른 전자 장치의 블록도이다.
도 3은 본 개시의 일 실시예에 따른 학습 영상을 샘플링하는 모습을 예시적으로 도시한 도면이다.
도 4a, 도 4b는 본 개시의 일 실시예에 따른 학습 모델 훈련 모습을 예시적으로 도시한 도면이다.
도 5는 본 개시의 일 실시예에 따른 이미지 그룹의 생성 방법을 도시한 도면이다.
도 6은 본 개시의 일 실시예에 따른 유사도 비교 방법을 도시한 도면이다.
도 7a, 도 7b는 본 개시의 일 실시예에 따른 유사도 판단 방법을 예시적으로 도시한 도면이다.
도 8은 본 개시의 일 실시예에 따른 입력 영상이 학습 영상과 유사한지를 판단하는 방법에 대한 흐름도이다.
본 개시에 기재된 다양한 실시예들은, 본 개시의 기술적 사상을 명확히 설명하기 위한 목적으로 예시된 것이며, 이를 특정한 실시 형태로 한정하려는 것이 아니다. 본 개시의 기술적 사상은, 본 개시에 기재된 각 실시예의 다양한 변경(modifications), 균등물(equivalents), 대체물(alternatives) 및 각 실시예의 전부 또는 일부로부터 선택적으로 조합된 실시예를 포함한다. 또한 본 개시의 기술적 사상의 권리 범위는 이하에 제시되는 다양한 실시예들이나 이에 대한 구체적 설명으로 한정되지 않는다.
개시에 사용되는 모든 기술적 용어들 및 과학적 용어들은, 달리 정의되지 않는 한, 본 개시가 속하는 기술 분야에서 통상의 지식을 가진 자에게 일반적으로 이해되는 의미를 갖는다. 본 개시에 사용되는 모든 용어들은 본 개시를 더욱 명확히 설명하기 위한 목적으로 선택된 것이며 본 개시에 따른 권리범위를 제한하기 위해 선택된 것이 아니다.
본 개시에서 사용되는 "포함하는", "구비하는", "갖는" 등과 같은 표현은, 해당 표현이 포함되는 어구 또는 문장에서 달리 언급되지 않는 한, 다른 실시예를 포함할 가능성을 내포하는 개방형 용어(open-ended terms)로 이해되어야 한다.
본 개시에서 기술된 단수형의 표현은 달리 언급하지 않는 한 복수형 의 의미를 포함할 수 있으며, 이는 청구의 범위에 기재된 단수형의 표현에도 마찬가지로 적용된다.
본 개시에서 사용되는 "제1", "제2" 등의 표현들은 복수의 구성요소들을 상호 구분하기 위해 사용되며, 해당 구성요소들의 순서 또는 중요도를 한정하는 것은 아니다.
본 개시에서 사용되는 "~에 기초하여"라는 표현은, 해당 표현이 포함되는 어구 또는 문장에서 기술되는, 결정, 판단의 행위 또는 동작에 영향을 주는 하나 이상의 인자를 기술하는데 사용되며, 이 표현은 결정, 판단의 행위 또는 동작에 영향을 주는 추가적인 인자를 배제하지 않는다.
본 개시에서, 어떤 구성요소가 다른 구성요소에 "연결되어" 있다거나 "접속되어" 있다고 언급된 경우, 상기 어떤 구성요소가 상기 다른 구성요소에 직접적으로 연결될 수 있거나 접속될 수 있는 것으로, 또는 새로운 다른 구성요소를 매개로 하여 연결될 수 있거나 접속될 수 있는 것으로 이해되어야 한다.
본 개시에서 사용된 표현 "~하도록 구성된(configured to)"은 문맥에 따라, "~하도록 설정된", "~하는 능력을 가지는", "~하도록 변경된", "~하도록 만들어진", "~를 할 수 있는" 등의 의미를 가질 수 있다. 이 표현은, "하드웨어적으로 특별히 설계된"의 의미로 제한되지 않으며, 예를 들어 특정 동작을 수행하도록 구성된 프로세서란, 소프트웨어를 실행함으로써 그 특정 동작을 수행할 수 있는 범용 프로세서(generic purpose processor)를 의미하거나, 그 특정 동작을 수행하도록 프로그래밍을 통해 구조화된 특수 목적 컴퓨터(special purpose computer)를 의미할 수 있다.
이하, 첨부한 도면들을 참조하여, 본 개시의 실시예들을 설명한다. 첨부된 도면에서, 동일하거나 대응하는 구성요소에는 동일한 참조번호가 부여되어 있다. 또한, 이하의 실시예들의 설명에 있어서, 동일하거나 대응하는 구성요소를 중복하여 기술하는 것이 생략될 수 있다. 그러나, 구성요소에 관한 기술이 생략되어도, 그러한 구성요소가 어떤 실시예에 포함되지 않는 것으로 의도되지는 않는다.
도 1은 본 개시의 일 실시예에 따른 영상의 유사 여부 판단을 위한 공유 플랫폼 시스템을 도시한 도면이다. 공유 플랫폼 시스템(100)은 전자 장치(103), 학습 모델(104), 서버(105)를 포함할 수 있다.
일 실시예에서 사용자(101)는 사용자 장치(미도시)를 사용하여 컨텐츠를 생성할 수 있다. 컨텐츠는 영상, 이미지, 음악 등을 포함할 수 있다. 영상 파일이란, 움직이는 영상물을 총칭하는 것으로서, 예를 들어, avi, mp4 또는 mkv의 확장자를 갖는 파일을 의미할 수 있다.
일 실시예에서, 사용자(101)는 영상(102)을 다른 사용자들과 공유하기 위해 공유 플랫폼 시스템(100)에 업로드할 수 있다. 예를 들어, 제1 사용자는 영상(102)을 공유 플랫폼 시스템(100)에 업로드할 수 있고, 제2 사용자는 제1 사용자가 입력한 영상(102)을 공유 플랫폼 시스템(100)에서 열람할 수도 있다.
일 실시예에서, 전자 장치(103)는 입력된 영상(102)을 서버(105)에 업로드할 수 있고, 서버(105)에 업로드하기 전 다양한 처리를 할 수도 있다. 복수의 구성 요소가 실제 물리적 환경에서는 서로 통합되는 형태로 구현될 수 있고, 필요에 따라 일부 구성 요소가 추가되거나 삭제될 수 있다.
본 개시의 전자 장치(103)는 다양한 형태의 장치가 될 수 있다. 예를 들어, 전자 장치(103)는 컴퓨팅 장치 또는 휴대용 통신 장치일 수 있다. 전자 장치(103)는 사용자(101)와 통신 연결되어 다양한 데이터를 송수신할 수 있다. 전자 장치(103)와 사용자(101)는 통신시 근거리 통신망(Local Area Network; LAN), 광역 통신망(Wide Area Network; WAN) 또는 부가가치 통신망(Value Added Network; VAN) 등과 같은 유선 네트워크나, 이동 통신망(mobile radio communication network), 위성 통신망, 블루투스(Bluetooth), Wibro(Wireless Broadband Internet), HSDPA(High Speed Downlink Packet Access) 등과 같은 모든 종류의 무선 네트워크를 통해 실행될 수 있다.
영상(102)이 다른 사람의 저작권과 관련되어 저작권을 침해할 가능성이 있는 경우, 전자 장치(103)는 영상(102)의 업로드를 금지할 수 있다. 예를 들어, 전자 장치(103)는 영상(102)을 수신하여 서버(105) 또는 다른 저장 장치(미도시)에 업로드하기 전에 다른 영상과의 유사 여부를 학습 모델(104)을 통해 확인할 수 있다. 입력되는 영상은 비교되는 다른 영상과 구분하여, 입력 영상 또는 타겟 영상 등으로 지칭될 수 있다.
학습 모델(104)은 전자 장치(103) 내에 통합되어 구현될 수도 있고, 별개로 분리되어 통신할 수도 있다. 전자 장치(103)는 영상(102)이 입력 되기 전에 학습 모델(104)을 미리 훈련(학습)시킬 수 있다. 학습 모델(104)은 이미지를 분류하기 위한 모델로, 다수의 이미지를 이용하여 학습될 수 있다. 이하에서는 설명의 편의를 위해 학습 모델(104)이 전자 장치(103)와 별도로 구현된 것으로 가정한다.
전자 장치(103)는 추후 명세서의 기재에 의해 구체화될 학습 영상(111)을 운영자(110)로부터 수신할 수 있다. 또는, 전자 장치(103)는 학습 영상(111)을 네트워크, 서버 등으로부터 다운로드 받을 수 있다. 학습 영상(111)은 시계열적으로 연속하는 다수의 프레임을 포함할 수 있다.
수신된 학습 영상(111)을 기초로 전자 장치(103)는 학습 모델(104)을 훈련시킬 수 있다. 전자 장치(103)는 학습 영상(111)을 복수의 프레임으로 샘플링할 수 있고, 샘플링된 복수의 프레임에 나타난 하나 이상의 객체를 검출하여 복수의 객체 이미지를 생성할 수 있다.
전자 장치(103)는 생성한 복수의 객체 이미지를 레이블링하여 학습 모델(104)을 훈련시킬 수 있다. 학습 모델(104)은 객체의 이미지와 객체의 레이블을 연관지어 상관 관계가 훈련될 수 있다. 학습 모델(104)은 훈련 후에 특정한 이미지에 포함된 객체를 보고, 정해진 레이블 중 하나로 분류할 수 있다.
학습이 된 학습 모델(104)은 영상 사이의 유사도를 판단할 수 있다. 전자 장치(103)는 학습 모델(104)이 판단한 영상 사이의 유사도를 기초로 저작권 침해 여부를 판단할 수 있다. 예를 들어, 학습 모델(104)은 사용자(101)가 업로드한 영상(102)의 객체에 대한 이미지들이 학습 영상(111) 속 객체의 이미지들과 얼마나 유사한지를 판단하여 영상(102)과 학습 영상(111)의 총 유사도를 산출할 수 있다. 전자 장치(103)는 산출된 총 유사도에 기초하여 영상(102)의 업로드 여부를 결정할 수 있다.
도 2는 본 개시의 일 실시예에 따른 전자 장치의 블록도이다. 도 2에서는 설명의 편의를 위해 전자 장치의 참조번호를 '200'으로 기재하였지만 전자 장치(103)과 전자 장치(200)은 서로 다른 장치 구성이 아니며 동일하거나 대응하는 장치일 수 있다. 도 2를 참조하면, 본 개시의 일 실시예에 따른 전자 장치(200)는 사용자에 의해 입력되는 영상(102)의 학습 영상(111)에 대한 유사 여부를 모니터링하여 저작권 침해 판단을 할 수 있는 서비스를 제공할 수 있다. 전자 장치(200)는 적어도 하나의 프로세서(210), 적어도 하나의 메모리(220), 및 적어도 하나의 네트워크 인터페이스(230)를 포함할 수 있다.
프로세서(210)는 메모리(220) 및/또는 네트워크 인터페이스(230)를 제어하며, 메모리(220)에 저장된 명령어를 실행하여 본 개시의 설명, 기능, 절차, 제안, 방법 및/또는 동작 순서도를 구현하도록 구성될 수 있다. 예를 들어, 프로세서(210)는 네트워크 인터페이스(230)를 통해 신호 및/또는 데이터를 수신하고, 신호 및/또는 데이터에 포함된 정보를 메모리(220)에 저장할 수 있다.
메모리(220)는 프로세서(210)와 연결될 수 있고, 프로세서(210)의 동작과 관련한 다양한 정보를 저장할 수 있다. 예를 들어, 메모리(220)는 프로세서(210)에 의해 제어되는 프로세스들 중 일부 또는 전부를 수행하거나, 본 개시의 설명, 기능, 절차, 제안, 방법 및/또는 동작 순서도들을 수행하기 위한 명령어들(instructions)을 포함하는 소프트웨어 코드를 저장할 수 있다. 메모리(220)는 비일시적인 컴퓨터 판독가능 매체, 예컨대 고속 랜덤 액세스 메모리 및/또는 비휘발성 컴퓨터 판독가능 저장 매체(예컨대, 하나 이상의 디스크 저장 장치, 플래쉬 메모리 장치, 또는 기타 비휘발성 솔리드 스테이트 메모리 장치)를 포함할 수 있다. 학습 모델(104)이 전자 장치(200) 내에 구현되는 경우 학습 모델(104)은 메모리(220) 내에 저장되어 구현될 수 있다.
네트워크 인터페이스(230)는 프로세서(210)와 연결될 수 있고, 유/무선 신호나 데이터를 전송 및/또는 수신할 수 있다. 예를 들어, 네트워크 인터페이스(230)는 유/무선 통신망을 통해 사용자 단말 및/또는 서버, 데이터베이스와 연결될 수 있다.
여기서, 무선 통신망은 이동 통신망, 무선 LAN, 근거리 무선 통신망 등을 포함할 수 있다. 예를 들어, 무선 통신망은 LTE, LTE-A(LTE Advance), CDMA(code division multiple access), WCDMA(wideband CDMA), UMTS(universal mobile telecommunications system), WiBro(Wireless Broadband), 및 GSM(Global System for Mobile Communications) 등 중 적어도 하나를 사용하는 셀룰러 통신을 포함할 수 있다. 예를 들어, 무선 통신망은 WiFi(wireless fidelity), 블루투스, 블루투스 저전력(BLE), 지그비 (Zigbee), NFC(near field communication), 및 라디오 프리퀀시(RF) 중 적어도 하나를 포함할 수 있다.
여기서, 유선 통신망은 USB(Universal Serial Bus), USART(Universal Synchronous/Asynchronous Receiver Transmitter), 및 이더넷(ethernet) 중 적어도 하나를 포함할 수 있다.
적어도 하나의 네트워크 인터페이스(230) 각각은 상술한 유/무선 통신망에 대응될 수 있다.
이하에서 개시되는 전자 장치(200)에 포함되는 기능 단위들은 상술한 프로세서(210), 메모리(220) 및 네트워크 인터페이스(230)를 포함하는 하드웨어나 명령어들을 구현하기 위한 소프트웨어 또는 하드웨어 및 소프트웨어의 결합으로 구현될 수 있다.
도 3은 일 실시예에 따른 프로세서가 학습 영상을 샘플링(sampling)하는 모습을 도시한 도면이다. 도 3을 참조하면 프로세서(210)는 수신한 영상(102)을 정해진 샘플링 주기로 샘플링할 수 있다. 샘플링은, 영상으로부터 프레임을 획득하는 것을 의미할 수 있다.
일 실시예에서, 샘플링의 주기는 미리 메모리(220)에 저장되어 있을 수 있고, 운영자(110)에 의해 입력을 받아 결정될 수 있다. 운영자(110)의 입력에 의해 샘플링 시점이 결정되는 경우, 샘플링의 주기는 주기적으로 이루어질 수 있고, 입력에 의해 비주기적으로 특정 시점의 프레임이 샘플링될 수 있다. 전체 프레임이 아닌 샘플링 된 프레임이 대해서만 처리를 함으로써 컴퓨팅 로드를 감소시킬 수 있다.
일 실시예에서, 프로세서(210)는 영상(102)에 있는 객체의 움직임에 대한 움직임 벡터를 취득하고 정해진 시간 동안 움직임 벡터의 크기를 측정할 수 있다. 정해진 시간 동안 움직임 벡터의 크기가 일정한 값 이상이면 객체의 움직임이 크다는 것을 의미할 수 있다. 예를 들어 움직임 벡터는 두드러진 움직임이 있는 물체를 배경에서 분리하기 위한 방법으로 가우시안 혼합 모델을 통해 생성된 배경 모델에 적응적 차영상을 통한 방법으로 구해질 수 있다. 정해진 시간 동안 움직임 벡터의 변화량이 제1 값 이상이면, 영상에서 객체의 중요성이 높다는 것을 의미할 수 있다. 프로세서(210)는 객체의 움직임에 대한 움직임 벡터의 변화량이 제1 값 이상이면 객체를 추적(tracking)해서 객체가 등장하는 프레임을 샘플링할 수 있다. 프로세서(210)는 객체를 추적하기 위해 점 추적, 커널(kernel) 추적, 실루엣 추적 등을 이용할 수 있다. 따라서, 프로세서(210)는 유의미한 프레임에 대해서만 샘플링을 함으로써, 컴퓨팅 로드를 감소시킬 수도 있다.
일 실시예에 따른 프로세서(210)는 영상(300)을 샘플링하여 서로 다른 시간에 대응하는 복수의 프레임(310)을 획득할 수 있다. 샘플링된 시간 T1, T2, T3, T4 사이의 간격은 일정할 수도 있고, 일정하지 않은 임의의 시간 간격일 수 있다. 복수의 프레임 중 각각의 프레임(311, 312, 313, 314)은 그 샘플링 된 시간에서의 영상의 프레임을 나타낸다.
도 4a, 도 4b는 본 개시의 일 실시예에 따른 학습 모델 훈련 모습을 예시적으로 도시한 도면이다. 도 4a에서는 설명의 편의를 위해 학습 모델의 참조번호를 '414'로 기재하였지만, 학습 모델(104)와 학습 모델(414)는 서로 다른 모델이 아니며 동일하거나 대응하는 학습 모델일 수 있다.
전술한 바와 같이, 프로세서(210)는 수신된 영상(102)으로부터 샘플링을 통해 복수의 프레임(311, 312, 313, 314)을 획득할 수 있다.
먼저 도 4a를 참조하면, 프레임 중 하나의 프레임(410)은 객체(411)를 포함할 수 있다. 예를 들어 객체(411)는 캐릭터, 인물, 특정한 형상, 의상, 건물, 또는 로고일 수 있다. 프로세서(210)는 프레임에서 객체(411)를 검출하기 위해 객체가 있는 영역(이하, 관심 영역(Region of Interest, ROI))을 식별(identify)할 수 있다.
일 실시예에서, 프로세서(210)는 프레임에서 색상, 질감, 패턴 중 적어도 어느 하나에 기초하여 유사한 영역을 찾고, 유사한 영역을 기초로 그룹화하여 ROI를 설정할 수 있다. 예를 들어, 객체(411)의 경계를 기초로 RGB 값에 급격한 차이가 날 수 있다. 즉, 프로세서(210)는 색이 비슷한 영역들을 유사한 영역으로 설정하고 영역화하여 ROI로 식별할 수 있다. 프로세서(210)는 유사한 영역을 판단하기 위해 프레임에서 색상값의 미분을 취하고, 미분값이 특정 임계값을 넘는 부분을 경계로 보고 유사한 영역을 그룹화할 수 있다. 이 경우, 유사 영역을 기반으로 ROI가 간편하게 설정될 수 있다.
예를 들어, 프로세서(210)는 객체의 윤곽을 검출하기 위해 소벨(Sobel) 필터를 사용하여 미분 연산자에 의한 픽셀의 밝기 값의 변화 또는 색상(RGB) 값의 변화를 이용하여 윤곽을 찾아낼 수 있다. 소벨 필터는 비선형 연산자로서 사용하는 마스크 창(Mask Window) 영역의 양 끝단에 속한 픽셀들 사이의 RGB 합의 차이를 구한 후 이를 수평과 수직 방향에 대하여 평균 크기를 구함으로써 윤곽 부위를 강조하는 기능을 수행한다. 소벨 필터의 기능에 대하여는 공지되어 있으므로 본 명세서에서는 상세한 설명을 생략한다.
일 실시예에서, 프로세서(210)는 프레임을 다수의 영역으로 분류하고 영역별로 질감 특징을 히스토그램화 하여 히스토그램의 값이 일정 크기 범위인 부분들을 하나의 영역으로 그룹화하여 ROI를 식별할 수도 있다.
일 실시예에서, 프로세서(210)는 운영자(110)로부터 입력을 받아서 ROI를 설정 할 수도 있다. 이는 운영자(110)가 관심 있어하는 영역, 즉 객체가 있다고 판단되는 영역을 설정하면 그 부분으로 ROI가 설정될 수 있다.
일 실시예에서, 프로세서(210)는 설정된 ROI를 기초로 배경으로부터 객체를 분리할 수 있다. 프로세서(210)는 ROI를 식별하여 배경과 객체를 분리하는 경우 ROI를 기초로 박스(411)를 생성하고, 박스(411)를 기초로 이미지를 크롭(crop)할 수 있다. 크롭된 이미지의 사이즈는 상이할 수 있다.
일 실시예에서 프로세서(210)는 크롭된 이미지들의 사이즈를 일정한 사이즈로 조정하는 처리를 추가적으로 할 수 있다. 예를 들어, 학습 데이터를 위해 정해진 사이즈로 이미지가 조절될 수 있다.
일 실시예에서 프로세서(210)는 크롭된 이미지들이 특정한 객체에 대응하는 이미지인지 여부에 따라 훈련 전 사전 분류를 할 수 있다. 예를 들어, 프로세서(210)는 이미지들 사이의 특징점을 찾아 유사한 이미지들끼리 그룹을 생성할 수 있다. 프로세서(210)는 특정 객체를 지시하는 이미지 그룹을 생성하고 이미지 그룹을 기초로 학습 모델(104)을 훈련시킬 수 있다. 이미지의 그룹 생성에 대해서는 이하의 도 5에서 자세히 설명한다.
다른 실시예에서, 프로세서(210)는 프레임으로부터 ROI를 식별하여 이미지를 취득하는 것뿐만 아니라 이미지를 가공하거나 추가 이미지를 생성할 수도 있다.
일 실시예에서, 크롭된 이미지에 복수의 객체가 오버랩(overlap) 되어 있을 수 있다. 예를 들어, 복수의 객체로 제1 객체, 제2 객체가 오버랩 되어 있을 수 있다. 여기서 프로세서(210)는 해당 프레임과 다른 프레임에서 제1 객체에 대한 이미지를 획득할 수 있다. 프로세서(210)는 다른 시점의 프레임에서의 제1 객체에 대한 이미지에 기초하여 해당 프레임의 이미지를 분리할 수 있다. 예를 들어, 다른 시점은 해당 프레임 시점 이전의 시점일 수 있고, 해당 프레임 시점 이후 일부 시간 경과 후의 시점일 수도 있다. 예를 들어, 프로세서(210)는 이전 시점의 프레임에서의 제1 객체에 대한 이미지를 해당 프레임으로부터 분리하고 남은 부분을 제2 객체의 이미지로 생성하도록 할 수 있다. 이미지를 분리한다는 것은 해당 이미지에 대한 색을 제거하거나, 해당 이미지의 레이어를 구분하는 것을 의미할 수 있다. 이 경우, 객체가 겹쳐져 있는 경우에서도 다른 시점의 객체 이미지를 통해 각 객체를 분리해내어 훈련을 위한 이미지를 다양하게 생성할 수 있다.
일 실시예에서 프로세서(210)는 이미지를 병합하여 학습 데이터를 생성할 수도 있다. 제1 프레임에서 얻어진 객체의 일부 이미지와 제2 프레임에서 획득된 객체의 일부 이미지를 병합할 수 있다. 예를 들어, 프로세서(210)는 이미지가 연속성이 있다고 판단하는 경우, 또는 이미지에 공통된 형상이 있다고 판단하는 경우 이미지를 병합하여 학습 데이터로 생성할 수 있다.
프로세서(210)는 학습 영상(111)으로부터 샘플링되고, 크롭되어 얻어진 복수의 이미지 및 추가적인 처리를 거친 이미지들을 기초로 학습 모델(414)을 훈련시킬 수 있다.
따라서, 학습 모델(414)을 훈련시킬 때, 운영자(110)가 방대한 양의 학습 이미지를 따로 입력할 필요 없이 학습 영상(111) 하나를 입력하는 것만으로 충분히 훈련시킬 수 있다는 효과를 갖을 수 있다.
일 실시예에서 프로세서(210)는 다양한 객체에 대한 이미지들을 기초로 학습 모델(414)을 훈련시킬 수 있다. 프로세서(210)는 학습 모델(414)에 지도 학습(supervised learning)을 시킬 수 있다. 지도 학습이란, 정해진 레이블(label)에 따라 정해진 이미지들을 모델에 훈련시켜 모델이 특정 이미지를 특정 레이블로 분류할 수 있도록 훈련시키는 것을 의미한다.
일 실시예에서, 프로세서(210)는 이미지를 훈련시킬 때마다 운영자(110)로부터 이미지에 대한 레이블을 입력받을 수 있다. 예를 들어, 프로세서(210)는 디스플레이(미도시)를 통해 운영자가 크롭된 이미지 중 하나를 확인할 수 있도록 할 수 있다. 운영자(110)는 이미지를 보고 해당하는 레이블을 타이핑(typing), 클릭(click) 등의 방식으로 입력할 수 있다. 학습 모델(414)은 각각의 이미지에 따른 레이블을 입력 받아 지도 학습될 수 있다.
다른 실시예에서, 프로세서(210)는 유사하게 분류된 이미지들에 대해 임의의 레이블로 분류하여 학습 모델(414)을 훈련시킬 수도 있다. 즉, 이미지 그룹별로 레이블이 정해질 수 있고, 동일한 이미지 그룹에 속한 이미지들을 동일한 레이블로 학습 모델(414)을 훈련시킬 수 있다. 예를 들어, 프로세서(210)는 복수의 이미지(412, 413, 414)를 동일한 레이블에 해당한다고 분류하고 이미지들(412, 413, 414)을 "Label 1"으로 레이블링을 통해 학습 모델(414)을 훈련시킬 수 있다. 예를 들어, 학습 모델(414)은 복수의 이미지(412, 413, 414)와 유사한 이미지는 "Label 1"이라고 분류할 수 있도록 학습될 수 있다.
일 실시예에서, 프로세서(210)는 메모리(220)에 미리 저장되어 있는 레이블과 각 레이블에 해당하는 이미지를 로드(load)하여 훈련에 사용할 수 있다. 프로세서(210)는 로드한 레이블에 해당하는 이미지를 찾아내고, 레이블에 따라 1차적으로 분류를 진행한 뒤 분류된 이미지들을 학습 모델(414)의 학습 데이터로 사용할 수 있다.
도 4b를 참조하면, 프로세서(210)는 특정한 객체의 이미지가 추출되는 확률을 계산하고 이를 기초로 학습 대상을 결정할 수도 있다. 예를 들어, 프로세서(210)는 객체 A에 대한 이미지(422)가 전체 객체 이미지에서 차지하는 비율을 계산할 수 있다. 또는, 프로세서(210)는 획득한 이미지들 중 객체 A에 대한 이미지(422)에 해당하는 이미지의 개수를 계산할 수 있다. 예를 들어, 운영자로부터 레이블을 입력받은 경우에는, 프로세서(210)는 특정 레이블에 해당하는 이미지의 개수를 계산할 수 있다. 예를 들어, 프로세서(210)가 유사한 이미지끼리 분류하여 이미지 그룹을 만드는 경우라면, 프로세서(210)는 해당하는 이미지 그룹에 속하는 이미지의 개수를 계산할 수 있다.
일 실시예에서, 프로세서(210)는 특정한 객체의 이미지에 해당하는 이미지의 개수가 정해진 개수 이상이면 해당하는 특정한 객체를 학습 대상으로 선택할 수 있다. 학습 대상으로 선택한다는 것은, 학습 모델(414)을 훈련시키기 위한 입력(input)으로 사용한다는 것을 의미한다.
일 실시예에서, 프로세서(210)는 특정한 객체의 이미지에 해당하는 이미지가 전체 이미지에서 차지하는 비중이 일정한 비율 이상이면 해당하는 특정한 객체를 학습 대상으로 선택할 수 있다. 예를 들어, 샘플링된 전체 프레임에서 크롭된 모든 이미지들 중 객체 A에 대한 이미지가 차지하는 비율이 70% 이상일 수 있다. 프로세서(210)는 비율이 30%가 넘어가는 이미지들을 학습 대상으로 선택할 수 있고, A에 대한 이미지가 차지하는 비율이 해당하는 임계치를 초과하므로, A에 대한 이미지들을 학습 대상으로 선택할 수 있다. 그 후 A에 대한 이미지들을 훈련시킬 수 있다.
일 실시예에서 프로세서(210)는 특정한 객체의 이미지에 해당하는 확률을 계산하기 위해 전체 크롭된 이미지의 개수를 측정(measure)할 수 있고, 그 중에서 A로 라벨링된 이미지, 또는 A와 유사한 이미지로 그룹화된 이미지들의 개수를 측정할 수 있다. 프로세서(210)는 동일한 종류로 라벨링된 또는 그룹화된 이미지들을 전체 이미지의 개수로 나누어 확률을 계산할 수 있다.
일 실시예에서 프로세서(210)는 어떤 객체의 이미지가 자주 등장하지 않은 경우라면 위에서 정의되는 확률보다 낮다고 판단할 수 있다. 또는, 프로세서(210)는 어떤 객체의 이미지가 자주 등장하지 않은 경우 그 객체에 대한 이미지의 절대적인 개수가 정해진 값보다 적다고 판단하여 해당하는 객체의 이미지를 학습 대상에서 제외할 수 있다. 프로세서(210)는 메모리(220)에 크롭된 모든 이미지를 저장하고 필요시 로드해서 사용할 수 있는데 학습 대상에서 제외하는 경우 해당하는 이미지를 메모리(220)에서 제거할 수 있다.
프로세서(210)는 학습 모델(414)을 학습하기 전에 미리 필터링을 통해 학습 대상을 선택하고 일정 빈도 이상의 객체에 대해서만 학습 대상으로 선택함으로써, 불필요한 리소스(resource) 낭비를 방지할 수 있다. 또한, 프로세서(210)는 학습 대상을 일부로 제한함으로써 지나친 레이블의 필요성을 방지하여 추후 학습된 모델이 분류를 할 때 소요되는 시간을 절약할 수도 있다.
도 5는 본 개시의 일 실시예에 따른 이미지 그룹의 생성 방법을 도시한 도면이다. 도 5에서는 설명의 편의를 위해 학습 모델의 참조번호를 '530'으로 기재하였지만 학습 모델(104), 학습 모델(414), 학습 모델(530)은 서로 다른 모델이 아니며 동일하거나 대응하는 학습 모델일 수 있다. 일 실시예에서 학습 모델을 훈련 시키기 전에 미리 유사한 이미지들끼리 그룹화하는 경우 프로세서(210)는 유사한 객체를 한 그룹으로 분류할 수 있다.
훈련을 위해 준비된 이미지 사이의 유사도를 판단하기 위해, 이미지에서 각 객체의 특징점(feature point)이 추출될 수 있다. 이미지의 특징점을 추출하는 것은, 모서리, 코너, 객체의 크기, 객체의 색상, 객체의 강도(rigid), 조도 등을 추출하는 것일 수 있다. 예를 들어, 이미지 안에서 픽셀의 값이 갑자기 변하는 경계선을 추출하기 위해 이미지를 미분한 그레디언트(gradient)벡터의 크기를 이용할 수 있다. 이미지 사이에서 추출된 특징점이 유사하면 유사한 이미지로 분류될 수 있다.
객체마다 그룹으로 분류되는 경우, 그룹 1(510)에는 객체 A에 대한 서로 다른 이미지(511, 512)가 포함될 수 있다. 서로 다른 이미지(511, 512)는, 예를 들어, 그룹의 분류를 특징점으로 한 경우, 특징점이 유사하지만 전체적으로는 사이즈, 표정, 행동, 포즈 등이 다른 이미지일 수 있다. 그룹 2(520)에는 객체 B에 대한 서로 다른 이미지(521, 522)가 포함될 수 있다.
일 실시예에서, 프로세서(210)는 이미지들을 일차적으로 그룹별로 분류한 후, 다른 특징점을 이용해 유사도를 한 번 더 체크하는 과정을 거쳐 그룹에 상이한 객체의 이미지가 포함되지 않도록 추가로 처리할 수 있다. 예를 들어, 프로세서(210)가 일차적으로 코너를 중심으로 이미지들을 분류하였으면 이차적으로 에지를 중심으로 객체의 이미지 사이의 유사 여부를 판단할 수 있다.
다른 실시예에서, 프로세서(210)는 이미지들 간의 유사도를 판단하는 기준값을 조절함으로써, 이미지의 분류를 조정할 수 있다. 프로세서(210)가 이미지 사이의 유사도를 넓게 보면 전체 그룹이 감소함에 따라 레이블의 개수가 줄어들고, 학습 모델(530)의 출력 복잡도를 감소시킬 수 있다.
일 실시예에서, 프로세서(210)는 그룹을 설정하고(즉, 분류를 하고), 그룹화된 이미지를 학습 모델(530)의 학습 데이터로 입력할 수 있다. 프로세서(210)는 이미지를 분류하여 학습 데이터를 정리함으로써 학습 모델(530)을 효율적으로 훈련시킬 수 있는 효과가 있다. 따라서, 프로세서(210)가 이미지를 분류하는 경우 사용자가 일일이 분류할 필요가 없게 되어 사용자의 편의를 도모할 수 있다.
일 실시예에서 프로세서(210)는 그룹마다 레이블을 배정할 수 있다. 레이블은 의미 없는 순서를 나타낼 수도 있지만, 다른 경우에서 레이블은 객체의 성질을 나타낼 수도 있다. 예를 들어, 레이블을 운영자(110)로부터 입력받는 경우라면 운영자(110)는 그룹 1(510), 그룹 2(520)에 대해 대응하는 캐릭터 이름을 레이블로 설정할 수 있다. 즉, 레이블은 그룹의 객체에 대한 특성(예를 들어, 캐릭터의 이름, 인물의 이름, 로고 명칭, 특정한 글자 등)으로 설정될 수 있고, 단순히 서수나 알파벳으로 설정될 수도 있다.
일 실시예에서 특정한 그룹에 대해, 학습을 위한 이미지의 개수가 부족할 수 있다. 예를 들어, 그룹 1에 대한 이미지의 개수가 학습에 필요한 개수 미만일 수 있다. 예를 들어, 학습에 필요한 이미지의 개수는 프로세서(210)에 미리 정해져 있을 수 있다. 예를 들어, 학습에 필요한 최소 이미지의 개수는 50개 일 수 있고, 그룹 1에 대한 이미지의 개수가 50개 미만일 수 있다.
일 실시예에서, 학습에 필요한 이미지 개수가 부족한 경우, 프로세서(210)는 해당하는 레이블의 그룹에 대한 이미지를 추가적으로 생성할 수 있다. 예를 들어, 프로세서(210)는 객체에 대한 이미지의 개수가 소정의 개수 미만이면 이미지 처리 모델을 통해 추가적인 이미지를 생성하고 추가적인 이미지를 학습 모델(530)의 학습 데이터로 사용할 수 있다. 예를 들어, 이미지 처리 모델은 미리 메모리(220)에 저장되어 있고, 프로세서(210)에 의해 로드될 수 있다. 예를 들어, 이미지 처리 모델은 이미지의 크기를 변경하거나, 색상을 변경하거나, 색 변환을 실시하거나, 채도를 조절하거나, 이미지의 일부를 자르거나, 이미지를 회전시키거나, 이미지를 다른 이미지와 합치거나 이미지의 종횡비를 조절하는 처리를 할 수 있다. 프로세서(210)는 이미지 처리 모델을 통해 특정한 객체에 대한 추가적인 이미지를 생성할 수 있다. 추가적인 이미지는 객체에 대한 이미지 수를 증가시켜, 훈련을 위한 데이터가 부족한 경우에 있어서도 추가적인 데이터를 통해 학습 모델(530)을 훈련시키는데 사용될 수 있다.
일 실시예에서, 추가적인 학습을 위한 이미지를 생성하기 위해 프로세서(210)는 미리 정의된 이미지 생성 알고리즘을 사용할 수 있다. 이미지 생성 알고리즘은 위에서 설명된 이미지 변환뿐만 아니라 이미지에 다른 이미지, 레이어를 합성하는 동작을 수행할 수도 있다.
일 실시예에서, 학습 모델(530)은 필요한 레이블에 해당하는 이미지를 운영자(110)에게 추가적으로 요청할 수도 있다.
일 실시예에서, 학습 모델(530)은 선형 판별 분석, KNN(K Nearest Neighbor), 의사 결정 트리, 신경망, SVM(Support Vector Machine) 방식 중 적어도 하나에 의해 학습되도록 구성될 수 있다.
학습 모델(530)이 신경망 모델일 때 학습 모델(530)은 복수의 레이어들을 포함하는 딥 뉴럴 네트워크(deep neural network, DNN)에 해당할 수 있고, 간단히 '신경망'으로 지칭될 수 있다. 복수의 레이어들은 입력 레이어(input layer), 히든 레이어 (hidden layer), 및 출력 레이어(output layer)를 포함할 수 있다. 신경망은 완전 연결 네트워크(fully connected network, FCN), 컨볼루셔널 뉴럴 네트워크(convolutional neural network, CNN), 및 리커런트 뉴럴 네트워크(recurrent neural network, RNN)를 포함할 수 있다. 예를 들어, 신경망 내 복수의 레이어들 중 어느 일부는 컨볼루셔널 뉴럴 네트워크(CNN)에 해당할 수 있고, 다른 일부는 완전 연결 네트워크(FCN)에 해당할 수 있다. 이 경우, 컨볼루셔널 뉴럴 네트워크(CNN)는 컨볼루셔널 레이어로 지칭될 수 있고, 완전 연결 네트워크 (FCN)는 완전 연결 레이어로 지칭될 수 있다.
컨볼루셔널 뉴럴 네트워크(CNN)의 경우, 각 레이어에 입력되는 데이터는 입력 특징 맵(input feature map)으로 지칭될 수 있고, 각 레이어에서 출력되는 데이터는 출력 특징 맵(output feature map)으로 지칭될 수 있다. 입력 특징 맵 및 출력 특징 맵은 액티베이션 데이터(activation data)로 지칭될 수도 있다.
신경망 모델은 딥 러닝에 기반하여 트레이닝된 후, 비선형적 관계에 있는 입력 데이터 및 출력 데이터를 서로 매핑함으로써 트레이닝 목적에 맞는 추론(inference)을 수행해낼 수 있다. 딥 러닝은 빅 데이터 세트(Big Dataset)로부터 영상 또는 음성 인식과 같은 문제를 해결하기 위한 기계 학습 기법이다.
도 6은 본 개시의 일 실시예에 따른 유사도 비교 방법을 도시한 도면이다. 학습 모델(104)은 유사도를 비교하기 위한 영상(102) 입력에 앞서 미리 학습 영상(111)을 통해 분류를 위한 훈련이 된 것으로 가정한다. 영상(102)은 전자 장치(103)에 전송되어 학습 모델(104)에 입력될 수 있다. 전자 장치(103)가 영상(102)을 복수의 프레임으로 샘플링한 뒤 복수의 프레임을 학습 모델(104)에 입력해서 학습 모델(104)은 분류를 진행하고 분류 결과에 기초하여 유사도를 판단할 수 있다. 다만, 전자 장치(103)가 샘플링하여 프레임에 대한 이미지 또는 객체에 대한 이미지를 학습 모델(104)에 입력하는 것에 제한되는 것은 아니고, 학습 모델(104)이 영상(102) 자체를 입력받아 프레임을 샘플링하고 그 후에 유사도를 판단하는 동작까지 수행하게 구성될 수도 있다.
일 실시예에서, 샘플링된 복수의 프레임 중 일 프레임(611)에서 객체(612)의 이미지가 추출될 수 있다. 객체(612)는 도 4a에서 설명된 훈련을 위한 데이터를 만드는 것과 마찬가지의 방식으로 프레임에서 식별되어 크롭된 뒤 이미지(613)로 추출될 수 있다. 이미지(613)는 학습 모델(104)에 입력되어 하나 이상의 레이블(620) 마다의 복수의 이미지(621, 622, 623)와 비교될 수 있다. 학습 모델(104)은 입력된 이미지를, 학습을 위한 레이블링에서 입력된 모든 레이블의 모든 이미지와 각각 비교할 수 있다.
다른 실시예에서, 이미지(612)는 따로 크롭되지 않고 프레임(611) 전체에 대한 이미지가 학습 모델(104)에 입력되어 레이블(620)에 해당하는 복수의 이미지(621, 622, 623)와 비교될 수 있다. 즉, 학습 모델(104)은 입력된 프레임에 대한 이미지 내에 레이블(620)에 해당하는 복수의 이미지(621, 622, 623)가 포함되어 있는지를 판단할 수 있다. 예를 들어, 일 레이블에 해당하는 이미지가 입력된 프레임에 포함되었다고 판단되면, 해당 프레임은 판단된 레이블로 분류될 수 있다. 프레임이 하나 이상의 객체 이미지로 추출되지 않고, 프레임 전체로서 학습 모델(104)에 입력된 경우, 해당 프레임에 여러 레이블의 이미지가 있을 수 있다. 프레임에 하나 이상의 레이블에 대응되는 이미지가 포함되어 있는 경우, 해당 프레임은 후의 도 7b에서 설명될 가중치가 가장 높은 레이블에 해당하는 이미지로 분류될 수 있다.
일 실시예에서, 학습 모델(104)은 유사도 판단을 위해 입력된 이미지에 변경을 가할 수 있다. 예를 들어, 프레임(611)에서 객체에 대한 이미지가 추출되어 학습 모델(104)에 입력된 경우에 있어서, 이미지(613)는 특정한 객체의 뒤집어진 형상일 수 있다. 학습 모델(104)은 이미지(613)를 변경하여 레이블에 대한 이미지들과 비교할 수 있다. 예를 들어, 학습 모델(104)은 이미지(613)에 크기 변경, 색상 변경, 이미지 자르기, 이미지 회전을 통해 이미지를 변경하고 변경된 이미지를 각각의 레이블에 대한 이미지들과 비교할 수 있다. 예를 들어, 학습 모델(104)은 이미지(613)를 180도 회전하여 이미지(614)를 만들고, 회전된 이미지를 레이블(620)에 대한 복수의 이미지(621, 622, 623)와 비교할 수 있다. 이를 통해 변형된 이미지에 대해서도 유사도가 정확하게 계산될 수 있다. 다만, 학습 모델(104)이 이미지를 변형시키는 것에 한정되는 것은 아니고 프로세서(210)가 이미지를 입력하기 전에 미리 이미지를 변형(예: 크기 변경, 색상 변경, 이미지 자르기, 이미지 회전 등)하여 입력할 수도 있다.
일 실시예에서, 학습 모델(104)은 이미지의 특징점을 획득하고 레이블에 대응하는 이미지와 비교할 수 있다. 예를 들어, 학습 모델(104)은 이미지 중에서 코너를 특징점으로 추출하여 레이블에 대한 이미지와 비교할 수 있다.
도 7a, 도 7b는 본 개시의 일 실시예에 따른 유사도를 판단하는 방법을 나타내는 도면이다.
학습 모델(104)은 사용자가 입력한 영상(102)이 학습 모델(104)의 학습 데이터로 사용되었던 학습 영상(111)과 얼마나 유사한지에 관한 유사도를 판단할 수 있다. 이를 위해 학습 모델(104)은 영상(102)의 프레임들 중 객체에 대한 제 1 이미지를 입력받은 후 이미지의 특징점을 추출할 수 있다. 특징점은 코너, 모서리 등일 수 있다. 추출된 객체의 제1 이미지(701) 또는 제1 이미지(701)의 추출된 특징점은 학습 모델(104)에 저장되어 있는 레이블들의 복수의 이미지의 특징점과 비교될 수 있다. 예를 들어, 추출된 제1 이미지(701)의 특징점과 레이블에 해당하는 이미지들 사이의 매칭되는 특징점 개수에 기초하여 레이블에 대한 유사도를 계산할 수 있다.
일 실시예에서, 하나의 레이블에 복수의 이미지가 있을 때, 가장 높은 유사도를 갖는 이미지를 기초로 해당하는 레이블과의 유사도를 결정할 수 있다. 다른 경우에서, 하나의 레이블에 복수의 이미지가 있을 때, 복수의 이미지에 대한 유사도 평균을 구해서 해당하는 레이블과의 유사도를 결정할 수도 있다.
다른 실시예에서, 학습 모델(104)은 각 레이블에 해당하는 학습된 이미지들을 기초로 레이블에 대한 대표 이미지를 생성할 수 있다. 이러한 경우 유사도를 판단할 때 영상(102)의 프레임들 중 추출된 객체의 이미지는 레이블에 해당되는 모든 이미지들이 아닌 학습 모델(104)의 각 레이블에 대한 대표 이미지와 제1 이미지(701) 사이의 유사도가 판단될 수 있다. 예를 들어, 학습 모델(104)은 입력된 제1 이미지(701)와 레이블의 대표 이미지의 특징점을 추출하고 매칭되는 특징점 개수를 산출하여 유사도를 판단할 수 있다. 대표 이미지를 이용함으로써 학습 모델(104)의 비교 속도를 향상시킬 수 있다.
일 실시예에서, 유사도는 신경망 모델을 사용하여 모든 레이블에 대한 유사도의 합이 1이 되도록 하나 이상의레이블 각각에 대해 소프트맥스(softmax) 함수를 통해 계산될 수 있다. 예를 들어, 제1 이미지(701)는 레이블1, 2, 3에 대해 각각 0.75, 0.24, 0.01의 유사도를 가질 수 있다. 이는 제1 이미지(701)는 레이블 1의 이미지와 0.75만큼 즉, 75%만큼의 유사한 이미지임을 나타낼 수 있다. 예를 들어, 제2 이미지(711)는 레이블 1, 2, 3에 대해 각각 0.11, 0.37, 0.52의 유사도를 가질 수 있다. 이는 제2 이미지(711)는 레이블 1의 이미지와는 11%만큼 유사하고, 레이블 2의 이미지와는 37% 유사하며 레이블 3의 이미지와는 52% 유사함을 의미할 수 있다.
일 실시예에서, 각각의 레이블에 대해 계산된 유사도에 기초하여, 학습 모델(104)은 특정한 이미지가 어떤 레이블에 속하는지 분류할 수 있다. 일 실시예에서, 학습 모델(104)은 레이블 중 유사도가 가장 높은 레이블로 해당 이미지를 분류할 수 있다. 예를 들어, 제1 이미지(701)에 대해서는 레이블 1의 유사도가 가장 높으므로 제1 이미지(701)는 레이블 1로 분류될 수 있다. 예를 들어, 학습 모델(104)은 제2 이미지(711)에 대해서는 레이블 3에 대한 유사도가 가장 높으므로 제2 이미지(711)는 레이블 3으로 분류될 수 있다.
다른 실시예에서, 학습 모델(104)은 특정 레이블과의 유사도가 가장 높고 유사도가 일정한 임계값을 초과하면 해당 레이블로 분류할 수 있다. 반면, 학습 모델(104)은 특정 레이블과의 유사도가 가장 높아도 해당 유사도가 임계값 이하이면 어떠한 레이블로도 분류하지 않을 수 있다. 예를 들어, 학습 모델(104)은 레이블과의 유사도가 0.5(50%)를 넘으면 해당 레이블로 분류를 할 수 있다. 예를 들어, 제1 이미지(701)는 레이블 1과의 유사도가 0.75이므로 레이블 1(702)로 분류될 수 있다. 즉, 제1 이미지(701)는 레이블 1에 해당하는 것으로 판정될 수 있다. 예를 들어, 특정한 이미지(미도시)는 모든 레이블에 대해 유사도가 0.5를 넘지 않을 수 있다. 이러한 경우 해당하는 이미지는 어떠한 레이블로도 분류되지 않을 수 있다.
이를 통해 레이블 중 어느 것과도 유사하지 않음에도 학습 모델(104)의 특성 상 레이블 중 어느 하나로 분류되어 추후 총 유사도 판단에 사용되는 경우, 영상(102)이 학습 영상(111)과 유사하지 않음에도 유사하다고 판단될 수 있는 문제를 해결할 수 있다.
도 7b는 본 개시의 일 실시예에 따른, 영상(102) 속 하나 이상의 이미지 각각의 학습 모델에 훈련된 레이블들에 대한 유사도를 기초로 입력된 영상(102)의 학습 영상(111)에 대한 총 유사도를 판단하는 방법을 나타내는 도면이다.
일 실시예에서, 학습 모델(104)은 입력 영상(102)에서 추출된 각 객체의 이미지들을 도 7a에 도시된 것과 같이 하나 이상의 레이블 각각의 복수의 이미지와 비교한 뒤 또는 각각의 레이블의 대표 이미지와 비교한 뒤 객체의 이미지의 각각의 레이블에 대한 유사도를 판단할 수 있다. 일 실시예에서, 프로세서(210)는 학습 모델(104)로부터 출력된 영상(102)의 프레임들 속의 추출된 객체 모두가 각각의 레이블에 대해 판단된 유사도를 기초로 총 유사도를 판단할 수 있다. 총 유사도는 하나 이상의 객체가 포함된 입력 영상(102)의 학습 영상(111)에 대한 유사도를 의미한다.
총 유사도는, 레이블의 가중치의 함수일 수 있다. 예를 들어, 프로세서(210)는 레이블마다 가중치를 다르게 설정하고, 하나 이상의 레이블 각각에 대한 유사도에 하나 이상의 레이블의 가중치를 곱하여 입력 영상(102)의 학습 영상(111)에 대한 총 유사도를 계산할 수 있다. 이를 통해 입력 영상(102)의 학습 영상(111)과의 유사한 정도를 유의미하고 편리하게 계산할 수 있다.
예를 들어, 프로세서(210)는 학습 영상(111)에서 해당 레이블에 대한 객체 이미지가 얼마나 등장하였는지 빈도를 계산할 수 있다. 빈도를 기초로 가중치를 계산함으로써 합리적인 유사도 판단이 이루어질 수 있다. 프로세서(210)는 계산된 빈도를 기초로 자주 등장한 객체의 레이블에는 가중치를 높게 둘 수 있고, 빈도가 낮은 객체의 레이블에 대해서는 가중치를 낮게 둘 수 있다. 프로세서(210)는 레이블에 대한 가중치를 운영자(110)로부터 입력 받을 수도 있다. 프로세서(103)에는 미리 가중치에 대한 정보가 저장되어 있을 수 있다.
일 실시예에서, 프로세서(210)는 입력된 영상(102)에 자주 등장하는 객체에 대해 매칭되는 레이블의 가중치를 높게 설정할 수 있다.
일 실시예에서, 각 레이블을 평가하는 레이블 평가 모델이 있고, 해당 레이블 평가 모델이 레이블의 가중치를 설정할 수 있다. 예를 들어, 레이블 평가 모델은 영상(102)에서 레이블에 해당하는 객체의 빈도 등에 기초한 함수로 학습될 수 있다.
다른 실시예에서, 프로세서(210)는 레이블에 대해 가중치를 설정하는 것이 아니고 입력된 영상(102)으로부터의 객체(721, 722, 723)에 직접 가중치를 설정할 수도 있다. 프로세서(210)는 객체(721, 722, 723)의 정해진 시간 동안의 움직임 벡터를 측정하고, 움직임 벡터의 양이 정해진 값보다 큰 경우 중요성이 높다고 판단하여 가중치를 높게 설정할 수 있다.
일 실시예에서, 레이블 또는 객체(721, 722, 723)에 정해진 가중치를 기초로 입력된 영상(102)의 학습 영상(111)에 대한 총 유사도가 계산될 수 있다. 총 유사도를 계산하기 위해 학습 모델(104)은 하나 이상의 레이블 각각의 유사도에 레이블 각각의 가중치를 곱하고 이를 기초로 총 유사도를 계산할 수 있다. 예를 들어, 학습 모델(104)은 객체(721, 722, 723)의 각 유사도를 기초로 입력된 영상(102)이 학습 영상(111)과 73% 유사하다는 총 유사도(724)를 계산할 수 있다.
일 실시예에 따르면, 전자 장치(103)는 계산된 총 유사도(724)가 임계값 이상이면 입력된 영상(102)의 업로드를 제한할 수 있다. 예를 들어, 전자 장치(103)는 사용자(101)에게 영상 침해에 대한 경고 문구를 디스플레이할 수 있다. 예를 들어, 전자 장치(103)는 영상(102)을 서버(105)에 업로드하는 것을 금지할 수 있다. 예를 들어, 전자 장치(103)는 영상(102)을 전자 장치(103)의 메모리에 저장하는 것을 금지할 수 있다. 예를 들어, 전자 장치(103)는 학습 영상(111)의 일례가 될 수 있는 타인의 저작권이 문제되는 영상과 유사한 영상의 경우 서버(105)에 업로드를 금지하여 저작권 침해 문제를 방지할 수 있다.
도 8은 본 개시의 일 실시예에 따른, 입력 영상이 학습 영상과 유사한지를 판단하는 방법에 대한 흐름도이다. 이하 실시예에서, 각 단계들은 순차적으로 수행될 수도 있으나, 반드시 순차적으로 수행되는 것은 아니다. 예를 들어, 각 단계들의 순서가 변경될 수도 있으며, 적어도 두 단계들이 병렬적으로 수행될 수도 있다.
단계(S810)에서 전자 장치(103)는 업로드 영상(102)으로부터 복수의 프레임을 샘플링한다. 예를 들어, 전자 장치(103)는 샘플링 주기가 정해진 경우에는 해당 주기마다 프레임을 샘플링하여 취득할 수 있다. 샘플링 주기가 정해지지 않은 경우에는 전자 장치(103)는 운영자(110)로부터 원하는 프레임 시간을 입력 받아 샘플링할 수 있다.
단계(S820)에서는 전자 장치(103)가 샘플링 된 복수의 프레임에서 객체가 있는 영역을 크롭하여 객체에 대응하는 이미지를 생성할 수 있다. 객체가 있는 영역은 ROI로서 식별될 수 있고, 운영자(110)에 의해 마우스 드래그로 지정될 수 있다. 운영자(110)의 입력이 아닌, ROI 지정에 의해 객체가 있는 영역이 식별되는 경우에는, ROI를 찾는 알고리즘에 의해 객체가 있는 영역이 식별될 수 있다.
단계(S830)에서 전자 장치(103)는 객체에 대응하는 이미지들을 레이블링하여 학습 모델을 지도 훈련시킬 수 있다. 전자 장치(103)는 학습 데이터로서 객체에 대응하는 이미지와, 해당하는 이미지에 대한 레이블을 학습 모델에 입력할 수 있다. 학습 데이터를 이용하여 학습 모델(104)은 입력 변수와 출력 변수를 맵핑시키는 함수를 찾을 수 있다. 훈련은 유사도를 판단하고자 하는 영상(102)이 전자 장치(103)에 입력 되기 전에 이루어질 수 있다.
단계(S840)에서 학습 모델(104)을 통해 입력 영상(102) 내의 객체들의 각각의 레이블에 대한 유사도를 획득할 수 있다. 입력된 영상은 복수 개의 프레임으로 분리되어 학습 모델(104)에 입력될 수 있다. 입력된 영상은 복수 개의 프레임으로 분리된 뒤 객체에 대한 이미지들이 크롭되어 학습 모델(104)에 입력될 수도 있다. 학습 모델(104)은 출력으로서 해당 이미지에 대해, 레이블들 중 가장 유사한 레이블을 출력할 수 있다. 예를 들어, 학습 모델(104)이 레이블 A, B, C로 지도 학습된 경우라면 학습 모델(104)에 입력된 객체의 이미지는 A, B, C 중 하나로 분류될 수 있다. 일 실시예에서, 학습 모델(104)은 추가적인 출력으로 객체의 이미지가 각 레이블과 갖는 유사도를 출력할 수 있다. 예를 들어, A, B, C 레이블과의 각각의 유사도를 출력할 수 있다.
단계(S850)에서, 전자 장치(103)는 레이블 각각에 대한 가중치를 다르게 설정하여 타겟 영상의 훈련 영상에 대한 총 유사도를 산출할 수 있다. 예를 들어, 총 유사도는 하나 이상의 레이블 각각에 대한 유사도에 하나 이상의 레이블 각각의 가중치를 곱하여 구할 수 있다.
일 실시예에서, 전자 장치(103)는 레이블에 따른 가중치를 미리 전자 장치(103)의 메모리(220)에 저장된 값을 로드해서 사용할 수 있다. 일 실시예에서, 전자 장치(103)는 레이블에 따른 객체 이미지의 빈도에 기초하여 가중치를 설정할 수 있다. 일 실시예에서, 전자 장치(103)는 운영자(110)의 입력에 기초하여 가중치를 설정할 수도 있다. 예를 들어, 가중치에 따라 특정한 레이블에 해당하는 객체의 이미지가 입력 영상(102)에서 다른 이미지들보다 적게 등장하는 경우라도 입력 영상과 학습 영상(111)의 총 유사도를 높게 하는데 기여할 수 있다.
전자 장치(103)는 특정한 레이블에 대한 가중치가 소정값 이상인 경우에만, 해당하는 레이블과 일치하는 객체 이미지들의 개수에 기초하여 총 유사도를 계산할 수도 있다. 이러한 경우, 모든 레이블에 대한 유사도를 고려하여 총 유사도를 계산하는 것보다 유의미한 레이블에 대해서만 계산하게 되므로, 계산 속도를 빠르게 할 수 있고 컴퓨팅 로드를 줄일 수 있다는 효과를 갖는다.
이상에서 설명된 실시예들은 하드웨어 구성요소, 소프트웨어 구성요소, 및/또는 하드웨어 구성요소 및 소프트웨어 구성요소의 조합으로 구현될 수 있다. 예를 들어, 실시예들에서 설명된 장치, 방법 및 구성요소는, 예를 들어, 프로세서, 콘트롤러, ALU(arithmetic logic unit), 디지털 신호 프로세서(digital signal processor), 마이크로컴퓨터, FPGA(field programmable gate array), PLU(programmable logic unit), 마이크로프로세서, 또는 명령(instruction)을 실행하고 응답할 수 있는 다른 어떠한 장치와 같이, 범용 컴퓨터 또는 특수 목적 컴퓨터를 이용하여 구현될 수 있다. 처리 장치는 운영 체제(OS) 및 상기 운영 체제 상에서 수행되는 소프트웨어 애플리케이션을 수행할 수 있다. 또한, 처리 장치는 소프트웨어의 실행에 응답하여, 데이터를 접근, 저장, 조작, 처리 및 생성할 수도 있다. 이해의 편의를 위하여, 처리 장치는 하나가 사용되는 것으로 설명된 경우도 있지만, 해당 기술분야에서 통상의 지식을 가진 자는, 처리 장치가 복수 개의 처리 요소(processing element) 및/또는 복수 유형의 처리 요소를 포함할 수 있음을 알 수 있다. 예를 들어, 처리 장치는 복수 개의 프로세서 또는 하나의 프로세서 및 하나의 컨트롤러를 포함할 수 있다. 또한, 병렬 프로세서(parallel processor)와 같은, 다른 처리 구성(processing configuration)도 가능하다. 소프트웨어는 컴퓨터 프로그램(computer program), 코드(code), 명령(instruction), 또는 이들 중 하나 이상의 조합을 포함할 수 있으며, 원하는 대로 동작하도록 처리 장치를 구성하거나 독립적으로 또는 결합적으로 (collectively) 처리 장치를 명령할 수 있다. 소프트웨어 및/또는 데이터는, 처리 장치에 의하여 해석되거나 처리 장치에 명령 또는 데이터를 제공하기 위하여, 어떤 유형의 기계, 구성요소(component), 물리적 장치, 가상 장치(virtual equipment), 컴퓨터 저장 매체 또는 장치, 또는 전송되는 신호 파(signal wave)에 영구적으로, 또는 일시적으로 구체화(embody)될 수 있다. 소프트웨어는 네트워크로 연결된 컴퓨터 시스템 상에 분산되어서, 분산된 방법으로 저장되거나 실행될 수도 있다. 소프트웨어 및 데이터는 컴퓨터 판독 가능 기록 매체에 저장될 수 있다.
실시예에 따른 방법은 다양한 컴퓨터 수단을 통하여 수행될 수 있는 프로그램 명령 형태로 구현되어 컴퓨터 판독 가능 매체에 기록될 수 있다. 컴퓨터 판독 가능 매체는 프로그램 명령, 데이터 파일, 데이터 구조 등을 단 독으로 또는 조합하여 저장할 수 있으며 매체에 기록되는 프로그램 명령은 실시예를 위하여 특별히 설계되고 구성된 것들이거나 컴퓨터 소프트웨어 당업자에게 공지되어 사용 가능한 것일 수도 있다. 컴퓨터 판독 가능 기록 매체의 예에는 하드 디스크, 플로피 디스크 및 자기 테이프와 같은 자기 매체(magnetic media), CD-ROM, DVD와 같은 광기록 매체(optical media), 플롭티컬 디스크(floptical disk)와 같은 자기-광 매체(magneto-optical media), 및 롬(ROM), 램(RAM), 플래시 메모리 등과 같은 프로그램 명령을 저장하고 수행하도록 특별히 구성된 하드웨어 장치가 포함된다. 프로그램 명령의 예에는 컴파일러에 의해 만들어지는 것과 같은 기계어 코드뿐만 아니라 인터프리터 등을 사용해서 컴퓨터에 의해서 실행될 수 있는 고급 언어 코드를 포함한다. 위에서 설명한 하드웨어 장치는 실시예의 동작을 수행하기 위해 하나 또는 복수의 소프트웨어 모듈로서 작동하 도록 구성될 수 있으며, 그 역도 마찬가지이다. 이상과 같이 실시예들이 비록 한정된 도면에 의해 설명되었으나, 해당 기술분야에서 통상의 지식을 가진 자라면 이를 기초로 다양한 기술적 수정 및 변형을 적용할 수 있다. 예를 들어, 설명된 기술들이 설명된 방법과 다른 순서로 수행되거나, 및/또는 설명된 시스템, 구조, 장치, 회로 등의 구성요소들이 설명된 방법과 다른 형태로 결합 또는 조합되거나, 다른 구성요소 또는 균등물에 의하여 대치되거나 치환되더라도 적절한 결과가 달성될 수 있다. 그러므로, 다른 구현들, 다른 실시예들 및 특허 청구범위와 균등한 것들도 후술하는 특허 청구의 범위의 범위에 속한다.
이하에서는, 본 개시의 다양한 실시예에 대해서 부기한다.
[부기 1]
하나 이상의 프로세서; 및
상기 프로세서에 의한 실행 시 상기 프로세서가 연산을 수행하도록 하는 명령어들이 저장된 하나 이상의 메모리를 포함하고,
상기 프로세서에 의해 상기 명령어들이 실행될 시, 상기 프로세서는,
이미지 프레임의 시퀀스로 구성된 학습 영상을 복수의 프레임으로 샘플링하고,
상기 복수의 프레임 각각에 나타난 하나 이상의 객체를 검출하여 상기 하나 이상의 객체에 대응하는 복수의 객체 이미지를 생성하고,
상기 복수의 객체 이미지를 레이블링(labeling)하여 상기 하나 이상의 객체 각각으로 분류되도록 학습 모델을 훈련하도록 구성되는, 전자 장치.
[부기 2]
제1항에 있어서,
상기 프로세서는,
상기 복수의 객체 이미지 사이의 특징점을 찾아 비교하여 특정 객체를 지시하는 유사한 이미지들인지 여부를 판단하고,
유사한 이미지들끼리 그룹화하여 상기 특정 객체를 지시하는 이미지 그룹을 생성하며,
상기 이미지 그룹에 기초하여 상기 학습 모델을 훈련시키도록 더 구성되는, 전자 장치.
[부기 3]
제1항에 있어서,
상기 하나 이상의 객체를 검출하여 상기 하나 이상의 객체에 대응하는 복수의 객체 이미지를 생성하는 것은, 프레임 내에 색상, 질감, 또는 패턴 중 적어도 하나에 기초한 유사한 값을 갖는 영역을 관심 영역으로 그룹화하고, 상기 관심 영역을 배경과 분리함으로써 객체 이미지를 생성하는 것을 포함하는, 전자 장치.
[부기 4]
제1항에 있어서,
상기 객체는 캐릭터, 인물, 특정한 형상, 의상, 건물, 또는 로고 중 적어도 하나에 해당하는, 전자 장치.
[부기 5]
제1항에 있어서,
상기 프로세서는, 상기 복수의 프레임 중 제1 시점에서 샘플링된 제1 프레임에 제1 객체 및 제2 객체가 오버랩(overlap)되어 있는 경우, 상기 제1 시점과 다른 제2 시점에서의 상기 제1 객체가 나타난 제2 프레임 및 상기 제1 프레임을 이용하여, 상기 제1 프레임으로부터 상기 제2 객체의 이미지를 생성하도록 구성되는, 전자 장치.
[부기 6]
제1항에 있어서,
상기 프로세서는,
상기 복수의 객체 이미지 중 상기 하나 이상의 객체 각각에 해당하는 이미지의 개수를 계산하고,
상기 이미지의 개수가 소정의 개수 이상인 객체에 대응하는 하나 이상의 객체의 이미지를 기초로 상기 학습 모델을 훈련하도록 구성되는, 전자 장치.
[부기 7]
제1항에 있어서,
상기 프로세서는,
상기 학습 영상에 나타난 특정 객체의 움직임에 대한 움직임 벡터를 취득하고, 정해진 시간 동안의 상기 움직임 벡터의 크기가 기준값 이상이면 상기 특정 객체를 추적(tracking)하고,
상기 특정 객체가 나타난 프레임을 샘플링하도록 더 구성되는, 전자 장치.
[부기 8]
제1항에 있어서,
상기 프로세서는, 미리 정해진 주기에 따라 상기 학습 영상을 샘플링하거나, 입력에 의해 특정 시점에 해당하는 프레임을 샘플링하도록 구성되는, 전자 장치.
[부기 9]
제6항에 있어서,
상기 프로세서는, 상기 학습 모델을 훈련시키기 위한 특정 객체에 대한 이미지의 개수가 소정의 개수 이하인 경우, 이미지 처리 모델을 통해 추가적인 이미지를 생성하고 상기 특정 객체에 대한 추가적인 이미지로 상기 특정 객체에 대해 학습 모델을 훈련시키도록 더 구성되는, 전자 장치.
[부기 10]
제9항에 있어서,
상기 이미지 처리 모델은 객체에 대한 이미지의 크기 변경, 색상 변경, 자르기, 또는 회전 중 적어도 하나를 수행하는, 전자 장치.
[부기 11]
제1항에 있어서,
상기 학습 모델은 선형 판별 분석, KNN(K Nearest Neighbor), 의사 결정 트리, 신경망, SVM(Support Vector Machine) 방식 중 적어도 하나에 의해 훈련시키도록 더 구성되는, 전자 장치.
[부기 12]
하나 이상의 프로세서;
네트워크 인터페이스;
학습 모델을 저장하는 하나 이상의 메모리를 포함하고,
상기 학습 모델은 학습 영상에 포함된 객체와 레이블의 상관 관계가 학습된 모델이고,
상기 프로세서는,
상기 메모리로부터 상기 학습 모델을 로드(load)하고,
입력된 타겟 영상으로부터 획득한 복수의 이미지를 상기 학습 모델에 입력하여 하나 이상의 레이블 각각에 대한 유사도를 획득하고,
상기 하나 이상의 레이블 각각에 대한 유사도에 상기 하나 이상의 레이블 각각의 가중치를 곱하여, 상기 타겟 영상의 상기 학습 영상에 대한 총 유사도를 산출하도록 구성되는, 전자 장치.
[부기 13]
제12항에 있어서,
상기 프로세서는,
상기 타겟 영상으로부터 복수의 프레임을 획득하고,
상기 복수의 프레임으로부터 하나 이상의 객체에 대응하는 복수의 객체 이미지를 추출하여 상기 학습 모델에 입력하고,
상기 학습 모델은 상기 복수의 객체 이미지를 상기 하나 이상의 레이블 각각에 대응하는 이미지들과 비교하도록 구성되는, 전자 장치.
[부기 14]
제13항에 있어서,
상기 학습 모델은 상기 복수의 객체 이미지에 크기 변경, 색상 변경, 자르기, 또는 회전 중 적어도 하나를 수행한 후, 상기 하나 이상의 레이블 각각에 대응하는 이미지들과 비교하도록 더 구성되는, 전자 장치.
[부기 15]
제13항에 있어서,
상기 학습 모델은,
상기 복수의 객체 이미지에서 상기 하나 이상의 객체 각각의 특징점을 추출하고,
상기 특징점을 이용하여 상기 하나 이상의 레이블 각각에 대응하는 이미지들과 비교하여 유사도를 계산하도록 구성되는, 전자 장치.
[부기 16]
제15항에 있어서,
상기 유사도는, 상기 특징점과 상기 하나 이상의 레이블 각각에 대응하는 이미지들의 특징점 사이의 매칭되는 특징점 개수에 기초하여 계산되는, 전자 장치.
[부기 17]
제12항에 있어서,
상기 유사도는, 신경망 모델을 사용하여, 상기 하나 이상의 레이블에 대한 유사도의 합이 1이 되도록 상기 하나 이상의 레이블 각각에 대해 소프트맥스(softmax) 함수를 통해 계산되는, 전자 장치.
[부기 18]
제12항에 있어서,
상기 하나 이상의 레이블 중 특정 레이블에 대한 가중치는, 상기 특정 레이블에 해당하는 객체의 빈도에 따라서 결정되거나, 입력을 받아 결정되는, 전자 장치.
[부기 19]
제12항에 있어서,
상기 프로세서는,
상기 총 유사도가 임계값 이상이면, 상기 타겟 영상의 업로드를 제한하도록 더 구성되는, 전자 장치.
[부기 20]
입력된 학습 영상을 복수의 프레임으로 샘플링하는 단계;
상기 복수의 프레임 각각에 나타난 하나 이상의 객체를 검출하여 상기 하나 이상의 객체에 대응하는 복수의 객체 이미지를 생성하는 단계;
상기 복수의 객체 이미지를 레이블링하여 상기 하나 이상의 객체 각각으로 분류되도록 학습 모델을 훈련시키는 단계;
입력된 타겟 영상으로부터 획득한 복수의 이미지를 상기 학습 모델에 입력하여 하나 이상의 레이블 각각에 대한 유사도를 획득하는 단계; 및
상기 하나 이상의 레이블 각각에 대한 유사도에 상기 하나 이상의 레이블 각각의 가중치를 곱하여, 상기 타겟 영상의 상기 학습 영상에 대한 총 유사도를 산출하는 단계
를 포함하는, 방법.
Claims (20)
- 하나 이상의 프로세서; 및상기 프로세서에 의한 실행 시 상기 프로세서가 연산을 수행하도록 하는 명령어들이 저장된 하나 이상의 메모리를 포함하고,상기 프로세서에 의해 상기 명령어들이 실행될 시, 상기 프로세서는,이미지 프레임의 시퀀스로 구성된 학습 영상을 복수의 프레임으로 샘플링하고,상기 복수의 프레임 각각에 나타난 하나 이상의 객체를 검출하여 상기 하나 이상의 객체에 대응하는 복수의 객체 이미지를 생성하고,상기 복수의 객체 이미지를 레이블링(labeling)하여 상기 하나 이상의 객체 각각으로 분류되도록 학습 모델을 훈련하도록 구성되는, 전자 장치.
- 제1항에 있어서,상기 프로세서는,상기 복수의 객체 이미지 사이의 특징점을 찾아 비교하여 특정 객체를 지시하는 유사한 이미지들인지 여부를 판단하고,유사한 이미지들끼리 그룹화하여 상기 특정 객체를 지시하는 이미지 그룹을 생성하며,상기 이미지 그룹에 기초하여 상기 학습 모델을 훈련시키도록 더 구성되는, 전자 장치.
- 제1항에 있어서,상기 하나 이상의 객체를 검출하여 상기 하나 이상의 객체에 대응하는 복수의 객체 이미지를 생성하는 것은, 프레임 내에 색상, 질감, 또는 패턴 중 적어도 하나에 기초한 유사한 값을 갖는 영역을 관심 영역으로 그룹화하고, 상기 관심 영역을 배경과 분리함으로써 객체 이미지를 생성하는 것을 포함하는, 전자 장치.
- 제1항에 있어서,상기 객체는 캐릭터, 인물, 특정한 형상, 의상, 건물, 또는 로고 중 적어도 하나에 해당하는, 전자 장치.
- 제1항에 있어서,상기 프로세서는, 상기 복수의 프레임 중 제1 시점에서 샘플링된 제1 프레임에 제1 객체 및 제2 객체가 오버랩(overlap)되어 있는 경우, 상기 제1 시점과 다른 제2 시점에서의 상기 제1 객체가 나타난 제2 프레임 및 상기 제1 프레임을 이용하여, 상기 제1 프레임으로부터 상기 제2 객체의 이미지를 생성하도록 구성되는, 전자 장치.
- 제1항에 있어서,상기 프로세서는,상기 복수의 객체 이미지 중 상기 하나 이상의 객체 각각에 해당하는 이미지의 개수를 계산하고,상기 이미지의 개수가 소정의 개수 이상인 객체에 대응하는 하나 이상의 객체의 이미지를 기초로 상기 학습 모델을 훈련하도록 구성되는, 전자 장치.
- 제1항에 있어서,상기 프로세서는,상기 학습 영상에 나타난 특정 객체의 움직임에 대한 움직임 벡터를 취득하고, 정해진 시간 동안의 상기 움직임 벡터의 크기가 기준값 이상이면 상기 특정 객체를 추적(tracking)하고,상기 특정 객체가 나타난 프레임을 샘플링하도록 더 구성되는, 전자 장치.
- 제1항에 있어서,상기 프로세서는, 미리 정해진 주기에 따라 상기 학습 영상을 샘플링하거나, 입력에 의해 특정 시점에 해당하는 프레임을 샘플링하도록 구성되는, 전자 장치.
- 제6항에 있어서,상기 프로세서는, 상기 학습 모델을 훈련시키기 위한 특정 객체에 대한 이미지의 개수가 소정의 개수 이하인 경우, 이미지 처리 모델을 통해 추가적인 이미지를 생성하고 상기 특정 객체에 대한 추가적인 이미지로 상기 특정 객체에 대해 학습 모델을 훈련시키도록 더 구성되는, 전자 장치.
- 제9항에 있어서,상기 이미지 처리 모델은 객체에 대한 이미지의 크기 변경, 색상 변경, 자르기, 또는 회전 중 적어도 하나를 수행하는, 전자 장치.
- 제1항에 있어서,상기 학습 모델은 선형 판별 분석, KNN(K Nearest Neighbor), 의사 결정 트리, 신경망, SVM(Support Vector Machine) 방식 중 적어도 하나에 의해 훈련시키도록 더 구성되는, 전자 장치.
- 하나 이상의 프로세서;네트워크 인터페이스;학습 모델을 저장하는 하나 이상의 메모리를 포함하고,상기 학습 모델은 학습 영상에 포함된 객체와 레이블의 상관 관계가 학습된 모델이고,상기 프로세서는,상기 메모리로부터 상기 학습 모델을 로드(load)하고,입력된 타겟 영상으로부터 획득한 복수의 이미지를 상기 학습 모델에 입력하여 하나 이상의 레이블 각각에 대한 유사도를 획득하고,상기 하나 이상의 레이블 각각에 대한 유사도에 상기 하나 이상의 레이블 각각의 가중치를 곱하여, 상기 타겟 영상의 상기 학습 영상에 대한 총 유사도를 산출하도록 구성되는, 전자 장치.
- 제12항에 있어서,상기 프로세서는,상기 타겟 영상으로부터 복수의 프레임을 획득하고,상기 복수의 프레임으로부터 하나 이상의 객체에 대응하는 복수의 객체 이미지를 추출하여 상기 학습 모델에 입력하고,상기 학습 모델은 상기 복수의 객체 이미지를 상기 하나 이상의 레이블 각각에 대응하는 이미지들과 비교하도록 구성되는, 전자 장치.
- 제13항에 있어서,상기 학습 모델은 상기 복수의 객체 이미지에 크기 변경, 색상 변경, 자르기, 또는 회전 중 적어도 하나를 수행한 후, 상기 하나 이상의 레이블 각각에 대응하는 이미지들과 비교하도록 더 구성되는, 전자 장치.
- 제13항에 있어서,상기 학습 모델은,상기 복수의 객체 이미지에서 상기 하나 이상의 객체 각각의 특징점을 추출하고,상기 특징점을 이용하여 상기 하나 이상의 레이블 각각에 대응하는 이미지들과 비교하여 유사도를 계산하도록 구성되는, 전자 장치.
- 제15항에 있어서,상기 유사도는, 상기 특징점과 상기 하나 이상의 레이블 각각에 대응하는 이미지들의 특징점 사이의 매칭되는 특징점 개수에 기초하여 계산되는, 전자 장치.
- 제12항에 있어서,상기 유사도는, 신경망 모델을 사용하여, 상기 하나 이상의 레이블에 대한 유사도의 합이 1이 되도록 상기 하나 이상의 레이블 각각에 대해 소프트맥스(softmax) 함수를 통해 계산되는, 전자 장치.
- 제12항에 있어서,상기 하나 이상의 레이블 중 특정 레이블에 대한 가중치는, 상기 특정 레이블에 해당하는 객체의 빈도에 따라서 결정되거나, 입력을 받아 결정되는, 전자 장치.
- 제12항에 있어서,상기 프로세서는,상기 총 유사도가 임계값 이상이면, 상기 타겟 영상의 업로드를 제한하도록 더 구성되는, 전자 장치.
- 입력된 학습 영상을 복수의 프레임으로 샘플링하는 단계;상기 복수의 프레임 각각에 나타난 하나 이상의 객체를 검출하여 상기 하나 이상의 객체에 대응하는 복수의 객체 이미지를 생성하는 단계;상기 복수의 객체 이미지를 레이블링하여 상기 하나 이상의 객체 각각으로 분류되도록 학습 모델을 훈련시키는 단계;입력된 타겟 영상으로부터 획득한 복수의 이미지를 상기 학습 모델에 입력하여 하나 이상의 레이블 각각에 대한 유사도를 획득하는 단계; 및상기 하나 이상의 레이블 각각에 대한 유사도에 상기 하나 이상의 레이블 각각의 가중치를 곱하여, 상기 타겟 영상의 상기 학습 영상에 대한 총 유사도를 산출하는 단계를 포함하는, 방법.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/KR2023/001254 WO2024158073A1 (ko) | 2023-01-27 | 2023-01-27 | 학습 모델을 이용한 영상들 간의 유사도 판단 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/KR2023/001254 WO2024158073A1 (ko) | 2023-01-27 | 2023-01-27 | 학습 모델을 이용한 영상들 간의 유사도 판단 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024158073A1 true WO2024158073A1 (ko) | 2024-08-02 |
Family
ID=91970749
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/KR2023/001254 Ceased WO2024158073A1 (ko) | 2023-01-27 | 2023-01-27 | 학습 모델을 이용한 영상들 간의 유사도 판단 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2024158073A1 (ko) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN119783068A (zh) * | 2024-11-22 | 2025-04-08 | 浙江天猫技术有限公司 | 数字内容检测方法、设备、存储介质和程序产品 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20140047331A (ko) * | 2012-10-12 | 2014-04-22 | 전남대학교산학협력단 | Block Clustering 을 이용한 관심영역기반 자동객체분할방법 및 자동객체분할시스템 |
| KR20160144149A (ko) * | 2015-06-08 | 2016-12-16 | 군산대학교산학협력단 | 다중 이동 물체의 겹침 제거 및 추적을 위한 영상 감시 장치 및 방법 |
| KR102179583B1 (ko) * | 2020-03-11 | 2020-11-17 | 주식회사 딥노이드 | 딥러닝 기반의 질환 진단 보조 시스템 및 딥러닝 기반의 질환 진단 보조 방법 |
| KR20210051473A (ko) * | 2019-10-30 | 2021-05-10 | 한국전자통신연구원 | 동영상 콘텐츠 식별 장치 및 방법 |
| KR20220079542A (ko) * | 2019-10-08 | 2022-06-13 | 유아이패스, 인크. | 컨볼루션 신경 네트워크를 사용하여 로봇 프로세스 자동화에서 사용자 인터페이스 엘리먼트들을 검출하기 |
-
2023
- 2023-01-27 WO PCT/KR2023/001254 patent/WO2024158073A1/ko not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR20140047331A (ko) * | 2012-10-12 | 2014-04-22 | 전남대학교산학협력단 | Block Clustering 을 이용한 관심영역기반 자동객체분할방법 및 자동객체분할시스템 |
| KR20160144149A (ko) * | 2015-06-08 | 2016-12-16 | 군산대학교산학협력단 | 다중 이동 물체의 겹침 제거 및 추적을 위한 영상 감시 장치 및 방법 |
| KR20220079542A (ko) * | 2019-10-08 | 2022-06-13 | 유아이패스, 인크. | 컨볼루션 신경 네트워크를 사용하여 로봇 프로세스 자동화에서 사용자 인터페이스 엘리먼트들을 검출하기 |
| KR20210051473A (ko) * | 2019-10-30 | 2021-05-10 | 한국전자통신연구원 | 동영상 콘텐츠 식별 장치 및 방법 |
| KR102179583B1 (ko) * | 2020-03-11 | 2020-11-17 | 주식회사 딥노이드 | 딥러닝 기반의 질환 진단 보조 시스템 및 딥러닝 기반의 질환 진단 보조 방법 |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN119783068A (zh) * | 2024-11-22 | 2025-04-08 | 浙江天猫技术有限公司 | 数字内容检测方法、设备、存储介质和程序产品 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020138745A1 (en) | Image processing method, apparatus, electronic device and computer readable storage medium | |
| WO2021080102A1 (en) | Method for training and testing adaption network corresponding to obfuscation network capable of processing data to be concealed for privacy, and training device and testing device using the same | |
| EP3908943A1 (en) | Method, apparatus, electronic device and computer readable storage medium for image searching | |
| EP4367628A1 (en) | Image processing method and related device | |
| WO2019031714A1 (ko) | 객체를 인식하는 방법 및 장치 | |
| WO2021006482A1 (en) | Apparatus and method for generating image | |
| WO2020027519A1 (ko) | 영상 처리 장치 및 그 동작방법 | |
| WO2022139111A1 (ko) | 초분광 데이터에 기반하는 해상객체 인식 방법 및 시스템 | |
| WO2023085862A1 (en) | Image processing method and related device | |
| WO2023243892A1 (en) | Image processing method and related device | |
| WO2022124865A1 (en) | Method, device, and computer program for detecting boundary of object in image | |
| WO2015137666A1 (ko) | 오브젝트 인식 장치 및 그 제어 방법 | |
| WO2019009664A1 (en) | APPARATUS FOR OPTIMIZING THE INSPECTION OF THE OUTSIDE OF A TARGET OBJECT AND ASSOCIATED METHOD | |
| WO2020116988A1 (ko) | 영상 분석 장치, 영상 분석 방법 및 기록 매체 | |
| WO2022080844A1 (ko) | 스켈레톤 분석을 이용한 객체 추적 장치 및 방법 | |
| WO2022045613A1 (ko) | 비디오 품질 향상 방법 및 장치 | |
| WO2024158073A1 (ko) | 학습 모델을 이용한 영상들 간의 유사도 판단 | |
| WO2021157880A1 (ko) | 전자 장치 및 데이터 처리 방법 | |
| WO2024186178A1 (ko) | 어노말리 디텍션을 위한 패치 피처 학습 방법 및 그 시스템 | |
| WO2024014870A1 (en) | Method and electronic device for interactive image segmentation | |
| WO2023287091A1 (en) | Method and apparatus for processing image | |
| WO2023146195A1 (ko) | 이미지를 분류하는 서버 및 그 동작 방법 | |
| WO2023059042A1 (en) | Method, system and apparatus for image orientation correction | |
| WO2021137533A1 (en) | Video sampling method and apparatus using the same | |
| WO2019074185A1 (en) | ELECTRONIC APPARATUS AND CONTROL METHOD THEREOF |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23918709 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 23918709 Country of ref document: EP Kind code of ref document: A1 |