WO2020111844A2 - 물체 레이블을 활용한 비주얼 슬램에서의 영상 특징점 강화 방법 및 장치 - Google Patents

물체 레이블을 활용한 비주얼 슬램에서의 영상 특징점 강화 방법 및 장치 Download PDF

Info

Publication number
WO2020111844A2
WO2020111844A2 PCT/KR2019/016641 KR2019016641W WO2020111844A2 WO 2020111844 A2 WO2020111844 A2 WO 2020111844A2 KR 2019016641 W KR2019016641 W KR 2019016641W WO 2020111844 A2 WO2020111844 A2 WO 2020111844A2
Authority
WO
WIPO (PCT)
Prior art keywords
frame
visual
feature point
image
word vector
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/KR2019/016641
Other languages
English (en)
French (fr)
Other versions
WO2020111844A3 (ko
Inventor
장병탁
이충연
이현도
황인준
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
SNU R&DB Foundation
Original Assignee
Seoul National University R&DB Foundation
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Priority claimed from KR1020190039736A external-priority patent/KR102285427B1/ko
Application filed by Seoul National University R&DB Foundation filed Critical Seoul National University R&DB Foundation
Publication of WO2020111844A2 publication Critical patent/WO2020111844A2/ko
Publication of WO2020111844A3 publication Critical patent/WO2020111844A3/ko
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05DSYSTEMS FOR CONTROLLING OR REGULATING NON-ELECTRIC VARIABLES
    • G05D1/00Control of position, course, altitude or attitude of land, water, air or space vehicles, e.g. using automatic pilots
    • G05D1/02Control of position or course in two dimensions

Definitions

  • Embodiments disclosed herein relates to a method and apparatus for strengthening an image feature point in Visual SLAM (Simultaneous Localization and Mapping) using an object label, and more specifically, for a robot to travel in an unknown space based on a camera image.
  • the present invention relates to a method and apparatus for generating a map and estimating its location.
  • Visual SLAM extracts recognizable feature points from the camera image and maps them to a 3D map. It uses the method, and compares the image feature points of the current view with the existing map to estimate the current position as the point with the highest matching result.
  • the image processing-based method has an advantage of high processing speed, but has an inherent limitation that the amount of information extractable from the image is insufficient.
  • Korean Patent Publication No. 10-2016-0111008 discloses translational motion in camera movement using at least one motion sensor while the camera is performing panoramic SLAM (simultaneous localization and mapping). And presents techniques for monocular visual SLAM based on initializing a 3D map for tracking of finite features, but does not solve the above-mentioned problem.
  • the above-mentioned background technology is the technical information acquired by the inventor for the derivation of the present invention or acquired in the derivation process of the present invention, and is not necessarily a known technology disclosed to the general public before filing the present invention. .
  • the embodiments disclosed herein have an object to present a method and apparatus for enhancing an image feature point in Visual SLAM using an object label.
  • the embodiments disclosed herein have an object of presenting a method and apparatus for strengthening an image feature point in Visual SLAM using an object label that extracts semantic information of an object based on object recognition and combines it with an image feature point. .
  • the embodiments disclosed herein have an object to present a method and apparatus for enhancing image feature points in Visual SLAM using an object label that improves recognition performance of a previous visited place by additionally using semantic information of the object.
  • a feature point is extracted from a frame constituting an image obtained through a camera, and the extracted feature point
  • a visual feature processing unit that generates a visual word vector that records the distribution of descriptors including information about characteristics
  • a semantic feature processing unit that recognizes an object located in the feature point region from the frame and extracts semantic information about the recognized object
  • a location estimation may be performed to identify whether the current location is a previously visited location based on the visual semantic word vector to which the semantic information is added to the visual word vector.
  • a feature point is extracted from a frame constituting an image acquired through a camera, and includes information on characteristics of the extracted feature point.
  • Generating a visual word vector recording the distribution of descriptors, recognizing an object located in the feature point area from the frame, extracting semantic information about the recognized object, and visually adding the semantic information to the visual word vector And identifying whether the current location is a previously visited location based on the semantic word vector.
  • the image feature enhancement method extracts and extracts feature points from a frame constituting an image acquired through a camera.
  • Generating a visual word vector recording the distribution of descriptors including information on the characteristics of the feature points, recognizing an object located in the feature point area from the frame, and extracting semantic information about the recognized object; and the visual And identifying whether the current location is a previously visited point based on the visual semantic word vector to which the semantic information is added to the word vector.
  • a computer program performed by an image feature enhancement device and stored in a recording medium to perform the image feature enhancement method, wherein the image feature enhancement method is a frame constituting an image acquired through a camera Extracting a feature point from and generating a visual word vector recording a distribution of descriptors including information on the characteristics of the extracted feature point, recognizing an object located in the feature point area from the frame, and semantic information about the recognized object
  • the method may include extracting and identifying whether the current location is a previously visited point based on the visual semantic word vector to which the semantic information is added to the visual word vector.
  • the method and apparatus for enhancing the image feature points in Visual SLAM using the object label that correctly recognizes the place of visit using the semantic information of the object can be presented.
  • FIG. 1 is a block diagram illustrating an image feature enhancement device according to an embodiment.
  • FIGS. 2 and 3 are flowcharts for explaining a method for enhancing an image feature point according to an embodiment.
  • SLAM Simultaneous Localization and Mapping
  • The'key frame' is a frame used to identify the current position of the robot among the frames constituting the image obtained through the camera provided in the robot.
  • 'Bag of Words' is a technique for converting text input into numbers using groups in which a plurality of texts are grouped.
  • The'feature point' is a point in which the intensity value differs from the surrounding points in an image captured by the camera.
  • FIG. 1 is a block diagram illustrating an image feature enhancement device 10 according to an embodiment.
  • the image feature enhancement device 10 may be implemented as a computer or a portable terminal, a television, a wearable device, or the like that connects to a remote server through a network N or connects to other terminals and servers.
  • the computer includes, for example, a laptop equipped with a web browser (WEB Browser), a desktop (desktop), a laptop (laptop), and the like
  • the portable terminal is, for example, a wireless communication device that guarantees portability and mobility.
  • the television may include Internet Protocol Television (IPTV), Internet Television (TV), terrestrial TV, and cable TV.
  • IPTV Internet Protocol Television
  • TV Internet Television
  • the wearable device is a type of information processing device that can be directly worn on the human body, for example, a watch, glasses, accessories, clothing, shoes, etc., to connect to a remote server or other terminal through a network directly or through another information processing device. And can be connected.
  • the image feature enhancement device 10 may include an input/output unit 110, a control unit 120, a communication unit 130, and a memory 140.
  • the input/output unit 110 may include an input unit for receiving input from a user, and an output unit for displaying information such as a result of performing a job or a state of the image feature enhancement device 10.
  • the input/output unit 110 may include an operation panel for receiving user input, a display panel for displaying the screen, and the like.
  • the input unit may include devices capable of receiving various types of user input, such as a keyboard, a physical button, a touch screen, a camera, or a microphone.
  • a camera sensor may be provided on the robot to photograph the surroundings of the robot.
  • the output unit may include a display panel or a speaker.
  • the present invention is not limited thereto, and the input/output unit 110 may include a configuration supporting various input/output.
  • the control unit 120 controls the overall operation of the image feature enhancement device 10, and may include a processor such as a CPU.
  • the control unit 120 may control other components included in the image feature enhancement device 10 to perform an operation corresponding to a user input received through the input/output unit 110.
  • the controller 120 may execute a program stored in the memory 140, read a file stored in the memory 140, or store a new file in the memory 140.
  • the control unit 120 may include a visual feature processing unit 121, a semantic feature processing unit 122, a location estimation unit 123, or a map updating unit 124.
  • the visual feature processing unit 121 may extract a feature point of a frame from a frame of an image photographed by the input/output unit 110, and store a characteristic of each feature point by calculating a descriptor for each extracted feature point Can be.
  • the visual feature processing unit 121 groups the descriptors into K-means clustering to generate a visual vocabulary included in a bag of words (Visual vocabulary vector recording the distribution of descriptors for each frame of the image) ), and the characteristics of each frame may be stored through the generated visual word vector.
  • the semantic feature processing unit 122 may generate a visual-semantic vocabulary vector by adding semantic information, which is an object label of the region to which the feature point belongs, to the visual word vector of the descriptor.
  • the visual-semantic vocabulary vector is expressed as follows.
  • the location estimator 123 may calculate the similarity between frames to identify whether the frame of the image captured at the current location is a previously visited location using the frame of the captured image.
  • the location estimator 123 may select a candidate key frame to compare the similarity with the frame of the current location, and compare the visual semantic word vector of the key frame to compare the key frame with a word distribution similar to the current frame. Similarity can be selected as a key frame to be compared.
  • the position estimator 123 can count the number of common object labels between each key frame of the captured image and the frame f of the current position, and if the number of common labels with the frame f is not 0, the key frame k It can be set as the first candidate frame.
  • the location estimator 123 extracts a key frame k at a preset ratio based on the number of object labels common to the frame f of the current position among the key frames k included in the first candidate frame, and can be set as the second candidate frame. have.
  • the location estimator 123 may set the key frame k as the second candidate frame based on the number of object labels common to the frame f of the current position in the key frame k included in the first candidate frame, which The formula is as follows.
  • C1 first candidate frame
  • C2 second candidate frame
  • the location estimator 123 is a neighboring frame of the key frame k belonging to the second candidate frame.
  • Set to, neighbor frame A frame similar to the current frame f among the key frames k'belonging to may be set as a third candidate frame.
  • the location estimator 123 is a neighboring frame
  • the frame k'most similar to the current frame f is calculated by using L1_norm whether the visual semantic word vector belonging to the current frame f and the visual semantic word vector belonging to the frame k'belong to the frame k'belonging to the third candidate frame Can be set to
  • the location estimator 123 may set a frame similar to the vector from the current frame f among the key frames belonging to the third candidate frame as a candidate key frame.
  • the location estimator 123 selects a candidate frame for comparing the similarity with the current frame in order to identify the current location, it is as follows.
  • the location estimator 123 may calculate the similarity with the current frame based on the selected candidate key frame.
  • the location estimator 123 may calculate the similarity between frames by calculating the similarity between descriptors using the visual semantic word vector of the descriptor of the feature points extracted from each of the current frame and the selected candidate key frame.
  • the location estimator 123 uses a Hamming distance method, but if the object labels included in the visual meaning word vectors of the two descriptors are different, the penalty is calculated. By multiplying by, the similarity between descriptors can be calculated according to the following formula.
  • the location estimator 123 may identify whether the current location is a previously visited location based on the similarity between the current frame and the candidate key frame.
  • the map update unit 124 may generate a map based on the captured image and update the map when the current location is identified as a previously visited point.
  • the map updater 124 may update the map by closing the regression point on the map.
  • the communication unit 130 may perform wired/wireless communication with other devices or networks.
  • the communication unit 130 may include a communication module supporting at least one of various wired and wireless communication methods.
  • the communication module may be implemented in the form of a chipset.
  • the wireless communication supported by the communication unit 130 may be, for example, Wi-Fi (Wireless Fidelity), Wi-Fi Direct, Bluetooth (Bluetooth), UWB (Ultra Wide Band), or NFC (Near Field Communication).
  • the wired communication supported by the communication unit 130 may be, for example, USB or HDMI (High Definition Multimedia Interface).
  • the controller 120 may access and use data stored in the memory 140, or may store new data in the memory 140. Also, the control unit 120 may execute a program installed in the memory 140.
  • the memory 140 may store an image captured by the input/output unit 110.
  • FIG. 2 is a flowchart illustrating an image feature enhancement method according to an embodiment.
  • the image feature point enhancement method according to the embodiment shown in FIG. 2 includes steps that are time-sequentially processed by the image feature enhancement device 10 shown in FIG. 1. Therefore, even if it is omitted below, the description of the image feature point enhancement device 10 shown in FIG. 1 above can also be applied to the image feature point enhancement method according to the embodiment shown in FIG. 2.
  • the image feature point enhancement device 10 may acquire an image through a camera sensor and extract a feature point from the acquired image (S2001).
  • the image feature point enhancement device 10 may extract a point having a difference between a surrounding point and brightness within each frame constituting the image as a feature point.
  • the image feature enhancement apparatus 10 may generate a visual word vector that records the distribution of descriptors including information on the characteristics of the extracted feature points (S2002).
  • the image feature enhancement apparatus 10 may generate visual vocabulary of a bag of words concept by clustering technicians with K-means.
  • the image feature enhancement device 10 may match the descriptor corresponding to the feature point on the frame to the visual word, and generate a visual word vector with a value corresponding to the matched visual word to store the characteristics of the frame.
  • the image feature enhancement device 10 may recognize an object located in the feature point region in the frame and generate a visual meaning word vector by adding semantic information as a label for the recognized object to the visual word vector (S2003).
  • the image feature enhancement device 10 may recognize an object located within a preset radius based on the feature point, and set a label corresponding to the recognized object among the pre-stored labels as semantic information.
  • the image feature enhancement apparatus 10 may generate a visual meaning word vector including meaning information in the visual word vector by including the object label as the last value of the descriptor for the feature point of the region where the object is located.
  • the image feature enhancement apparatus 10 may calculate the similarity between the current frame and the frame included in the pre-stored image based on the visual semantic word vector to identify whether the current location is the pre-visited location.
  • the feature point enhancement device 10 may determine a candidate key frame for calculating the similarity with the current frame (S2004).
  • the image feature enhancement apparatus 10 may set a first candidate frame in a key frame of a pre-stored image based on the object label of the current frame (S3001).
  • the image feature enhancement apparatus 10 may set a frame including at least one object label identical to an object label included in a current frame among key frames of a pre-stored image as a first candidate frame.
  • the image feature enhancement apparatus 10 may set a frame in which the number of the first candidate frames including the object label equal to the object label of the current frame exceeds a preset number as the second candidate frame (S3002).
  • the image feature enhancement apparatus 10 may set the first candidate frame belonging to the top 20% of the first candidate frame as the second candidate frame based on the number of object labels equal to the object label of the current frame. .
  • the image feature enhancement apparatus 10 may set a neighboring frame of the second candidate frame similar to the current frame as the third candidate frame based on the size of the visual word vector of each of the current frames among the neighboring frames of the second candidate frame. (S3003).
  • the image feature enhancement apparatus 10 may set a neighbor frame of the second candidate frame and a neighbor frame having the largest vector size, which is an L1_Norm value between current frames, as the third candidate frame.
  • the image feature enhancement apparatus 10 may finally set the candidate key frame using the size of the current frame and the visual semantic word vector from the third candidate frame (S3004).
  • the image feature enhancement apparatus 10 may align the third candidate frame based on the L1_Norm value of the current frame and the neighboring frame, and the size of the visual semantic word vector with the current frame exceeds a preset value.
  • a candidate key frame may be selected as the third candidate frame.
  • the image feature enhancement apparatus 10 may calculate whether the current location is a previously visited location by calculating similarity based on the difference between the determined candidate key frame and the value of each visual semantic word vector of the frame at the current location ( S2005).
  • the image feature enhancement apparatus 10 may calculate the similarity between the current frame and the candidate key frame by calculating a Hamming distance between the visual semantic word vector of the current frame and the visual semantic word vector of the candidate key frame. .
  • the image feature enhancement device 10 may update the map generated based on the image (S2006).
  • the image feature point enhancement device 10 may generate a map based on the captured image, detect a regression point according to whether the current location is a previously visited location in step S2005, and close the regression point. You can update the map.
  • the term' ⁇ unit' used in the above embodiments means software or hardware components such as a field programmable gate array (FPGA) or an ASIC, and' ⁇ unit' performs certain roles. However,' ⁇ wealth' is not limited to software or hardware.
  • The' ⁇ unit' may be configured to be in an addressable storage medium or may be configured to reproduce one or more processors. Thus, as an example,' ⁇ unit' refers to components such as software components, object-oriented software components, class components and task components, processes, functions, attributes, and procedures. , Subroutines, segments of program patent code, drivers, firmware, microcode, circuitry, data, database, data structures, tables, arrays, and variables.
  • the functions provided within the components and' ⁇ units' may be combined into a smaller number of components and' ⁇ units', or separated from additional components and' ⁇ units'.
  • components and' ⁇ unit' may be implemented to play one or more CPUs in the device or secure multimedia card.
  • the image feature point enhancement method according to the embodiment described with reference to FIGS. 2 and 3 may also be implemented in the form of a computer-readable medium that stores instructions and data executable by a computer.
  • instructions and data may be stored in the form of program code, and when executed by a processor, a predetermined program module may be generated to perform a predetermined operation.
  • the computer-readable medium can be any available medium that can be accessed by a computer, and includes both volatile and nonvolatile media, removable and non-removable media.
  • the computer-readable medium may be a computer recording medium, which is implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data. It can include both volatile, removable and non-removable media.
  • computer recording media are magnetic storage media such as HDDs and SSDs, optical recording media such as CDs, DVDs and Blu-ray Discs, or are accessible over a network. It may be a memory included in the server.
  • the image feature point enhancement method according to the embodiment described with reference to FIGS. 2 and 3 may be implemented as a computer program (or computer program product) including instructions executable by a computer.
  • the computer program includes programmable machine instructions processed by a processor and may be implemented in a high-level programming language, object-oriented programming language, assembly language, or machine language.
  • the computer program may be recorded on a tangible computer-readable recording medium (eg, memory, hard disk, magnetic/optical medium, or solid-state drive (SSD)).
  • the image feature point enhancement method according to the embodiment described with reference to FIGS. 2 and 3 may be implemented by executing a computer program as described above by a computing device.
  • the computing device may include at least some of a processor, a memory, a storage device, a high-speed interface connected to the memory and a high-speed expansion port, and a low-speed interface connected to the low-speed bus and the storage device.
  • a processor may include at least some of a processor, a memory, a storage device, a high-speed interface connected to the memory and a high-speed expansion port, and a low-speed interface connected to the low-speed bus and the storage device.
  • Each of these components is connected to each other using various buses, and can be mounted on a common motherboard or mounted in other suitable ways.
  • the processor is capable of processing instructions within the computing device, such as to display graphical information for providing a graphical user interface (GUI) on an external input or output device, such as a display connected to a high-speed interface. Examples are commands stored in memory or storage devices. In other embodiments, multiple processors and/or multiple buses may be used in conjunction with multiple memories and memory types as appropriate. Also, the processor may be implemented as a chipset formed by chips including a plurality of independent analog and/or digital processors.
  • Memory also stores information within computing devices.
  • the memory may consist of volatile memory units or a collection thereof.
  • the memory may consist of non-volatile memory units or a collection thereof.
  • the memory may also be other forms of computer-readable media, such as magnetic or optical disks.
  • the storage device can provide a large storage space for the computing device.
  • the storage device may be a computer-readable medium or a configuration including such a medium, and may include, for example, devices within a storage area network (SAN) or other configurations, and may include floppy disk devices, hard disk devices, optical disk devices, Or a tape device, flash memory, or other similar semiconductor memory device or device array.
  • SAN storage area network

Landscapes

  • Engineering & Computer Science (AREA)
  • Aviation & Aerospace Engineering (AREA)
  • Radar, Positioning & Navigation (AREA)
  • Remote Sensing (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Automation & Control Theory (AREA)
  • Image Analysis (AREA)

Abstract

물체 레이블을 활용한 Visual SLAM에서의 영상 특징점 강화 방법 및 장치를 제시하며, 물체 레이블을 활용한 Visual SLAM에서의 영상 특징점 강화 방법 및 장치는 카메라를 통해 획득된 영상을 구성하는 프레임으로부터 특징점을 추출하고, 추출된 특징점의 특성에 대한 정보를 포함하는 기술자의 분포를 기록한 시각적 단어 벡터를 생성하는 시각적특징처리부, 상기 프레임으로부터 상기 특징점 영역에 위치한 물체를 인식하고, 인식된 물체에 대한 의미정보를 추출하는 의미적특징처리부 및 상기 시각적 단어 벡터에 상기 의미정보가 추가된 시각적 의미 단어 벡터를 기초로 현재 위치가 기 방문한 지점인지 식별하는 위치추정부를 포함할 수 있다.

Description

[규칙 제26조에 의한 보정 26.12.2019] 물체 레이블을 활용한 비주얼 슬램에서의 영상 특징점 강화 방법 및 장치
본 명세서에서 개시되는 실시예들은 물체 레이블을 활용한 Visual SLAM(Simultaneous Localization and Mapping)에서의 영상 특징점 강화 방법 및 장치에 관한 것으로, 보다 상세하게는 카메라 영상 기반으로 미지의 공간에서 로봇이 주행하기 위해 지도를 생성하는 동시에 자신의 위치를 추정하는 방법 및 장치에 관한 것이다.
1-1. 과제고유번호 : 1711081135
1-2. 사사표기 : 본 연구는 과학기술정보통신부 및 정보통신기획평가원의 혁신성장동력프로젝트사업의 연구결과로 수행되었음(IITP-2017-0-01772-003).
2-1. 과제고유번호 : 1711081008
2-2. 사사표기 : 본 연구는 과학기술정보통신부 및 정보통신기획평가원의 SW컴퓨팅산업원천기술개발사업의 연구결과로 수행되었음(IITP-2015-0-00310-005).
일반적으로 사람이 공간 인식 과정에서 주변 환경에 배치된 물체의 종류와 색상 등 다양한 의미적 특징들을 기억하고 활용하는 것과 달리, Visual SLAM은 카메라의 영상으로부터 인식 가능한 특징점들을 추출하여 3차원 지도에 매핑하는 방식을 사용하며, 현재 시점의 영상 특징점들을 기존 지도와 비교하여 매칭 결과가 가장 높은 지점으로 현재 위치를 추정한다.
하지만 영상처리 기반의 방법은 처리 속도가 빠른 장점이 있지만, 영상으로부터 추출 가능한 정보량이 부족한 태생적인 한계를 지닌다.
즉, 특정 픽셀과 그 주변 픽셀들의 밝기값으로부터 추출된 국소적인 정보만을 사용하기 때문에 모바일 로봇의 움직임에 따른 영상 변화에 민감하며, 이로 인해 충분한 영상 특징점이 추출되지 않는 경우 위치 추정에 실패하고, 이전에 방문한 장소를 올바로 인식하지 못하게 되는 문제가 있다.
관련하여 선행기술 문헌인 한국특허공개번호 제 10-2016-0111008 호에서는 카메라가 파노라마 SLAM(simultaneous localization and mapping)을 수행하고 있는 동안, 적어도 하나의 모션 센서를 사용하여 카메라의 움직임에서 병진 모션을 검출하고 유한 피쳐들의 추적을 위해 3차원 맵을 초기화하는 것에 기초하는 단안 시각적 SLAM을 위한 기술들을 제시할 뿐, 상술된 문제점을 해결하지 못한다.
따라서 상술된 문제점을 해결하기 위한 기술이 필요하게 되었다.
한편, 전술한 배경기술은 발명자가 본 발명의 도출을 위해 보유하고 있었거나, 본 발명의 도출 과정에서 습득한 기술 정보로서, 반드시 본 발명의 출원 전에 일반 공중에게 공개된 공지기술이라 할 수는 없다.
본 명세서에서 개시되는 실시예들은, 물체 레이블을 활용한 Visual SLAM에서의 영상 특징점 강화 방법 및 장치를 제시하는 데 목적이 있다.
본 명세서에서 개시되는 실시예들은, 물체 인식을 기반으로 물체의 의미적 정보를 추출하여 영상의 특징점에 결합하는 물체 레이블을 활용한 Visual SLAM에서의 영상 특징점 강화 방법 및 장치를 제시하는 데 목적이 있다.
본 명세서에서 개시되는 실시예들은, 물체의 의미적 정보를 추가적으로 이용함으로써 이전 방문 장소의 인식 성능을 향상시키는 물체 레이블을 활용한 Visual SLAM에서의 영상 특징점 강화 방법 및 장치를 제시하는 데 목적이 있다.
상술한 기술적 과제를 달성하기 위한 기술적 수단으로서, 일 실시예에 따르면, SLAM에서 영상의 특징점을 강화하는 장치에 있어서, 카메라를 통해 획득된 영상을 구성하는 프레임으로부터 특징점을 추출하고, 추출된 특징점의 특성에 대한 정보를 포함하는 기술자의 분포를 기록한 시각적 단어 벡터를 생성하는 시각적특징처리부, 상기 프레임으로부터 상기 특징점 영역에 위치한 물체를 인식하고, 인식된 물체에 대한 의미정보를 추출하는 의미적특징처리부 및 상기 시각적 단어 벡터에 상기 의미정보가 추가된 시각적 의미 단어 벡터를 기초로 현재 위치가 기 방문한 지점인지 식별하는 위치추정부를 포함할 수 있다.
다른 실시예에 따르면, 영상특징점강화장치가 SLAM에서 영상의 특징점을 강화하는 방법에 있어서, 카메라를 통해 획득된 영상을 구성하는 프레임으로부터 특징점을 추출하고, 추출된 특징점의 특성에 대한 정보를 포함하는 기술자의 분포를 기록한 시각적 단어 벡터를 생성하는 단계, 상기 프레임으로부터 상기 특징점 영역에 위치한 물체를 인식하고, 인식된 물체에 대한 의미정보를 추출하는 단계 및 상기 시각적 단어 벡터에 상기 의미정보가 추가된 시각적 의미 단어 벡터를 기초로 현재 위치가 기 방문한 지점인지 식별하는 단계를 포함할 수 있다.
또 다른 실시예에 따르면, 영상특징점강화방법을 수행하는 프로그램이 기록된 컴퓨터 판독이 가능한 기록매체로서, 상기 영상특징점강화방법은, 카메라를 통해 획득된 영상을 구성하는 프레임으로부터 특징점을 추출하고, 추출된 특징점의 특성에 대한 정보를 포함하는 기술자의 분포를 기록한 시각적 단어 벡터를 생성하는 단계, 상기 프레임으로부터 상기 특징점 영역에 위치한 물체를 인식하고, 인식된 물체에 대한 의미정보를 추출하는 단계 및 상기 시각적 단어 벡터에 상기 의미정보가 추가된 시각적 의미 단어 벡터를 기초로 현재 위치가 기 방문한 지점인지 식별하는 단계를 포함할 수 있다.
다른 실시예에 따른면, 영상특징점강화장치에 의해 수행되며, 상기 영상특징점강화방법을 수행하기 위해 기록매체에 저장된 컴퓨터프로그램으로서, 상기 영상특징점강화방법은, 카메라를 통해 획득된 영상을 구성하는 프레임으로부터 특징점을 추출하고, 추출된 특징점의 특성에 대한 정보를 포함하는 기술자의 분포를 기록한 시각적 단어 벡터를 생성하는 단계, 상기 프레임으로부터 상기 특징점 영역에 위치한 물체를 인식하고, 인식된 물체에 대한 의미정보를 추출하는 단계 및 상기 시각적 단어 벡터에 상기 의미정보가 추가된 시각적 의미 단어 벡터를 기초로 현재 위치가 기 방문한 지점인지 식별하는 단계를 포함할 수 있다.
전술한 과제 해결 수단 중 어느 하나에 의하면, 물체 레이블을 활용한 Visual SLAM에서의 영상 특징점 강화 방법 및 장치를 제시할 수 있다.
전술한 과제 해결 수단 중 어느 하나에 의하면, 영상의 특징점이 추출되지 않는 경우라도 물체의 의미적 정보를 이용하여 방문장소를 올바르게 인식하는 물체 레이블을 활용한 Visual SLAM에서의 영상 특징점 강화 방법 및 장치를 제시할 수 있다.
전술한 과제 해결 수단 중 어느 하나에 의하면, 물체의 의미적 정보를 추가적으로 이용함으로써 이전 방문 장소의 인식 성능을 향상시키는 물체 레이블을 활용한 Visual SLAM에서의 영상 특징점 강화 방법 및 장치를 제시
개시되는 실시예들에서 얻을 수 있는 효과는 이상에서 언급한 효과들로 제한되지 않으며, 언급하지 않은 또 다른 효과들은 아래의 기재로부터 개시되는 실시예들이 속하는 기술분야에서 통상의 지식을 가진 자에게 명확하게 이해될 수 있을 것이다.
도 1 은 일 실시예에 따른 영상특징점강화장치를 도시한 블록도이다.
도 2 및 도 3 은 일 실시예에 따른 영상특징점강화방법을 설명하기 위한 순서도이다.
아래에서는 첨부한 도면을 참조하여 다양한 실시예들을 상세히 설명한다. 아래에서 설명되는 실시예들은 여러 가지 상이한 형태로 변형되어 실시될 수도 있다. 실시예들의 특징을 보다 명확히 설명하기 위하여, 이하의 실시예들이 속하는 기술분야에서 통상의 지식을 가진 자에게 널리 알려져 있는 사항들에 관해서 자세한 설명은 생략하였다. 그리고, 도면에서 실시예들의 설명과 관계없는 부분은 생략하였으며, 명세서 전체를 통하여 유사한 부분에 대해서는 유사한 도면 부호를 붙였다.
명세서 전체에서, 어떤 구성이 다른 구성과 "연결"되어 있다고 할 때, 이는 ‘직접적으로 연결’되어 있는 경우뿐 아니라, ‘그 중간에 다른 구성을 사이에 두고 연결’되어 있는 경우도 포함한다. 또한, 어떤 구성이 어떤 구성을 "포함"한다고 할 때, 이는 특별히 반대되는 기재가 없는 한, 그 외 다른 구성을 제외하는 것이 아니라 다른 구성들을 더 포함할 수도 있음을 의미한다.
이하 첨부된 도면을 참고하여 실시예들을 상세히 설명하기로 한다.
다만 이를 설명하기에 앞서, 아래에서 사용되는 용어들의 의미를 먼저 정의한다.
‘SLAM(Simultaneous Localization and Mapping)’이란 미지의 공간에서 로봇이 주행하기 위해 지도를 생성하는 동시에 자신의 위치를 추정하는 기술이다. 그리고 카메라 영상 기반의 SLAM을 Visual SLAM 이라고 한다.
‘키프레임’은 로봇에 구비된 카메라를 통해 획득된 영상을 구성하는 프레임 중 로봇의 현재 위치를 식별하는데 이용되는 프레임이다.
‘단어의 가방(Bag Of Words)’은 복수의 텍스트가 그룹핑된 그룹을 이용하여 입력되는 텍스트를 숫자로 변환하는 기법이다.
‘특징점’은 카메라를 통해 촬영된 영상에서 주변 점들과 강도(intensity)값 차이가 큰 점이다.
위에 정의한 용어 이외에 설명이 필요한 용어는 아래에서 각각 따로 설명한다.
도 1은 일 실시예에 따른 영상특징점강화장치(10)를 설명하기 위한 블럭도이다.
영상특징점강화장치(10)는 네트워크(N)를 통해 원격지의 서버에 접속하거나, 타 단말 및 서버와 연결 가능한 컴퓨터나 휴대용 단말기, 텔레비전, 웨어러블 디바이스(Wearable Device) 등으로 구현될 수 있다. 여기서, 컴퓨터는 예를 들어, 웹 브라우저(WEB Browser)가 탑재된 노트북, 데스크톱(desktop), 랩톱(laptop)등을 포함하고, 휴대용 단말기는 예를 들어, 휴대성과 이동성이 보장되는 무선 통신 장치로서, PCS(Personal Communication System), PDC(Personal Digital Cellular), PHS(Personal Handyphone System), PDA(Personal Digital Assistant), GSM(Global System for Mobile communications), IMT(International Mobile Telecommunication)-2000, CDMA(Code Division Multiple Access)-2000, W-CDMA(W-Code Division Multiple Access), Wibro(Wireless Broadband Internet), 스마트폰(Smart Phone), 모바일 WiMAX(Mobile Worldwide Interoperability for Microwave Access) 등과 같은 모든 종류의 핸드헬드(Handheld) 기반의 무선 통신 장치를 포함할 수 있다. 또한, 텔레비전은 IPTV(Internet Protocol Television), 인터넷 TV(Internet Television), 지상파 TV, 케이블 TV 등을 포함할 수 있다. 나아가 웨어러블 디바이스는 예를 들어, 시계, 안경, 액세서리, 의복, 신발 등 인체에 직접 착용 가능한 타입의 정보처리장치로서, 직접 또는 다른 정보처리장치를 통해 네트워크를 경유하여 원격지의 서버에 접속하거나 타 단말과 연결될 수 있다.
도 1 을 참조하면, 일 실시예에 따른 영상특징점강화장치(10)는, 입출력부(110), 제어부(120), 통신부(130) 및 메모리(140)를 포함할 수 있다.
우선, 입출력부(110)는 사용자로부터 입력을 수신하기 위한 입력부와, 작업의 수행 결과 또는 영상특징점강화장치(10)의 상태 등의 정보를 표시하기 위한 출력부를 포함할 수 있다. 예를 들어, 입출력부(110)는 사용자 입력을 수신하는 조작 패널(operation panel) 및 화면을 표시하는 디스플레이 패널(display panel) 등을 포함할 수 있다.
구체적으로, 입력부는 키보드, 물리 버튼, 터치 스크린, 카메라 또는 마이크 등과 같이 다양한 형태의 사용자 입력을 수신할 수 있는 장치들을 포함할 수 있다. 예를 들어, 카메라 센서는 로봇에 구비되어 로봇의 주위를 촬영할 수 있다.
또한, 출력부는 디스플레이 패널 또는 스피커 등을 포함할 수 있다. 다만, 이에 한정되지 않고 입출력부(110)는 다양한 입출력을 지원하는 구성을 포함할 수 있다.
제어부(120)는 영상특징점강화장치(10)의 전체적인 동작을 제어하며, CPU 등과 같은 프로세서를 포함할 수 있다. 제어부(120)는 입출력부(110)를 통해 수신한 사용자 입력에 대응되는 동작을 수행하도록 영상특징점강화장치(10)에 포함된 다른 구성들을 제어할 수 있다.
예를 들어, 제어부(120)는 메모리(140)에 저장된 프로그램을 실행시키거나, 메모리(140)에 저장된 파일을 읽어오거나, 새로운 파일을 메모리(140)에 저장할 수도 있다.
이러한 제어부(120)는 시각적특징처리부(121), 의미적특징처리부(122), 위치추정부(123) 또는 지도갱신부(124)를 포함할 수 있다.
우선, 시각적특징처리부(121)는 입출력부(110)에서 의해 촬영된 영상의 프레임으로부터 프레임의 특징점을 추출할 수 있고, 추출된 특징점 각각에 대해 기술자(descriptor)를 계산하여 각 특징점의 특성을 저장할 수 있다.
이를 위해, 시각적특징처리부(121)는 기술자를 K-means 클러스터링으로 그룹핑하여 단어의 가방에 포함되는 시각적 단어(Visual vocabulary)를 생성하여 영상의 프레임별 기술자의 분포를 기록한 시각적 단어 벡터(Visual vocabulary vector)를 생성할 수 있고, 생성된 시각적 단어 벡터를 통해 각 프레임의 특성을 저장할 수 있다.
그리고 의미적특징처리부(122)는 특징점이 속한 영역의 물체 레이블인 의미 정보를 기술자의 시각적 단어 벡터에 추가하여 시각적 의미 단어 벡터(visual-semantic vocabulary vector)를 생성할 수 있다.
이때, 시각적 의미 단어 벡터(visual-semantic vocabulary vector)를 수식으로 나타내면 아래와 같다.
프레임 f 에 포함된 의미 단어(semantic vocabulary)와 시각적 단어(visual vocabulary)의 집합을 각각
Figure PCTKR2019016641-appb-img-000001
,
Figure PCTKR2019016641-appb-img-000002
로 두고, f에 포함된 특징점들 중 의미 단어 l을 가지는 집합을
Figure PCTKR2019016641-appb-img-000003
, 시각적 단어 i 를 가지는 집합을
Figure PCTKR2019016641-appb-img-000004
로 두었을 때, 시각적 의미 단어 벡터(visual-semantic vocabulary vector)
Figure PCTKR2019016641-appb-img-000005
를 아래의 수식과 같이 나타낼 수 있다.
[수학식 1]
Figure PCTKR2019016641-appb-img-000006
여기서
Figure PCTKR2019016641-appb-img-000007
는 단어의 가방에서 단어 i의 가중치이고,
Figure PCTKR2019016641-appb-img-000008
는 시각적 단어 벡터와 의미 단어 벡터간의 가중치를 조절하는 상수이다.
그리고 위치추정부(123)는 촬영된 영상의 프레임을 이용하여 현재 위치에서 촬영된 영상의 프레임이 이전에 방문한 지점인지를 식별하기 위해 프레임간 유사도를 계산할 수 있다.
이때, 위치추정부(123)는 현재 위치의 프레임과 유사도를 비교할 후보 키 프레임(key frame)을 선택할 수 있으며, 키 프레임의 시각적 의미 단어 벡터를 비교하여 현재 프레임과 비슷한 단어 분포를 가진 키 프레임을 유사도를 비교할 키 프레임으로 선택할 수 있다.
우선, 위치추정부(123)는 촬영된 영상의 모든 키 프레임 각각과 현재 위치의 프레임 f간 공통된 물체 레이블의 수를 카운팅할 수 있고, 프레임 f와의 공통된 레이블의 수가 0이 아니면, 키 프레임 k 를 제 1 후보프레임으로 설정할 수 있다.
그리고 위치추정부(123)는 제 1 후보프레임에 포함된 키 프레임 k 중 현재 위치의 프레임 f와 공통된 물체 레이블의 수를 기초로 기 설정된 비율의 키 프레임 k를 추출하여 제 2 후보프레임으로 설정할 수 있다.
예를 들어, 위치추정부(123)는 제 1 후보프레임에 포함된 키 프레임 k 에서 현재 위치의 프레임 f와 공통된 물체 레이블의 수를 기준으로 키 프레임 k를 제 2 후보프레임으로 설정할 수 있으며, 이를 수식으로 나타내면 아래와 같다.
[수학식 2]
Figure PCTKR2019016641-appb-img-000009
C1 : 제 1 후보프레임, C2 : 제 2 후보프레임
그리고 위치추정부(123)는 제 2 후보프레임에 속한 키 프레임 k의 이웃프레임
Figure PCTKR2019016641-appb-img-000010
로 설정하고, 이웃프레임
Figure PCTKR2019016641-appb-img-000011
에 속하는 키 프레임 k’ 중 현재 프레임 f와 유사한 프레임을 제 3 후보프레임으로 설정할 수 있다.
예를 들어, 위치추정부(123)는 이웃프레임
Figure PCTKR2019016641-appb-img-000012
에 속하는 프레임 k’에 대해 현재 프레임 f 에 속하는 시각적 의미 단어 벡터와 프레임 k’에 속하는 시각적 의미 단어 벡터간 동일여부를 L1_norm을 이용하여 계산하여 현재 프레임 f와 가장 유사한 프레임 k’을 제 3 후보프레임으로 설정할 수 있다.
이후, 위치추정부(123)는 제 3 후보프레임에 속한 키 프레임 중 현재 프레임 f와의 벡터와 유사한 프레임을 후보 키 프레임으로 설정할 수 있다.
이와 같이 위치추정부(123)가 현재 위치를 식별하기 위해 현재 프레임과 유사도를 비교할 후보프레임을 선택하는 알고리즘을 수도코드로 표시하면 아래와 같다.
Figure PCTKR2019016641-appb-img-000013
이와 같이 현재 프레임과 유사도를 비교할 후보 키 프레임을 시각적 단어 벡터를 이용하여 선택함으로써 기존에 지도상에 저장된 모든 키 프레임(key frame)에 대하여 현재 프레임과의 유사도를 비교하지 않아 계산량을 줄일 수 있다.
그리고 위치추정부(123)는 선택된 후보 키 프레임을 기준으로 현재 프레임과의 유사도를 계산할 수 있다.
즉, 위치추정부(123)는 현재 프레임과 선택된 후보 키 프레임 각각에서 추출된 특징점의 기술자의 시각적 의미 단어 벡터를 이용하여 기술자간 유사도를 계산함으로써 프레임간 유사도를 계산할 수 있다.
예를 들어, 위치추정부(123)는 해밍거리(Hamming distance) 방법을 이용하되, 두 기술자의 시각적 의미 단어 벡터에 포함된 물체 레이블이 다르면 패널티
Figure PCTKR2019016641-appb-img-000014
를 곱함으로써 기술자간 유사도를 아래의 수식에 따라 계산할 수 있다.
[수학식 3]
Figure PCTKR2019016641-appb-img-000015
그리고 위치추정부(123)는 현재 프레임과 후보 키 프레임간의 유사도를 기초로 현재 위치가 이전에 방문했던 지점인지 여부를 식별할 수 있다.
이를 통해, 프레임의 특징점이 위치한 영역의 물체를 인식하여 의미 정보를 포함한 시각적 의미 단어 벡터를 이용하여 프레임간의 유사여부를 식별함으로써 계산량을 줄이면서도 이전 방문 지점인지 여부를 정확하게 판단할 수 있다.
한편, 지도갱신부(124)는 촬영된 영상에 기초하여 지도를 생성하고, 현재 위치가 기 방문한 지점으로 식별되면 지도를 갱신할 수 있다.
예를 들어, 현재 위치에서 촬영된 영상의 프레임이 기 저장된 영상의 프레임과 유사하여 기 방문 지점으로 판단되면, 지도갱신부(124)는 지도상의 회귀점을 폐쇄하여 지도를 갱신할 수 있다.
통신부(130)는 다른 디바이스 또는 네트워크와 유무선 통신을 수행할 수 있다. 이를 위해, 통신부(130)는 다양한 유무선 통신 방법 중 적어도 하나를 지원하는 통신 모듈을 포함할 수 있다. 예를 들어, 통신 모듈은 칩셋(chipset)의 형태로 구현될 수 있다.
통신부(130)가 지원하는 무선 통신은, 예를 들어 Wi-Fi(Wireless Fidelity), Wi-Fi Direct, 블루투스(Bluetooth), UWB(Ultra Wide Band) 또는 NFC(Near Field Communication) 등일 수 있다. 또한, 통신부(130)가 지원하는 유선 통신은, 예를 들어 USB 또는 HDMI(High Definition Multimedia Interface) 등일 수 있다.
메모리(140)에는 파일, 어플리케이션 및 프로그램 등과 같은 다양한 종류의 데이터가 설치 및 저장될 수 있다. 제어부(120)는 메모리(140)에 저장된 데이터에 접근하여 이를 이용하거나, 또는 새로운 데이터를 메모리(140)에 저장할 수도 있다. 또한, 제어부(120)는 메모리(140)에 설치된 프로그램을 실행할 수도 있다.
이러한 메모리(140)는 입출력부(110)에 의해 촬영된 영상을 저장할 수 있다.
도 2 는 일 실시예에 따른 영상특징점강화방법을 설명하기 위한 순서도이다.
도 2 에 도시된 실시예에 따른 영상특징점강화방법은 도 1 에 도시된 영상특징점강화장치(10)에서 시계열적으로 처리되는 단계들을 포함한다. 따라서, 이하에서 생략된 내용이라고 하더라도 도 1 에 도시된 영상특징점강화장치(10)에 관하여 이상에서 기술한 내용은 도 2 에 도시된 실시예에 따른 영상특징점강화방법에도 적용될 수 있다.
우선, 영상특징점강화장치(10)는 카메라센서를 통해 영상을 획득할 수 있고, 획득된 영상으로부터 특징점을 추출할 수 있다(S2001).
예를 들어, 영상특징점강화장치(10)는 영상을 구성하는 각 프레임 내에서 주변의 지점과 밝기 등의 차이가 있는 지점을 특징점으로 추출할 수 있다.
그리고 영상특징점강화장치(10)는 추출된 특징점의 특성에 대한 정보를 포함하는 기술자의 분포를 기록한 시각적 단어 벡터를 생성할 수 있다(S2002).
예를 들어, 영상특징점강화장치(10)는 기술자를 K-means 클러스터링하여 단어의 가방(Bag of words) 개념의 시각적 단어(visual vocabulary)를 생성할 수 있다. 그리고 영상특징점강화장치(10)는 프레임 상의 특징점에 대응되는 기술자를 시각적 단어에 매칭할 수 있고, 매칭된 시각적 단어에 대응되는 값으로 시각적 단어 벡터를 생성하여 프레임의 특성을 저장할 수 있다.
그리고 영상특징점강화장치(10)는 프레임에서 특징점 영역에 위치한 물체를 인식하고, 인식된 물체에 대한 레이블인 의미정보를 시각적 단어 벡터에 추가하여 시각적 의미 단어 벡터를 생성할 수 있다(S2003).
예를 들어, 영상특징점강화장치(10)는 특징점을 기준으로 기 설정된 반경 내에 위치한 물체를 인식할 수 있고, 기 저장된 레이블 중 인식된 물체에 대응되는 레이블을 의미정보로 설정할 수 있다. 그리고 영상특징점강화장치(10)는 물체가 위치한 영역의 특징점에 대한 기술자의 마지막 값으로 물체 레이블을 포함시킴으로써 시각적 단어 벡터에 의미정보가 포함된 시각적 의미 단어 벡터를 생성할 수 있다.
이후, 영상특징점강화장치(10)는 현재 위치가 기 방문한 위치인지 여부를 식별하기 위해 시각적 의미 단어 벡터를 기초로 현재 프레임과 기 저장된 영상에 포함된 프레임의 유사도를 계산할 수 있으며, 이를 위해, 영상특징점강화장치(10)는 현재 프레임과의 유사도를 계산할 후보 키 프레임을 결정할 수 있다(S2004).
도 3 은 후보 키 프레임을 결정하는 순서를 도시한 순서도이다. 도 3 을 참조하면, 영상특징점강화장치(10)는 현재 프레임의 물체 레이블에 기초하여 기 저장된 영상의 키 프레임에서 제 1 후보프레임을 설정할 수 있다(S3001).
예를 들어, 영상특징점강화장치(10)는 기 저장된 영상의 키 프레임 중 현재 프레임에 포함된 물체 레이블과 동일한 적어도 하나의 물체 레이블을 포함하고 있는 프레임을 제 1 후보프레임으로 설정할 수 있다.
그리고 영상특징점강화장치(10)는 제 1 후보프레임 중 현재 프레임의 물체 레이블과 동일한 물체 레이블을 포함하는 수가 기 설정된 수를 초과하는 프레임을 제 2 후보프레임으로 설정할 수 있다(S3002).
예를 들어, 영상특징점강화장치(10)는 현재 프레임의 물체 레이블과 동일한 물체 레이블의 포함 개수를 기준으로 제 1 후보프레임에서 상위 20% 내에 속하는 제 1 후보프레임을 제 2 후보프레임으로 설정할 수 있다.
이후, 영상특징점강화장치(10)는 제 2 후보프레임의 이웃 프레임 중 현재 프레임 각각의 시각적 단어 벡터의 크기를 기초로 현재 프레임과 유사한 제 2 후보프레임의 이웃 프레임을 제 3 후보프레임으로 설정할 수 있다(S3003).
예를 들어, 영상특징점강화장치(10)는 제 2 후보프레임의 이웃프레임과 현재 프레임간 L1_Norm 값인 벡터크기가 가장 큰 이웃 프레임을 제 3 후보프레임으로 설정할 수 있다.
그리고 영상특징점강화장치(10)는 제 3 후보프레임으로부터 현재 프레임과 시각적 의미 단어 벡터의 크기를 이용하여 최종적으로 후보 키 프레임을 설정할 수 있다(S3004).
예를 들어, 영상특징점강화장치(10)는 현재 프레임과 이웃 프레임의 L1_Norm 값을 기초로 제 3 후보프레임을 정렬할 수 있고, 현재 프레임과의 시각적 의미 단어 벡터의 크기가 기 설정된 값을 초과하는 제 3 후보프레임을 후보 키 프레임을 선택할 수 있다.
이후, 영상특징점강화장치(10)는 결정된 후보 키 프레임과 현재 위치의 프레임 각각의 시각적 의미 단어 벡터를 구성하는 값의 차이를 기초로 유사도를 계산하여 현재 위치가 기 방문한 지점인지 식별할 수 있다(S2005).
예를 들어, 영상특징점강화장치(10)는 현재 프레임의 시각적 의미 단어 벡터와 후보 키 프레임의 시각적 의미 단어 벡터간의 해밍거리(Hamming distance)를 계산하여 현재 프레임과 후보 키 프레임간의 유사도를 계산할 수 있다.
그리고 현재 위치가 상기 기 방문한 지점으로 식별되면, 영상특징점강화장치(10)는 영상을 기초로 생성된 지도를 갱신할 수 있다(S2006).
예를 들어, 영상특징점강화장치(10)는 촬영된 영상을 기초로 지도를 생성할 수 있고, S2005단계에서 현재 위치가 기 방문한 위치인지 여부에 따라 회귀점을 검출할 수 있고, 회귀점을 폐쇄하여 지도를 갱신할 수 있다.
이상의 실시예들에서 사용되는 '~부'라는 용어는 소프트웨어 또는 FPGA(field programmable gate array) 또는 ASIC 와 같은 하드웨어 구성요소를 의미하며, '~부'는 어떤 역할들을 수행한다. 그렇지만 '~부'는 소프트웨어 또는 하드웨어에 한정되는 의미는 아니다. '~부'는 어드레싱할 수 있는 저장 매체에 있도록 구성될 수도 있고 하나 또는 그 이상의 프로세서들을 재생시키도록 구성될 수도 있다. 따라서, 일 예로서 '~부'는 소프트웨어 구성요소들, 객체지향 소프트웨어 구성요소들, 클래스 구성요소들 및 태스크 구성요소들과 같은 구성요소들과, 프로세스들, 함수들, 속성들, 프로시저들, 서브루틴들, 프로그램특허 코드의 세그먼트들, 드라이버들, 펌웨어, 마이크로코드, 회로, 데이터, 데이터베이스, 데이터 구조들, 테이블들, 어레이들, 및 변수들을 포함한다.
구성요소들과 '~부'들 안에서 제공되는 기능은 더 작은 수의 구성요소들 및 '~부'들로 결합되거나 추가적인 구성요소들과 '~부'들로부터 분리될 수 있다.
뿐만 아니라, 구성요소들 및 '~부'들은 디바이스 또는 보안 멀티미디어카드 내의 하나 또는 그 이상의 CPU 들을 재생시키도록 구현될 수도 있다.
도 2 및 도 3 을 통해 설명된 실시예에 따른 영상특징점강화방법은 컴퓨터에 의해 실행 가능한 명령어 및 데이터를 저장하는, 컴퓨터로 판독 가능한 매체의 형태로도 구현될 수 있다. 이때, 명령어 및 데이터는 프로그램 코드의 형태로 저장될 수 있으며, 프로세서에 의해 실행되었을 때, 소정의 프로그램 모듈을 생성하여 소정의 동작을 수행할 수 있다. 또한, 컴퓨터로 판독 가능한 매체는 컴퓨터에 의해 액세스될 수 있는 임의의 가용 매체일 수 있고, 휘발성 및 비휘발성 매체, 분리형 및 비분리형 매체를 모두 포함한다. 또한, 컴퓨터로 판독 가능한 매체는 컴퓨터 기록 매체일 수 있는데, 컴퓨터 기록 매체는 컴퓨터 판독 가능 명령어, 데이터 구조, 프로그램 모듈 또는 기타 데이터와 같은 정보의 저장을 위한 임의의 방법 또는 기술로 구현된 휘발성 및 비휘발성, 분리형 및 비분리형 매체를 모두 포함할 수 있다.예를 들어, 컴퓨터 기록 매체는 HDD 및 SSD 등과 같은 마그네틱 저장 매체, CD, DVD 및 블루레이 디스크 등과 같은 광학적 기록 매체, 또는 네트워크를 통해 접근 가능한 서버에 포함되는 메모리일 수 있다.
또한 도 2 및 도 3 을 통해 설명된 실시예에 따른 영상특징점강화방법은 컴퓨터에 의해 실행 가능한 명령어를 포함하는 컴퓨터 프로그램(또는 컴퓨터 프로그램 제품)으로 구현될 수도 있다. 컴퓨터 프로그램은 프로세서에 의해 처리되는 프로그래밍 가능한 기계 명령어를 포함하고, 고레벨 프로그래밍 언어(High-level Programming Language), 객체 지향 프로그래밍 언어(Object-oriented Programming Language), 어셈블리 언어 또는 기계 언어 등으로 구현될 수 있다. 또한 컴퓨터 프로그램은 유형의 컴퓨터 판독가능 기록매체(예를 들어, 메모리, 하드디스크, 자기/광학 매체 또는 SSD(Solid-State Drive) 등)에 기록될 수 있다.
따라서 도 2 및 도 3 을 통해 설명된 실시예에 따른 영상특징점강화방법은 상술한 바와 같은 컴퓨터 프로그램이 컴퓨팅 장치에 의해 실행됨으로써 구현될 수 있다. 컴퓨팅 장치는 프로세서와, 메모리와, 저장 장치와, 메모리 및 고속 확장포트에 접속하고 있는 고속 인터페이스와, 저속 버스와 저장 장치에 접속하고 있는 저속 인터페이스 중 적어도 일부를 포함할 수 있다. 이러한 성분들 각각은 다양한 버스를 이용하여 서로 접속되어 있으며, 공통 머더보드에 탑재되거나 다른 적절한 방식으로 장착될 수 있다.
여기서 프로세서는 컴퓨팅 장치 내에서 명령어를 처리할 수 있는데, 이런 명령어로는, 예컨대 고속 인터페이스에 접속된 디스플레이처럼 외부 입력, 출력 장치상에 GUI(Graphic User Interface)를 제공하기 위한 그래픽 정보를 표시하기 위해 메모리나 저장 장치에 저장된 명령어를 들 수 있다. 다른 실시예로서, 다수의 프로세서 및(또는) 다수의 버스가 적절히 다수의 메모리 및 메모리 형태와 함께 이용될 수 있다. 또한 프로세서는 독립적인 다수의 아날로그 및(또는) 디지털 프로세서를 포함하는 칩들이 이루는 칩셋으로 구현될 수 있다.
또한 메모리는 컴퓨팅 장치 내에서 정보를 저장한다. 일례로, 메모리는 휘발성 메모리 유닛 또는 그들의 집합으로 구성될 수 있다. 다른 예로, 메모리는 비휘발성 메모리 유닛 또는 그들의 집합으로 구성될 수 있다. 또한 메모리는 예컨대, 자기 혹은 광 디스크와 같이 다른 형태의 컴퓨터 판독 가능한 매체일 수도 있다.
그리고 저장장치는 컴퓨팅 장치에게 대용량의 저장공간을 제공할 수 있다. 저장 장치는 컴퓨터 판독 가능한 매체이거나 이런 매체를 포함하는 구성일 수 있으며, 예를 들어 SAN(Storage Area Network) 내의 장치들이나 다른 구성도 포함할 수 있고, 플로피 디스크 장치, 하드 디스크 장치, 광 디스크 장치, 혹은 테이프 장치, 플래시 메모리, 그와 유사한 다른 반도체 메모리 장치 혹은 장치 어레이일 수 있다.
상술된 실시예들은 예시를 위한 것이며, 상술된 실시예들이 속하는 기술분야의 통상의 지식을 가진 자는 상술된 실시예들이 갖는 기술적 사상이나 필수적인 특징을 변경하지 않고서 다른 구체적인 형태로 쉽게 변형이 가능하다는 것을 이해할 수 있을 것이다. 그러므로 상술된 실시예들은 모든 면에서 예시적인 것이며 한정적이 아닌 것으로 이해해야만 한다. 예를 들어, 단일형으로 설명되어 있는 각 구성 요소는 분산되어 실시될 수도 있으며, 마찬가지로 분산된 것으로 설명되어 있는 구성 요소들도 결합된 형태로 실시될 수 있다.
본 명세서를 통해 보호 받고자 하는 범위는 상기 상세한 설명보다는 후술하는 특허청구범위에 의하여 나타내어지며, 특허청구범위의 의미 및 범위 그리고 그 균등 개념으로부터 도출되는 모든 변경 또는 변형된 형태를 포함하는 것으로 해석되어야 한다.

Claims (16)

  1. SLAM에서 영상의 특징점을 강화하는 장치에 있어서,
    카메라를 통해 획득된 영상을 구성하는 프레임으로부터 특징점을 추출하고, 추출된 특징점의 특성에 대한 정보를 포함하는 기술자(descriptor)의 분포를 기록한 시각적 단어 벡터를 생성하는 시각적특징처리부;
    상기 프레임으로부터 상기 특징점 영역에 위치한 물체를 인식하고, 인식된 물체에 대한 의미정보를 추출하는 의미적특징처리부; 및
    상기 시각적 단어 벡터에 상기 의미정보가 추가된 시각적 의미 단어 벡터를 기초로 현재 위치가 기 방문한 지점인지 식별하는 위치추정부를 포함하는, 영상특징점강화장치.
  2. 제 1 항에 있어서,
    상기 의미적특징처리부는,
    기 저장된 레이블에서 상기 인식된 물체에 대응되는 레이블을 의미정보로써 설정하고, 상기 프레임 내의 물체 레이블의 분포를 상기 시각적 단어 벡터에 추가하여 상기 시각적 의미 단어 벡터를 생성하는, 영상특징점강화장치.
  3. 제 1 항에 있어서,
    상기 위치추정부는,
    기 저장된 영상의 키 프레임과 현재 위치에서의 프레임 간 유사도를 계산하는, 영상특징점강화장치.
  4. 제 3 항에 있어서,
    상기 위치추정부는,
    상기 현재 위치의 프레임에 포함된 물체의 레이블을 이용하여 상기 기 저장된 영상의 키 프레임 중 유사도를 계산할 후보프레임을 추출하는, 영상특징점강화장치.
  5. 제 4 항에 있어서,
    상기 위치추정부는,
    추출된 후보프레임 중 상기 현재 위치의 프레임의 시각적 의미 단어 벡터와의 유사도에 기초하여 후보 키 프레임을 결정하는, 영상특징점강화장치.
  6. 제 5 항에 있어서,
    상기 위치추정부는,
    결정된 후보 키 프레임과 상기 현재 위치의 프레임 각각의 시각적 의미 단어 벡터를 구성하는 값의 차이를 기초로 유사도를 계산하여 현재 위치가 기 방문한 지점인지 식별하는, 영상특징점강화장치.
  7. 제 6 항에 있어서,
    상기 영상특징점강화장치는,
    상기 현재 위치가 상기 기 방문한 지점으로 식별되면, 상기 영상을 기초로 생성된 지도를 갱신하는 지도갱신부를 더 포함하는, 영상특징점강화장치.
  8. 영상특징점강화장치가 SLAM에서 영상의 특징점을 강화하는 방법에 있어서,
    카메라를 통해 획득된 영상을 구성하는 프레임으로부터 특징점을 추출하고, 추출된 특징점의 특성에 대한 정보를 포함하는 기술자의 분포를 기록한 시각적 단어 벡터를 생성하는 단계;
    상기 프레임으로부터 상기 특징점 영역에 위치한 물체를 인식하고, 인식된 물체에 대한 의미정보를 추출하는 단계; 및
    상기 시각적 단어 벡터에 상기 의미정보가 추가된 시각적 의미 단어 벡터를 기초로 현재 위치가 기 방문한 지점인지 식별하는 단계를 포함하는, 영상특징점강화방법.
  9. 제 8 항에 있어서,
    상기 의미정보를 추출하는 단계는,
    기 저장된 레이블에서 상기 인식된 물체에 대응되는 레이블을 의미정보로써 설정하는 단계; 및
    상기 프레임 내의 물체 레이블의 분포를 상기 시각적 단어 벡터에 추가하여 상기 시각적 의미 단어 벡터를 생성하는 단계를 포함하는, 영상특징점강화방법.
  10. 제 8 항에 있어서,
    상기 기 방문한 지점인지 식별하는 단계는,
    기 저장된 영상의 키 프레임과 현재 위치에서의 프레임 간 유사도를 계산하는 단계를 포함하는, 영상특징점강화방법.
  11. 제 10 항에 있어서,
    상기 영상특징점강화방법은,
    상기 현재 위치의 프레임에 포함된 물체의 레이블을 이용하여 상기 기 저장된 영상의 키 프레임 중 유사도를 계산할 후보프레임을 추출하는 단계를 더 포함하는, 영상특징점강화방법.
  12. 제 11 항에 있어서,
    상기 영상특징점강화방법은,
    추출된 후보프레임 중 상기 현재 위치의 프레임의 시각적 의미 단어 벡터와의 유사도에 기초하여 후보 키 프레임을 결정하는 단계를 더 포함하는, 영상특징점강화방법.
  13. 제 12 항에 있어서,
    상기 영상특징점강화방법은,
    결정된 후보 키 프레임과 상기 현재 위치의 프레임 각각의 시각적 의미 단어 벡터를 구성하는 값의 차이를 기초로 유사도를 계산하여 현재 위치가 기 방문한 지점인지 식별하는 단계를 더 포함하는, 영상특징점강화방법.
  14. 제 13 항에 있어서,
    상기 영상특징점강화방법은,
    상기 현재 위치가 상기 기 방문한 지점으로 식별되면, 상기 영상을 기초로 생성된 지도를 갱신하는 단계를 더 포함하는, 영상특징점강화방법.
  15. 제 8 항에 기재된 방법을 수행하는 프로그램이 기록된 컴퓨터 판독 가능한 기록 매체.
  16. 영상특징점강화장치에 의해 수행되며, 제 8 항에 기재된 방법을 수행하기 위해 매체에 저장된 컴퓨터 프로그램.
PCT/KR2019/016641 2018-11-28 2019-11-28 물체 레이블을 활용한 비주얼 슬램에서의 영상 특징점 강화 방법 및 장치 Ceased WO2020111844A2 (ko)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
KR20180149674 2018-11-28
KR10-2018-0149674 2018-11-28
KR1020190039736A KR102285427B1 (ko) 2018-11-28 2019-04-04 물체 레이블을 활용한 Visual SLAM에서의 영상 특징점 강화 방법 및 장치
KR10-2019-0039736 2019-04-04

Publications (2)

Publication Number Publication Date
WO2020111844A2 true WO2020111844A2 (ko) 2020-06-04
WO2020111844A3 WO2020111844A3 (ko) 2020-10-15

Family

ID=70853388

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2019/016641 Ceased WO2020111844A2 (ko) 2018-11-28 2019-11-28 물체 레이블을 활용한 비주얼 슬램에서의 영상 특징점 강화 방법 및 장치

Country Status (1)

Country Link
WO (1) WO2020111844A2 (ko)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112016612A (zh) * 2020-08-26 2020-12-01 四川阿泰因机器人智能装备有限公司 一种基于单目深度估计的多传感器融合slam方法
CN113343982A (zh) * 2021-06-16 2021-09-03 北京百度网讯科技有限公司 多模态特征融合的实体关系提取方法、装置和设备

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR100834577B1 (ko) * 2006-12-07 2008-06-02 한국전자통신연구원 스테레오 비전 처리를 통해 목표물 검색 및 추종 방법, 및이를 적용한 가정용 지능형 서비스 로봇 장치
WO2015194864A1 (ko) * 2014-06-17 2015-12-23 (주)유진로봇 이동 로봇의 맵을 업데이트하기 위한 장치 및 그 방법
KR101976424B1 (ko) * 2017-01-25 2019-05-09 엘지전자 주식회사 이동 로봇

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112016612A (zh) * 2020-08-26 2020-12-01 四川阿泰因机器人智能装备有限公司 一种基于单目深度估计的多传感器融合slam方法
CN113343982A (zh) * 2021-06-16 2021-09-03 北京百度网讯科技有限公司 多模态特征融合的实体关系提取方法、装置和设备
CN113343982B (zh) * 2021-06-16 2023-07-25 北京百度网讯科技有限公司 多模态特征融合的实体关系提取方法、装置和设备

Also Published As

Publication number Publication date
WO2020111844A3 (ko) 2020-10-15

Similar Documents

Publication Publication Date Title
WO2022197136A1 (en) System and method for enhancing machine learning model for audio/video understanding using gated multi-level attention and temporal adversarial training
WO2022124701A1 (en) Method for producing labeled image from original image while preventing private information leakage of original image and server using the same
WO2018135881A1 (en) Vision intelligence management for electronic devices
WO2017213398A1 (en) Learning model for salient facial region detection
WO2017078361A1 (en) Electronic device and method for recognizing speech
WO2019031714A1 (ko) 객체를 인식하는 방법 및 장치
WO2016027983A1 (en) Method and electronic device for classifying contents
WO2019045244A1 (ko) 시각 대화를 통해 객체의 위치를 알아내기 위한 주의 기억 방법 및 시스템
WO2019050360A1 (en) ELECTRONIC DEVICE AND METHOD FOR AUTOMATICALLY SEGMENTING TO BE HUMAN IN AN IMAGE
WO2022139327A1 (en) Method and apparatus for detecting unsupported utterances in natural language understanding
WO2016013885A1 (en) Method for retrieving image and electronic device thereof
EP3752978A1 (en) Electronic apparatus, method for processing image and computer-readable recording medium
WO2021107592A1 (en) System and method for precise image inpainting to remove unwanted content from digital images
WO2016006728A1 (ko) 화상을 이용해 3차원 정보를 처리하는 전자 장치 및 방법
WO2019004754A1 (en) ADVERTISEMENTS WITH INCREASED REALITY ON OBJECTS
WO2021221490A1 (en) System and method for robust image-query understanding based on contextual features
WO2020256517A2 (ko) 전방위 화상정보 기반의 자동위상 매핑 처리 방법 및 그 시스템
EP3997625A1 (en) Electronic apparatus and method for controlling thereof
EP3459009A2 (en) Adaptive quantization method for iris image encoding
KR102285427B1 (ko) 물체 레이블을 활용한 Visual SLAM에서의 영상 특징점 강화 방법 및 장치
WO2024043490A1 (en) Video see-through (vst) augmented reality (ar) device and operating method for the same
WO2021251652A1 (ko) 영상 분석 장치 및 방법
CN116524532B (zh) 动作识别方法、装置、存储介质以及电子设备
WO2023182794A1 (ko) 검사 성능을 유지하기 위한 메모리 기반 비전 검사 장치 및 그 방법
WO2016021829A1 (ko) 동작 인식 방법 및 동작 인식 장치

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19888280

Country of ref document: EP

Kind code of ref document: A2

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 19888280

Country of ref document: EP

Kind code of ref document: A2