APPARATUSES AND METHODS FOR POSITIONING A DEVICE
TECHNICAL FIELD
Generally, the following description relates to the field of positioning systems. More specifically, the following description relates to apparatuses and methods for positioning a device.
BACKGROUND
Modern cars, mobile phones and many other devices are often equipped with a positioning system. The positioning system is used to measure a location of the device. The location information may be used in many applications. For example, navigation and other map related applications require that the position of the device is known so that navigation instructions can be formed.
The most common way to derive the exact location information is to use satellite-based navigation systems. In satellite navigation, a device determines the current location by receiving signals transmitted along a line of sight from a plurality of satellites. The device, such as a mobile phone or a vehicle, has electronic receivers for receiving the signals that can be used in positioning of the device. Typically, a global navigation satellite system does not require Internet reception; however, in some circumstances it may be used in assisting the positioning process.
As mentioned above, positioning signals are received along a line of sight. Thus, positioning is possible only when there is a line of sight to at least four satellites. This limits positioning in many locations to outdoors only. Furthermore, sometimes buildings, mountains or other obstacles are preventing visibility to at least four satellites, and the positioning is not possible. This is particularly a problem underground, for example in parking garages and road tunnels, or inside buildings.
The visibility problem has been addressed by developing indoor positioning systems. When coarse positioning is sufficient, the mobile communication network may be used. Special purpose indoor positioning systems are capable of providing the exact location. However, they are normally based on geolocated beacons, which transmit signals using Bluetooth, Wi-Fi or other suitable radio technology. These beacons are expensive and require maintenance. Another option is to use a plurality of different sensors and maps, either separately or together. For example, a vehicle can use a map and odometry to support navigation when entering a tunnel; however, it typically drifts quite quickly. Furthermore, when a mobile phone or a tablet computer is used instead of an integrated navigation device, the odometry information is not accessible. Thus, there is a need for improved positioning systems.
SUMMARY
In the following disclosure, apparatuses and methods for determining a position of a device are disclosed. The positioning is based on acquiring images that have unique position identifying features, which are either natural or man built. These position identifying features are used in generating a multi-dimensional signature. The multi-dimensional signature is processed so that a compact signature with reduced dimensions is achieved. The position of the device can be retrieved using the compact signature.
In a first aspect a method for determining a position of an apparatus is disclosed. The method comprises acquiring an image; extracting a position-identifying feature from the acquired image; generating a multi-dimensional signature of the acquired image using the position-identifying feature; generating a compact signature by reducing dimensions of the multidimensional signature of the acquired image; and retrieving the position of the apparatus using the compact signature. It is beneficial to use position-identifying features from acquired images to determine the location of a device or a vehicle. This provides a possibility for accurate positioning also in locations where satellite-based positioning is not possible.
In an implementation of the first aspect, acquiring the image comprises acquiring the image in a road direction. It is beneficial to use images that are acquired in the road direction either
forward or backward. This facilitates accurately measuring the distance between images and reduces the motion blur.
In an implementation of the first aspect, the method further comprises acquiring a plurality of images, extracting a position-identifying feature from each of the acquired images; generating a multi-dimensional signature of each of the acquired images using the positionidentifying feature; generating a compact signature by reducing dimensions of each of the multidimensional signatures; and retrieving the position of the apparatus using a plurality of compact signatures. It is beneficial to use a plurality of images as it improves the accuracy of the method in case of occlusion in one of the images.
In an implementation of the first aspect, the method further comprises measuring a geographical distance between the first and the last image. It is beneficial to measure the distance between the first and the last image, as the distance information improves the accuracy of the method by determining the proper length of the multidimensional signature.
In an implementation of the first aspect, retrieving the position of the apparatus comprises: generating a request comprising the compact signature; matching the compact signature of the request with a database comprising a plurality of compact signatures, wherein each of the compact signatures is associated with a position; and determining the position of an apparatus based on positions of the compact signatures matching the at least one signature. It is beneficial to have a database in the device so that the positioning can be done without a network connection, as it may not be available when the positioning is needed.
In an implementation of the first aspect, matching comprises finding a sequence of compact signatures that is the most similar to the compact signatures of the request. It is beneficial to try to find a sequence that is the most similar to the compact signatures of the request. The compact signature of the request may also have features that are not positionidentifying, and the compact signature does not need to be an exact match.
In an implementation of the first aspect, retrieving the position comprises generating a request comprising the compact signature; transmitting the request to an external database suitable for matching the compact signature of the request with a database comprising a
plurality of compact signatures associated with the position of an apparatus based on positions of the compact signatures matching the at least one signature; and receiving the position of the apparatus as a response to the request. It is beneficial to use an external database that might be larger and also easier to maintain so that it is always updated.
In an implementation of the first aspect, the position identifying feature is extracted using a machine learning arrangement configured to find a relevant position-identifying feature. It is beneficial to use a machine learning arrangement in detecting position-identifying features, as the machine learning arrangement may be trained further. This improves positioning in all possible locations instead of only the trained position.
In a second aspect a computer program comprising computer program code is disclosed. The computer program code is configured to cause performing a method as described above when the computer program code is executed in a computing apparatus. It is beneficial to implement the positioning method as software using a circuitry for determining the position.
In a third aspect an apparatus is disclosed. The apparatus comprises a camera configured to acquire an image; and processing circuitry configured to: extract a position identifying feature from the acquired image; generate a multi-dimensional signature of the acquired image using the position identifying feature; generate a compact signature by reducing dimensions of the multidimensional signature of the acquired image; and retrieve the position of the apparatus using at least one compact signature including the compact signature. It is beneficial to use position-identifying features from acquired images to determine the location of a device or a vehicle. This provides a possibility for accurate positioning also in locations where satellite-based positioning is not possible.
In an implementation of the third aspect, the camera is configured to acquire the image in a road orientation. It is beneficial to use to use images that are acquired in the road direction either forward or backward. This facilitates accurately measuring the distance between images and reduces the motion blur.
In an implementation of the third aspect, the camera is configured to acquire a plurality of images; and the processing circuitry is configured to extract a position-identifying feature
from each of the acquired images; generate a multi-dimensional signature of each of the acquired images using the position-identifying feature; generate a compact signature by reducing dimensions of each of the multidimensional signatures; and retrieve the position of the apparatus using a plurality of compact signatures. It is beneficial to use a plurality of images as it improves the accuracy of the method in case of occlusion in one of the images.
In an implementation of the third aspect, the processing circuitry is further configured to: measure a geographical distance between the first and the last image. It is beneficial to measure the distance between the first and the last image, as the distance information improves the accuracy of the method by determining the proper length of the multidimensional signature.
In an implementation of the third aspect, when retrieving the position of the apparatus, the processing circuitry is further configured to: generate a request comprising the compact signature; match one or more compact signatures of the request with a database comprising a plurality of compact signatures, wherein each of the compact signatures is associated with a position; and determine the position of an apparatus based on positions of the compact signatures matching the at least one signature. It is beneficial to have a database in the device so that the positioning can be done without a network connection, as it may not be available when the positioning is needed.
In an implementation of the third aspect, when matching, the processing circuitry is further configured to find a sequence of compact signatures that is the most similar to the compact signatures of the request. It is beneficial to try to find a sequence that is the most similar to the compact signatures of the request. The compact signature of the request may also have features that are not position-identifying, and the compact signature does not need to be an exact match.
In an implementation of the third aspect, when retrieving the position of an apparatus, the processing circuitry is further configured to: generate a request comprising the compact signature; transmit the request to an external service suitable for matching one or more compact signatures of the request with a database comprising a plurality of compact signatures associated with the position of an apparatus based on positions of the compact signatures matching at least one compact signature; and receive the position of the
apparatus as a response to the request. It is beneficial to use an external database that might be larger and also easier to maintain so that it is always updated.
In an implementation of the third aspect, the processing circuitry is configured to extract the position identifying feature using a machine learning arrangement configured to find a relevant position identifying feature. It is beneficial to use a machine learning arrangement in detecting position-identifying features, as the machine learning arrangement may be trained further. This improves positioning in all possible locations instead of only the trained position.
In a fourth aspect a server for positioning is disclosed. The server comprises a database comprising a plurality of compact signatures associated with a position; and a processing circuitry configured to: receive a request comprising at least one compact signature; match at least one compact signature of the request with a database comprising a plurality of compact signatures, wherein each of the compact signatures is associated with a position; and determine the position of a requesting apparatus based on positions of the compact signatures matching the at least one compact signature. It is beneficial to use an external server receiving requests from a positioning device. The external server may have a database that is larger and also easier to maintain so that it is always updated.
In an implementation of the fourth aspect, matching comprises finding a sequence of compact signatures that is the most similar to the compact signatures of the request. It is beneficial to try to find a sequence that is the most similar to the compact signatures of the request. The compact signature of the request may also have features that are not positionidentifying, and the compact signature does not need to be an exact match.
The foregoing and other objects are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description and the figures.
The principles discussed in the present description can be implemented in hardware and/or software.
BRIEF DESCRIPTION OF THE DRAWINGS
Further example embodiments will be described with respect to the following figures, wherein:
Fig. 1 shows an example of a method for generating compact signatures;
Fig. 2 shows an example of a scheme for generating compact signatures;
Fig. 3 shows an example of a training process for machine learning involved in road signature generation;
Fig. 4 shows an example of a scheme for an acquisition vehicle; and
Fig. 5 shows an example of an arrangement for positioning a car.
In the figures, identical reference signs will be used for identical or at least functionally equivalent features.
DETAILED DESCRIPTION OF THE EMBODIMENTS
In the following description, reference is made to the accompanying drawings, which form part of the disclosure, and in which are shown, by way of illustration, specific aspects in which the present apparatuses and methods may be provided. It is understood that other aspects may be utilized and structural or logical changes may be made without departing from the scope of the claims. Thus, the following detailed description is not to be taken in a limiting sense.
For instance, it is understood that a disclosure in connection with a described method may also hold true for a corresponding device or system configured to perform the method and vice versa. For example, if a specific method step is described, a corresponding device may
include a unit to perform the described method step, even if such unit is not explicitly described or illustrated in the figures. Further, it is understood that the features of the various exemplary aspects described herein may be combined with each other, unless specifically noted otherwise.
Fig. 1 discloses an example of a method for retrieving a position of an apparatus, such as a mobile phone, vehicle or any similar apparatus comprising a camera and an access to a database comprising position associated compact signatures or other similar references that can be derived from an image. In the example of Fig. 1 , the method is initiated by acquiring an image, step 100. The image is acquired using a conventional camera, for example a camera unit of a mobile phone or a camera of a vehicle. The acquired image is then processed in order to extract position-identifying features from the image, step 110. When the features are extracted, it is possible that also non-position identifying features are extracted. It is possible to use additional analysis tools for removing them; however, the positioning is possible when acquired images include enough position identifying features. In some cases, even one position identifying feature may be enough. For example, some of the road signs are unique, and when they are extracted the position can be accurately determined, even if the image comprises additional features. The extracting step may include identifying the extracted features. For example, when the feature is identified to be a vehicle, it may be discarded as it is highly unlikely that a vehicle can be used for determining the position. From the extracted features, a multi-dimensional signature is generated, step 120. The multi-dimensional signature provides very accurate identifying information of the extracted feature. It is used as a source for generating a compact signature, step 130. Finally, the position of the device is retrieved using one or more generated compact signatures, step 140.
In the above, an example of a method is disclosed. The method may be implemented in one device comprising a processing circuitry. The processing circuitry comprises at least one processor for executing computer programs and at least one memory for storing computer programs and related data. The memory may be volatile or non-volatile depending on the task required. The processor is configured to execute computer programs that are configured to cause performing the method as explained. Instead of one device, two or more devices may be used as explained in the following disclosure.
The example method of Fig. 1 may be completely performed in the device. This implementation may be considered as an offline implementation. In an online
implementation, one or more steps are executed by a remote service, such as a cloud service or a positioning server. Step 100 of acquiring images is performed in the device that is being positioned. As modern cameras are of high resolution, it is typically not desired to send the acquired images over the network to a remote service. In order to reduce network traffic, position-identifying features are typically extracted in the requesting device. Furthermore, generating multi-dimensional signatures and also generating compact signatures is typically done at the requesting device. This is beneficial, as compact signatures are smaller in size, and thus the request transmitted to a cloud or positioning service does not require as much bandwidth as it would require with images or multidimensional signatures. The position can be retrieved from an internal database stored in a memory of the device, or it can be retrieved from an external service. Retrieving comprises generating an appropriate request and receiving a response. Furthermore, retrieving comprises processing the request and transmitting a response. In an offline implementation, these are internal steps and signals within the device. In an online implementation, these tasks are performed by the external service, and the device does not need any further information on them.
Fig. 2 discloses an example of a scheme for generating compact signatures. The scheme may be used in several different devices or apparatuses, for example in mobile phones or cars. The following disclosure uses a car as an example; however, the scheme may be applied in any device having a camera unit for acquiring images and an access to a database, which may be located in a device itself or in a network service, such as a cloud service or server.
In the scheme, firstly one or more images are acquired, step 200. In the example of a car, these images are acquired in a road orientation or direction. This orientation may be forward or backward. One image may be sufficient for coarse positioning. However, when a method according to the scheme is used, for example in a navigation application, a plurality of images increases the accuracy of the method. Furthermore, the distance traveled between the first and the last or between each consecutive image may be measured and used in positioning. The acquired images can be ordinary images acquired with a camera that is already integrated into the acquiring device. For example, modern cars and mobile phones typically have one or more cameras.
The acquired images are transmitted to an image encoder 210. In the image encoder, the image is processed for finding position identifying features in the image. A position identifying feature is a feature that can be used in positioning. Thus, the position identifying feature is an object, landmark or the like. The position identifying feature does not change or move rapidly over time. Examples of position identifying features include trees, rock formations, lakes, buildings, road signs, roads and streets as such and the like. Examples of other features include the sun, persons, animals and other cars. These features typically move and change their place, so they cannot be used in determining the position. As an outcome, the image encoder 210 provides a set of image features 220 that includes one or more position identifying features.
The image encoder comprises a convolutional neural network (CNN) followed by a global pooling layer. The CNN outputs k feature maps that are summarized to one value by a feature map by the global pooling layer. The obtained k dimensional vector is finally normalized.
In the example of Fig. 2, the outcome 220 from the image encoder 210 is fed to the signature generator 230. The position identifying features 220, together with the distance information, are then used in generating a signature. The road signature generation module uses as inputs a continuous sequence of images from time t to t + n and the distance traveled dt+n from the position of the camera at time t and the position of the camera at time t + n. The number of /vectors generated is controlled by the spatial resolution factor c/A of the signature generator and can be computed using the relation: I =
. The attention mechanism
inside the signature generator uses a temporal filter to combine the input image features into road signature vectors.
The signature generator 230 returns a set of I vectors 240 that is called a road signature. Each vector of the signature has a dimension k and two consecutive vectors are separated by a distance of c/A. The set of I vectors 240 is then fed to dimension reduction 250. In the example of Fig. 2, the dimension reduction is an unsupervised machine learning method that is configured to reduce the dimension of the road signature. Dimension reduction is applied to each road signature vector to reduce its size from k to m, with m « k. The outcome is a compact road signature 260, which is a set of I vectors, wherein each of the vectors has a dimension m, which is considerably less than k, and two consecutive vectors are separated by a distance of c/A.
In the above described processes of image encoding, signature generation and dimension reduction are differentiable operations. That means the signature generation process is a fully differentiable operation, and thus road signatures can be optimized end to end for localization purposes and all the road signature generator components may be optimized jointly during the training. The training process is described below.
The process as described above is an example of a process that may be used in the examples discussed below. Other similar processes providing an association between one or more acquired images and a location may be used.
In Fig. 3, an example of a training process for machine learning involved in the road signature generator is disclosed. When producing discriminative compact signatures for precise localization, the road signature generator can be trained with the following principles.
The example of Fig. 3 shows two separate pipelines involving road signature generation. The first pipeline involves receiving a reference sequence of images and distance 300 as a reference input, which is processed by a road signature generator 302. As a result, one or more reference compact signatures 304 are computed. The process of generating road signatures may be similar to the process described above in the example of Fig. 2. The second pipeline receives a target sequence of images and distance 310 as a target input, which is processed by a road signature generator 304. As a result, one or more target compact signatures 314 are computed.
The first pipeline uses a reference input of a continuous sequence of images from time t to t + n and the distance traveled dt+n from the position of the camera at time t and the position of the camera at time t + n. The second pipeline uses a target input of a continuous sequence of images from time f’to t’ + m and the distance traveled df+m from the position of the camera at time t’ and the position of the camera at time t ’ + m. The target sequence is recorded on the same road as the reference sequence at a different time t
t The target sequence is longer than the reference sequence: df+m > dt+n.
The road signature generators 302 and 312 both use a Siamese network sharing the same set of parameter weights to compute the reference and the target signatures from the image sequences. These Siamese networks return the compact signature of the reference and the target sequence of images. As the distance traveled in the target sequence is longer than the distance traveled in the reference sequence, the target signature is longer than the reference signature.
A road signature matcher 320 computes signal similarity 322 between the reference and the target signature using sliding-window functions, for example Zero-mean Normalized Cross Correlation. The road signature matcher 320 returns a similarity measurement 330 of various sizes depending on the signature length. High similarity results in a low value in the similarity measurement signal. The road signature matcher 320 creates a ground truth signal from the ground truth spatial and temporal alignment of the reference and the target sequence 360. It is computed by generating a narrow 1 D Gaussian function centered at the overlapping position of the reference and target sequence. Ground truth alignment of the reference and target sequence are obtained from GPS measurement or by content-based image retrieval. A training signal is obtained by computing the cross-entropy loss 350 between the ground truth signal and the similarity measurement obtained from the road signature matcher 320. Before computing the loss, the similarity measurement obtained from the road signature matcher 320 is converted to probability 340 using a Softmax function 335 on the opposed signal. A gradient of the loss is computed regarding the input sequences and back-propagated through the whole pipeline to correct the weight parameters of the road signature generator.
Fig. 4 discloses a method for an acquisition vehicle 400 that is used in collecting location associated data and building a database to be used in image-based positioning. The acquisition vehicle of the example of Fig. 4 comprises an acquisition subsystem including an inertial navigation system, a satellite-based 410 navigation system antenna and an odometry system. The inertial navigation system is coupled to a camera, oriented along the longitudinal axis of the vehicle. Additionally, a subsystem may be used to ensure the synchronization of the image capture process and the inertial navigation system output. The additional subsystem saves images in a dedicated memory. The acquisition subsystem acquires images precisely geolocated and geotagged along the traveled road.
From the recorded sequence of images 420 and the traveled distance, the road signature generator 430 computes a compact road signature. The road signature computation can be done on the embedded vehicle to reduce the inboard storage usage or on a remote server to reduce the computational load on the acquisition vehicle. The road signature generator 430 produces a compact road signature 440 precisely longitudinally localized on the road being traveled by the acquisition vehicle. The precisely localized compact road signature 440 is stored in a geographic database 450. The geographic database 450 is then used as a source of information, or the information stored can be packaged for later use.
In Fig. 5, an example of an arrangement for positioning a car 500 is shown. The user is in a vehicle, equipped with a device that combines a GNSS application and a camera module, with a geodatabase in a memory. When entering GNSS restricted areas, the GNSS application will trigger an online process which will determine the position of the car using an image-based position system. The car 500 acquires images 510 from its camera module as well as the traveled distance, or vehicle speed to derive the traveled distance, of the image sequence. The vehicle uses a road signature generator 512 to compute a compact road signature 514 from the online-captured image sequence. The compact road signature 514 computed online by the user device or vehicle is called a target road signature.
A geographic database 520 of precomputed compact road signatures may be stored or partially stored on the user vehicle or device, and it is queried when entering a GNSS restricted area. In another implementation, the geographical database is accessed using an internet connection of the vehicle.
The geographic database 520 is used to retrieve a reference georeferenced compact road signature. This is done from a last valid GNSS position received before entering the GNSS restricted area. The user device or vehicle retrieves from the embedded geographic database 520 the precomputed compact road signature 522 geolocated on the road being traveled by the user. This compact road signature 522, precisely longitudinally located on the road, is called a reference compact road signature.
The vehicle 500 proceeds to target and reference signatures matching using a road signature matcher 530. The road signature matcher 530 produces a similarity signal between the target and the reference signatures for similarity measurement 532. The pick
in the similarity signal represents the optimal alignment between the target and the reference signatures. Thus, it provides the relative position of the target signature to the reference signature that produces the highest similarity measurement.
Finally, the road signature georeferencing module 540 of the user device or vehicle uses the similarity signal from the road signature matcher 530 and the precise position of the reference compact road signature to compute the precise position 550 of the target compact road signature along the driven road and to derive the current position of the user vehicle. Finally, the device location may be shown in an application that normally relies on satellite positioning.
As explained above, the arrangements using positioning as described above may be implemented in hardware, such as a mobile telephone, tablet computer, computer, telecommunication network base station or any other network connected device, or as a method. The method may be implemented as a computer program. The computer program is then executed in a computing device.
The apparatus, such as apparatus for positioning, is configured to perform one of the methods described above. The apparatus comprises necessary hardware components. These may include at least one processor, at least one memory, at least one network connection, a bus and the like. Instead of dedicated hardware components, it is possible to share for example memories or processors with other components or access at a cloud service, centralized computing unit or other resource that can be used over a network connection.
The apparatus for positioning and the corresponding method have been described in conjunction with various embodiments herein. However, other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention from a study of the drawings, the disclosure, and the appended claims. In the claims, the word “comprising” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude a plurality. A single processor or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. A computer program may be
stored/distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.