EP4677570A1 - Device and method for controlling augmented reality - Google Patents
Device and method for controlling augmented realityInfo
- Publication number
- EP4677570A1 EP4677570A1 EP23712617.2A EP23712617A EP4677570A1 EP 4677570 A1 EP4677570 A1 EP 4677570A1 EP 23712617 A EP23712617 A EP 23712617A EP 4677570 A1 EP4677570 A1 EP 4677570A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- physical object
- image
- key features
- spatial representation
- virtual spatial
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/20—Scenes; Scene-specific elements in augmented reality scenes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/70—Determining position or orientation of objects or cameras
- G06T7/73—Determining position or orientation of objects or cameras using feature-based methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/60—Type of objects
- G06V20/64—Three-dimensional [3D] objects
- G06V20/647—Three-dimensional [3D] objects by matching two-dimensional images to three-dimensional objects
Definitions
- the disclosure relates generally to augmented reality (AR).
- AR augmented reality
- the disclosure relates to a device, method, AR headset, AR system and non-transitory computer- readable storage medium for controlling AR.
- Augmented reality relates to mixed reality or computer-mediated reality.
- AR enhances a user's ongoing perception of a real world.
- AR provides perceptually enriched experiences by supplementing an object or an element in a natural and real- world with virtual information.
- the supplemental information may be computer-generated perceptual information about the environment and objects thereof, like scores and comments over a live video of a sporting event.
- AR provides an interactive experience and immersive sensation by blending the digital world into the user's perception of the real world.
- AR’s characteristics may require augmentation techniques to be performed in real time, accurate three-dimensional recognition and manipulation of virtual and real objects.
- AR technologies may be complex, cumbersome, and/or require high performance requirements of devices in AR applications.
- Examples and different aspects of this disclosure may simplify determination of an orientation (e.g., an upright orientation) of a spatial representation of a physical object in the digital space, such as by use of automatic recognition of the physical object. Accordingly, as compared to existing heavy- computational techniques (e.g., requiring use of a pre-existing computer model), the computations of the examples and different aspects of this disclosure may be faster and/or more efficient, thereby reducing processing requirements of relevant AR devices. In addition, this may allow smaller and/or more lightweight AR head-sets to be employed, improving use and wearability.
- a device for controlling augmented reality comprising: a non-transitory computer memory operable to store computer-readable instructions; and a processor operable to read the computer-readable instructions, the processor when executing the computer-readable instructions is configured to: obtain an image of a physical object captured by an image capturer; identify, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtain a virtual spatial representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determine an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
- the first aspect of the disclosure may seek to employ a simplified way to obtain the orientation information of the physical object in the real world. By extracting several key features from the image and projecting those key features from the image onto the spatial representation of the physical object, the orientation may be determined faster and more efficiently.
- Such device may be used in AR applications for positioning supplemental digital information, components and objects overlayed and relative to the physical object.
- the processor when executing the computer-readable instructions is further configured to detect a pixel position in the image for each of the plurality of the key features.
- the plurality of key features are projected, using a projection matrix, onto the virtual spatial representation by aligning the image with the virtual spatial representation.
- the processor when executing the computer-readable instructions is further configured to overlay, in the virtual spatial representation, the image relative to the identified respective position of the plurality of key features of the physical object.
- the plurality of key features of the physical object and the respective position for each of the plurality of key features are identified by a remote server, and wherein the processor when executing the computer-readable instructions is further configured to: transmit the image of the physical object to the remote server; and receive the plurality of key features and the respective position from the remote server. By transferring this task from the device to a remote server, the device may be further off loaded.
- the processor when executing the computer-readable instructions is further configured to: project the rays through the image onto the virtual spatial representation of the physical object to detect a position intersection for each of the plurality of key features of the physical object between the image and the virtual spatial representation; and determine the orientation of the virtual spatial representation based on the detected position intersection.
- the processor when executing the computer-readable instructions is further configured to: obtain characteristic information of at least one of the plurality of key features of the physical object from the remote server; and determine the orientation of the virtual spatial representation based on the characteristic information and the projection of the respective positions of the plurality of key features.
- the processor when executing the computer-readable instructions is further configured to: obtain information indicating a position or the orientation of the physical object from a sensor mounted on the physical object; and determine the orientation of the virtual spatial representation based on the information received from the sensor and the projection of the respective positions of the plurality of key features. Then a more accurate or precise orientation determination may be obtained with the sensor data.
- the virtual spatial representation is a spatial mesh.
- a moveable device for tracking a physical object comprising: an image capturer, for capturing an image of a physical object; and the device according to the first aspect of the disclosure.
- the moveable device comprising: a non-transitory computer memory operable to store computer-readable instructions; and a processor operable to read the computer-readable instructions, the processor when executing the computer-readable instructions is configured to: obtain an image of a physical object captured by an image capturer; identify, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtain a virtual spatial representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determine an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
- the moveable device is an augmented reality headset.
- the moveable device further comprises, a transceiver, for transmitting the image of the physical object to the remote server and receiving the plurality of key features of the physical object and the respective position from the remote server.
- the physical object is a vehicle
- the plurality of key features are at least one of a headlight, logo, tire and/or chassis.
- the image of the physical object is a two-dimensional image captured by the image capturer when the moveable device is moving relative to the physical object.
- the image capturer captures a series of images of the physical object; and project the rays through at least one of the captured images to the virtual spatial representation of the physical object; wherein, the respective position for each of the plurality of key features is an average central position identified from the series of images.
- the transceiver further receives information indicating a position or the orientation of the physical object from a sensor mounted on the physical object; and the processor when executing the computer-readable instructions is configured to, determines the orientation of the virtual spatial representation based on the information received from the sensor and the projection of the respective positions of the plurality of key features.
- a method for controlling augmented reality comprising: obtaining, by a processor of an augmented reality device, an image of a physical object captured by an image capturer; identifying, by the processor, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtaining, by the processor, a virtual spatial representation of the physical object captured by a depth sensor; projecting, by the processor, rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determining, by the processor, an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
- an augmented reality system comprising: a remote server for identifying, by using a trained object recognition model, a plurality of key features of a physical object and a respective position for each of the plurality of key features from an image of the physical object; and an augmented reality headset, for transmitting the image of the physical object to the remote server, and for receiving the plurality of key features and the respective positions from the remote server, wherein, the augmented reality headset comprises: an image capturer, for capturing the image of the physical object; a transceiver, for transmitting the image to the remote server and receiving the plurality of key features and the respective positions from the remote server; a non-transitory computer memory operable to store computer-readable instructions; and a processor operable to read the computer- readable instructions, the processor when executing the computer-readable instructions is configured to: obtain a virtual spatial representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the pluralit
- the image of the physical object is captured by the image capturer when the augmented reality headset is moving relative to the physical object
- the remote server receives a series of images of the physical object from the augmented reality headset, and identifies, using the trained object recognition model, the plurality of key features of the physical object and the respective position for each of the plurality of key features from the received series of images, wherein the respective position is an average central position identified from the received series of images.
- a non-transitory computer-readable storage medium comprising instructions, which when executed by a processor device, cause the processor device to: obtain an image of a physical object captured by an image capturer; identify, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtain a virtual spatial representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determine an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
- a computer program product comprising program code for performing, when executed by the processor device, the method of the third aspect of the disclosure.
- FIG. 1 shows an exemplary AR application system according to an example of the disclosure.
- FIG. 2 shows an exemplary interaction and illustrating diagram of an AR device according to the exemplary AR application system shown in FIG. 1, according to an example of the disclosure.
- FIG. 3A shows an exemplary annotated image of a training data set according to an example of the disclosure.
- FIG. 3B shows exemplary different types of physical objects that may be trained according to an example of the disclosure.
- FIG. 3C shows an exemplary coordinate system of a physical object according to an example of the disclosure.
- FIG. 3D shows an exemplary virtual spatial representation of a physical object according to an example of the disclosure.
- FIG. 4A shows an exemplary image of a physical object and key features thereof according to an example of the disclosure.
- FIG. 4B shows another exemplary image of a physical object and key features thereof according to another example of the disclosure.
- FIG. 5 shows an exemplary projection of a key feature onto a virtual spatial representation of a physical object according to an example of the disclosure.
- FIG. 6A shows an exemplary projection of rays through an image as shown in FIG. 4A onto a virtual spatial representation according to an example of the disclosure.
- FIG. 6B shows an enlarged view of the exemplary projection of rays as shown in FIG. 6A according to an example of the disclosure.
- FIG. 7 shows an exemplary flow chart of an exemplary method for being implemented on the AR device as shown in FIG. 2 according to an example of the disclosure.
- FIG. 8A shows another exemplary AR application system with an illustrated user’s AR visual view according to an example of the disclosure.
- FIG. 8B shows an exemplary AR view displayed to the user according to an example of the disclosure.
- FIG. 9 is a schematic diagram of an exemplary computer system for implementing examples disclosed herein in relation to an AR device according to an example of the disclosure.
- technologies such as simultaneous localization and mapping (SLAM) and computer-aided design (CAD) modules may be utilized to build a map, localize a physical object, and/or carry out a series of AR tasks.
- SLAM simultaneous localization and mapping
- CAD computer-aided design
- a change within the user’s view of the real world is tracked (e.g., the latest position of a physical object in the real world).
- the user’s AR visual display or view may be adjusted accordingly, to reflect the change.
- CAD computer-aided design
- the use of AR may be impractical.
- the use of AR may be limited.
- an AR device may capture an image of a physical object with an image capturer.
- the image capturer may be, for example, an AR camera.
- a plurality of key features of the physical object from the image may be identified by using a trained object recognition model. This may be performed at the AR device or at a remote device. Such identification is possibly carried out without using or depending on the existing complicated computational models, such as CAD models.
- a respective position in the image, for each of the plurality of key features, may also be identified.
- the type of object may be identified, such as a particular maker and model of a vehicle. This may be done based on the key features and their locations.
- the AR device may also obtain a virtual spatial representation of the physical object, which may be captured by a depth sensor of the AR device.
- the AR device may then project rays through the image onto the virtual spatial representation thereby projecting the plurality of key features onto the virtual spatial representation. Then an orientation of the virtual spatial representation may be determined by the AR device based on the projection.
- a faster and more efficient way e.g., consuming less computing power
- it may be more efficient to only identify a few key parts of physical object and project rays through the image to project those key features onto the virtual spatial representation by the AR device.
- identifying key parts or features of physical object may be more efficient in terms of trained datasets if such key parts are the same or similar across an entire product family of the physical object.
- Training may last the whole design lifecycle of the physical object when the key parts are the same. The lifecycle may be as long as 2-4 years or 10 years (e.g., for passenger car or truck industry, respectively) or even longer, depending on the type of the physical object.
- the computations may be made lighter, thus faster, since some examples of the invention do not require the use of heavy-computational models, such as CAD/SLAM models. Furthermore, this may also potentially reduce the amount of trained data of the physical object required.
- the computational burden on the AR device may be alleviated and the identification time may be reduced. This may be especially true for an AR headset or head mounted see through AR glasses, since such AR devices have relatively low computation performance.
- the result e.g., the determined orientation of the physical object
- the result may be utilized by other AR technologies, such as CAD model and other CAD based technologies (e.g., PTC Vuforia engine). For example, the orientation can be used to secure or improve the quality of the tracking of a physical object.
- the key features may be noticeable parts of the physical object in the image.
- the noticeable parts may be distinguishable, identifiable, or unique parts specific to the physical object, or specific to a particular type of objects to which the physical object belonging. Therefore, such noticeable parts may be used to match and find the position of the physical object in the real world.
- aligning the physical world with the digital world displayed to the user in this simplified manner a faster and less computationally intensive process may be obtained. Accordingly, this may reflect changes in the physical world more fully and timely, thereby a more accurate AR depiction may be provided.
- additional information received from a sensor mounted on the physical object may be utilized to increase precision.
- the sensor may be an onboard gyroscope and/or an accelerometer sensor. As data on the position of the vehicle may be obtained directly from the vehicle mounted sensor, precision of the position of the vehicle may be improved.
- FIG. 1 shows an exemplary AR system according to an example of the disclosure.
- the aforementioned AR device and the physical object are illustrated as AR headset 10 with AR lens 10a and vehicle 20, respectively.
- the AR headset 10 and vehicle 20 may communicate with each other via a network 50.
- the network 50 may be any kind of applicable networks that provide appropriate network connections therebetween.
- such exemplary AR device may be a head mounted see through AR glass and the like.
- the AR headset 10 may capture an image of the vehicle 20 which is observed by the user. Then operations described in examples and aspects of this disclosure may be carried out by the AR headset 10, based on the image it captured.
- some operations may be performed by another device, such as a remote server 30.
- the AR headset 10 may communicate with the remote server 30 via the network 50, to transmit information, such as the image it captured, to the remote server and receive information from the remote server 30.
- the remote server 30 may process the image received from the AR headset 10 to identify the plurality of key features of the physical object by using a trained object recognition model.
- the remote server 30 may train the object recognition model, which may be maintained by the remote server 30 or loaded into the AR headset 10.
- the remote server 30 may access database 40 as needed to perform those relevant operations.
- FIG. 2 illustrates an exemplary diagram of the AR headset 10 in FIG. 1.
- the AR headset 10 may include a system bus 100 to communicatively connect the components included therein.
- a non-transitory computer memory 110 may store necessary data and information for operation of the AR headset 10.
- computer-readable instructions 112 may be stored in the non-transitory computer memory 110.
- Processor 120 may read the computer-readable instructions 112 stored in the non-transitory computer memory 110, for example, via the system bus 100.
- the processor 120 may be configured to perform various operations described herein in relation to the AR headset 10.
- an image capturer 130 may be provided for capturing an image or photo 202 of the vehicle 20.
- the image capturer 130 may have a built in AR lens 10a as shown in FIG. 1 for certain camera effects as needed.
- AR lens 10a may, such as simulate the size, position, and field of view of human eyes, and/or add AR elements to the digital world to be displayed by the headset 10.
- a transceiver 140 communicates with the remote server 30 when needed.
- transceiver 140 may transmit the image 202 captured by the AR headset 10 to the remote server 30.
- transceiver 140 may also receive information from the remote server 30, for example, a plurality of key features 204 in the image 202 which are identified by the remote server 30.
- More detailed information about the key features 204 may be stored in the non-transitory computer memory 110, a memory built in or communicatively connected with the remote server 30, or any other appropriate storage devices. Such detailed information, as will be described below, may indicate, for example, characteristic information of the key features 204, such as what is an identified key feature 204. The detailed information may be retrieved as needed, when, for example, it may assist with aligning the physical world with the digital world displayed by the AR headset 10, or it may facilitate to retrieve AR content to be added to the displayed digital world, etc.
- a depth sensor 150 may be used to conduct measurements during the process of creating a virtual spatial representation 151 of a physical object in the real world.
- the virtual spatial representation 151 represents what a user sees through the AR headset 10.
- the depth sensor 150 may acquire information relating to multiple-point distance across, for example, a wide field-of-view (FOV) of surroundings of the AR headset 10 and help to create the virtual spatial representation 151.
- FOV wide field-of-view
- the depth sensor 150 may update its measurements to capture all the changes in the scene.
- FOV wide field-of-view
- relevant devices alternatively or additionally to the ones shown in FIG. 2 and described herein may be employed according to the particular use.
- other components may also be provided therein.
- the AR headset 10 may also maintain or access a trained object recognition model 160, which may be stored in the non-transitory computer memory 110.
- the trained object recognition model 160 may comprise one or more training data sets 160a.
- a training set 160a may be a plurality of images (e.g., hundreds or thousands of images). Features of interest in the images may be annotated or labeled.
- An exemplary annotated image is shown in FIG. 3 A. The image is a street view. In this image, traffic related information is annotated and labeled, including people, traffic lights and vehicles.
- FIG. 3B shows exemplary types of physical objects that may be used as training in some examples. Each type of physical objects may be characterized by one or more key parts or features of the physical object.
- the key parts may be headlights, logo, or the like of a vehicle. Detected differences in size, shape, location, etc., may characterize a different key part.
- FIG. 3B five different types of vehicles, which may represent five different truck families including thousands of trucks, which may be used by the trained object recognition model 160. The training data may then cover the full range of trucks within the same family during the lifecycle of the product family, as long as the corresponding key feature(s) has not been changed.
- the exemplary different types of vehicles in FIG. 3B may be considered as specific examples of the vehicle 20.
- FIG. 3C shows an exemplary coordinate system of a physical object according to an example of the disclosure.
- a coordinate system may be used to position a physical object in the real world, and used to calculate a precise positioning of any digital information to be overlaid on the physical object.
- FIG. 3C shows a 3D physical object 301 and a coordinate system 302.
- the 3D physical object 301 is illustrated as a tractor.
- the coordinate system 302 is a 3D coordinate system XYZ, wherein the origin is (0,0,0).
- the trained object recognition model 160 may place the 3D physical object 301 relative to the origin of the coordinate system 302. In this way, relative positions of all the components of the 3D physical object 301 learned by the trained object recognition model 160.
- the trained object recognition model 160 may know where the headlights, the logo, and other key features are located in the coordinate system 302.
- the AR headset 10 may recognize key features and their respective positions on the physical object. Those positions may be correlated to the coordinate system 302 in the physical world (e.g., by knowing relative positions of the key features to the origin). The origin of the coordinate system 302 may be used to determine relative positioning of those key features and thereby positioning the physical object and optionally any digital information on the physical object. By the correlation, the AR headset 10 may display the physical object to the user and display relevant digital information in the proper location. The AR headset 10 may merge the physical world and digital world. For example, the AR headset 10 may display additional information regarding the physical object, as shown in Fig. 8B
- the features of interest may also be the key features 204 herein, such as the lights and logo in figures 4A and 4B.
- the object recognition model 160 may contain an algorithm for a processing of the stored training data sets 160a.
- the object recognition model 160 may include weights and relevant calculations for processing an image of the training data sets 160a, in order to determine if the image comprises a key feature 204 and where that key feature is located.
- the trained object recognition model 160 may support the AR headset 10 to identify the key features 204 from the image 202, and also their respective positions in the image 202.
- the trained object recognition model 160 may be trained in a plurality of manners.
- a first training data set 160a comprising images of the physical vehicle 20 may be obtained.
- the first training data set 160a may include, for example, images of the different vehicle models available from a particular manufacturer.
- Key features of the physical vehicle 20 in the first training data set 160a may be identified by processing the images of the physical vehicle 20 using an object detection algorithm.
- the first training data set 160a may then be reduced into a second training data set 160a based on the identified restricted number of key features of the physical vehicle 20.
- the trained object recognition model 160 may be obtained by training an object recognition model to recognize the key features of the physical vehicle 20 in real-time image streams based on the reduced second training data set 160a.
- the exemplary training may be done in advance, for example by the remote server 30 or other appropriate processing apparatus.
- the trained object recognition model 160 may also be trained during use. In another example, the trained object recognition model 160 may be updated periodically or as needed when stored in the AR headset 10.
- the trained object recognition model 160 may also be trained with genericized data, such as the genericized key features which will be described in detail in the following. Upon being trained with relevant genericized information for a plurality of similar objects, such as for multiple types of trucks, the trained object recognition model 160 may identify a truck captured by the image capturer without requiring the user to indicate or input the specific type of the truck within his/her FOV. With the genericized data training, computations by the trained object recognition model 160 and therefore the overall processing time by the AR headset 10 may faster, more efficient. Furthermore, this renders the application of the trained object recognition model 160 and accordingly the AR headset possibly to various scenarios to identify an object in the physical world without preexisting knowledge or identification of the specific type of the object in advance.
- a user will not be interrupted by being required to input details of the object that he/she is observing (e.g., the type of a truck) when looking through the AR headset.
- details of the object that he/she is observing e.g., the type of a truck
- it possible that a user may enjoy a real-time AR display of a truck in front of him/her, then when turning around and facing another truck (either the same or different type of truck) nearby, they may continue to enjoy an AR display of that truck without any interruptions.
- the image capturer 130 may be triggered to take an image of the vehicle 20, for example when a user facing the vehicle 20 puts the AR headset 10 on or when the user moves relative to the vehicle 20.
- the image capturer 130 may be an AR camera.
- the image capturer 130 captures a two-dimensional (2D) image 202 of a physical object.
- the image may take many other forms, such as views from multiple AR cameras, multi-dimensional data from a three-dimensional (3D) image capturer, etc.
- the image capturer 130 may optionally capture live images of vehicle 20 in a real-time manner.
- the depth sensor 150 mounted on the AR headset 10 may capture and create the virtual spatial representation 151 of vehicle 20.
- the virtual spatial representation 151 for example, may be a spatial mesh.
- FIG. 3D shows an exemplary virtual spatial representation of a physical object according to an example of the disclosure, with different views of the physical object.
- the depth sensor 150 creates a spatial mesh 151a of the physical environment including a vehicle.
- the AR headset 10 may identify a plurality of key features 204 together with corresponding positions (e.g., pixel positions) from the image 202 captured by the image capturer 130, with the trained object recognition model 160, without using a CAD model.
- a key feature may be a predefined part of the physical object.
- a key feature may be a noticeable design part of the physical object.
- a key feature may represent a unique part of a physical object or a particular type of physical objects, as compared to others.
- key features may be one or more of headlights, logo, chassis, tire, etc. FIGs.
- the noticeable design parts may be one or more of wind deflector 204g, tire 204h, chassis 204i, the door 204j on the driver side, the knob 204k on the door, the tanks 2041or the like. Therefore, the key features 204 may be selected from a plurality of those components.
- the key features 204 may be genericized features of the object. Such genericized features may be common features for the object or similar objects. For example, for vehicles being the objects that observed by a user, the genericized features may be any features that a vehicle commonly has, such as a front windshield, a logo, a right headlight, a left headlight, a front bumper, etc. The genericized features may also be features that shared by certain objects, including the inter relationships among those shared features. For example, for trucks that belong to the same series, they may have a logo, a right and left headlights which are located at the same or similar approximate locations of the trucks.
- the distance between the right and left headlight, the distance between the right headlight and the logo, and the distance between the left headlight and the logo may also be the same or similar for the same series of trucks.
- genericizing features may allow an identification of similar physical objects. For example, key features (e.g., headlights and logo) genericized from multiple vehicles may be used for identifying those different vehicles.
- some different vehicles may have the same genericized features if they belong to a same series, a same brand, or if they are considered as a same classification or type produced by same/different manufactures, or the like.
- the distance between the right and left headlights, the distance from the right headlight to the logo, and the distance from the left headlight to the logo are approximately the same or within a reasonable difference value.
- such genericized key features e.g., the right and left headlights and the logo
- an object recognition model may be trained with the same genericized physical features for a number of different series of trucks of the same brand, or different brands.
- the trained object recognition model 160 may be used for identifying multiple types of vehicles, for example, without preexisting knowledge of the specific vehicle type.
- the user may not be interrupted and required to input further information about the physical object, such as the specific type or model of a truck.
- Various key features genericized in a similar way may be used for identifying a plurality of different, but similar physical objects.
- the trained object recognition model 160 may be developed with appropriate training algorithm(s) with the help of advanced AR technologies, such as computer vision to extract information from an image.
- Computer vision may be employed for acquiring, processing, analyzing and understanding digital images like image 202, and as well as extracting high-dimensional data (e.g., 3D digital information) from the digital images taken from the real world. Then numerical or symbolic information may be produced accordingly.
- a transformation of visual images into descriptions of the physical object being observed in the real world is carried out by using models constructed under the aid of geometry, physics, statistics, leaning theories, etc.
- sub-domains of computer vision may be utilized as needed, for example object detection, object recognition, three-dimensional pose estimation, learning, indexing, motion estimation, three-dimensional scene modeling, image restoration, or the like.
- the object recognition model 160 may be trained in advance, and then stored on the AR headset 10. In this scenario, the trained object recognition model 160 on the AR headset 10 may be updated periodically, or as needed. In some examples wherein an off-loading of processing from the AR headset 10 is desired, the trained object recognition model 160 may be developed and maintained at the remote server 30. In this case, it may be stored in the database 40. Then when the AR headset 10 is in use (e.g., powered on, or triggered to work, etc.), a communicative connection between the remote server 30 and the AR headset 10 may be established via the transceiver 140. The remote server 30 may perform the operations to identify the key features 204 and corresponding positions from the image 202 received from the transceiver 140. Subsequently, the transceiver 140 receives the identified key features 204 together with the respective position of each of the key features 204 in the captured image 202.
- the AR headset 10 may project rays through image 202 onto the virtual spatial representation 151 of vehicle 20.
- the plurality of key features 204 are projected onto the virtual spatial representation 151. Examples of such projection are shown in FIGs. 5, 6A and 6B respectively.
- FIG. 5 an exemplary projection of a key feature onto a virtual spatial representation of a vehicle as shown in FIG. 3D is illustrated.
- the key feature herein is a right headlight 204b of a vehicle, and its location in one of its captured image 202c is indicated as 205.
- the AR headset 10 may project the image 202 (e.g., image 202a in FIGs. 6 A and 6B) onto the virtual spatial representation 151 (e.g., virtual spatial representation 151b in FIGs. 6A and 6B) by using a projection matrix.
- the projection matrix may compensate for visual difference between a visual scene of the physical world from the location of the camera and the AR view displayed to the user by the user looking through AR headset 10. In other words, the visual difference may be due to a location difference of an AR camera and the location of human eyes.
- the orientation of vehicle 20 in the virtual spatial representation 151 may be determined. In this manner, the obtained 2D image 202 is aligned to the virtual spatial representation 151.
- the orientation of vehicle 20 is determined without identifying the particular vehicle 20.
- the type of vehicle 20 may be identified during the process of identification of the key features, projection and alignment mentioned above, but not the specific model.
- more details of vehicle 20 may be identified during the process, such as the specific model, the plate number, etc., if such details are needed later on for such as retrieving relevant AR content or the like.
- the positions of those key features may be used by the AR headset 10 to determine for example an orientation of vehicle 20 in digital space.
- the respective location for each of those three or more key features in image 202 is identified as well.
- the AR headset 10 may determine an upright orientation of vehicle 20 by using those identified locations of the three or more key features. For example, to determine the upright orientation of vehicle 20, the AR headset 10 may triangulate its logo, left and right headlights.
- the AR headset 10 may position the digital image 202 in an AR visual display for the user.
- the digital image 202 overlays the AR visual display (e.g., the virtual spatial representation 151) relative to the identified positions of the key features 204.
- images of a physical object may be taken in a real-time manner resulting in a real-time image stream including a series of images. In other examples, it may be not a real-time image stream, but any series of images, for example, images captured close in time.
- a cluster of positions for a specific key feature may be obtained.
- the trained object recognition model 160 may consider positions extracted from different images as belonging to a cluster of a particular key feature of vehicle 20 if such as the difference(s) in relation to the location of the same particular key feature in those different images is/are within a certain reasonable range.
- the reasonable range may be predetermined or predefined, it may also be trained along with the training of the trained object recognition model 160.
- an average central position calculated based on the positions identified from the series of images and considered as within the same cluster of a particular key feature is taken as the position of this key feature.
- the projection of rays may be performed through a selected image, some images, or the series of images, onto the virtual spatial representation 151.
- the introduction of such redundancy may help improve the accuracy of identification of key features and relevant positions thereof. For example, this may help the AR device or the remote server to differentiate and discard extraneous data.
- data may be, for example, information that relating to other physical objects nearby the physical object that the user is observing.
- information about a headlight or other parts of another vehicle close to the vehicle being observed by the user may be captured by the image capturer 130 and thus present in the image 202.
- the identified key features are shown as logo and headlights.
- the key features may be other components of the vehicle as mentioned above.
- other noticeable components may be used to train an object recognition model.
- a communicative connection between the AR headset 10 and vehicle 20 may be established via the transceiver 140. Then the transceiver 140 may obtain information from a sensor mounted on the vehicle 20. For example, when a gyroscope or accelerometer sensor is mounted on vehicle 20, optionally, information on position or the orientation of vehicle 20 may be transmitted to the transceiver 140 from the sensor. Then such information may help the determination of the orientation of the virtual spatial representation 151. Such determination may be made based on the information received from the sensor and the projection of the features 204. By obtaining supplemental or additional information from the sensor, the accuracy of orientation determination may be improved.
- FIG. 7 shows an exemplary flow chart of an exemplary method according to an example of the disclosure.
- This illustrated method may be performed by an AR device, for example, by the AR headset 10 of FIG. 2.
- An AR device may be triggered to work at power-on, or by other triggering mechanisms such as unfolding the legs of an AR glass, or any feasible or possible ways.
- the AR headset 10 obtains an image of a physical object (e.g., vehicle 20 of FIG. 2) captured by an image capturer (e.g., the image capturer 130 of FIG. 2).
- the AR headset 10 identifies, using the trained object recognition model 160, a plurality of key features 204 of vehicle 20 in the image 202 and a respective position for each of the key features 204.
- the AR headset 10 obtains a virtual spatial representation 151 of vehicle 20 captured by the depth sensor 150.
- the AR headset 10 projects rays through the image 202 onto the virtual spatial representation 151 of the physical object (i.e., vehicle 20) such that the plurality of key features 204 are projected onto the virtual spatial representation 151.
- the AR headset 10 determines an orientation of the virtual spatial representation 151 based on the projection of the plurality of key features 204 onto the virtual spatial representation 151.
- FIG. 8A shows another exemplary AR system with a user’s AR visual view according to an example of the disclosure.
- the illustrated AR system is similar as the one shown in FIG. 1, with additional functions performed by the remote server 30.
- the remote server 30 performs some of the aforementioned operations conducted by the AR headset 10.
- the remote server 30 may perform computing vision training 30a for developing models as needed, such as the trained object recognition model 160.
- the remote server 30 may also have a database for records of physical objects 30b (such as physical vehicles) for training sets.
- the remote server 30 may maintain other information and functional models or the like as needed in an AR application context.
- the remote server 30 may also provide characteristic information of the identified key features 204.
- the remote server may indicate that the identified features 204 are a logo of a vehicle, a left headlight and a right headlight thereof, respectively.
- This information may assist in the determination of the orientation of the vehicle 20.
- a positional relationship among the three key features 204 may help determine the vehicle’s orientation in digital world.
- the positional relationship could be the distance between two headlights, the distance between one headlight and the logo, etc.
- FIG. 8A also shows an AR visual view displayed to the user.
- an orientation of the virtual spatial representation of a physical object i.e., a vehicle herein
- the vehicle is located in the digital world in alignment with where the vehicle actually is in the real world.
- the virtual spatial representation is ready for the AR content.
- relevant information to be overlaid on, or additions to the AR visual view may be located more efficiently. For example, when a key feature is a logo, information on this logo or brand may be selected to combine with the surrounding real world of the user.
- FIG. 8B shows an exemplary AR view displayed to the user, according to an example of the disclosure.
- a user wearing an AR headset 10 may walk around a vehicle.
- the AR headset 10 may identify the key features of the vehicle, as discussed above. Then by relating the respective positions of the key features to the coordinate system, the AR headset 10 may recognize that the user is on the left side of the vehicle and looking down at the wheelbase of the vehicle. Then the AR headset 10 may display a digital object 801 in the headset to the user.
- the digital object 801 may illustrate that the user is on the left of a 4 wheelbase, with additional details of nearby components presented in the appropriate location.
- relevant digital information may be overlaid on the current view of the physical object in the real world. Therefore, the virtual world may be seamlessly interwoven with the physical world, which reflects the latest change in user’s view of the physical world in the AR digital world displayed to the user, may be achieved such that the user may perceive an immersive aspect of the real physical world.
- FIG. 9 is a schematic diagram of an exemplary computer system for implementing examples disclosed herein in relation to an AR device according to an example of the disclosure.
- the schematic diagram of a computer system 1000 may be utilized for implementing examples disclosed herein. Therefore, the computer system 1000 may be incorporated into an existing AR device, like the AR headset 10 in FIG. 2, to fulfil the functions of processor 120 and non-transitory computer memory 110 thereof.
- the computer system 1000 is adapted to execute instructions from a computer-readable medium to perform these and/or any of the functions or processing described herein.
- the computer system 1000 may be connected (e.g., networked) to other machines in a LAN, an intranet, an extranet, or the Internet.
- control system may include a single control unit or a plurality of control units connected or otherwise communicatively coupled to each other, such that any performed function may be distributed between the control units as desired.
- control system may include a single control unit or a plurality of control units connected or otherwise communicatively coupled to each other, such that any performed function may be distributed between the control units as desired.
- such devices may communicate with each other or other devices by various system architectures, such as directly or via a Controller Area Network (CAN) bus, etc.
- CAN Controller Area Network
- the computer system 1000 may comprise at least one computing device or electronic device capable of including firmware, hardware, and/or executing software instructions to implement the functionality described herein.
- the computer system 1000 may include processing circuitry 1002 (e.g., processing circuitry including one or more processor devices or control units), a memory 1004, and a system bus 1006.
- the computer system 1000 may include at least one computing device having the processing circuitry 1002.
- the system bus 1006 provides an interface for system components including, but not limited to, the memory 1004 and the processing circuitry 1002.
- the processing circuitry 1002 may include any number of hardware components for conducting data or signal processing or for executing computer code stored in memory 1004.
- the processing circuitry 1002 may, for example, include a general- purpose processor, an application specific processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), a circuit containing processing components, a group of distributed processing components, a group of distributed computers configured for processing, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein.
- the processing circuitry 1002 may further include computer executable code that controls operation of the programmable device.
- the system bus 1006 may be any of several types of bus structures that may further interconnect to a memory bus (with or without a memory controller), a peripheral bus, and/or a local bus using any of a variety of bus architectures.
- the memory 1004 may be one or more devices for storing data and/or computer code for completing or facilitating methods described herein.
- the memory 1004 may include database components, object code components, script components, or other types of information structure for supporting the various activities herein. Any distributed or local memory device may be utilized with the systems and methods of this description.
- the memory 1004 may be communicab ly connected to the processing circuitry 1002 (e.g., via a circuit or any other wired, wireless, or network connection) and may include computer code for executing one or more processes described herein.
- the memory 1004 may include non-volatile memory 1008 (e.g., read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.), and volatile memory 1010 (e.g., random-access memory (RAM)), or any other medium which can be used to carry or store desired program code in the form of machineexecutable instructions or data structures and which can be accessed by a computer or other machine with processing circuitry 1002.
- a basic input/output system (BIOS) 1012 may be stored in the non-volatile memory 1008 and can include the basic routines that help to transfer information between elements within the computer system 1000.
- BIOS basic input/output system
- the computer system 1000 may further include or be coupled to a non-transitory computer- readable storage medium such as the storage device 1014, which may comprise, for example, an internal or external hard disk drive (HDD) (e.g., enhanced integrated drive electronics (EIDE) or serial advanced technology attachment (SATA)), HDD (e.g., EIDE or SATA) for storage, flash memory, or the like.
- HDD enhanced integrated drive electronics
- SATA serial advanced technology attachment
- the storage device 1014 and other drives associated with computer-readable media and computer-usable media may provide non-volatile storage of data, data structures, computer-executable instructions, and the like.
- Computer-code which is hard or soft coded may be provided in the form of one or more modules.
- the module(s) can be implemented as software and/or hard-coded in circuitry to implement the functionality described herein in whole or in part.
- the modules may be stored in the storage device 1014 and/or in the volatile memory 1010, which may include an operating system 1016 and/or one or more program modules 1018.
- All or a portion of the examples disclosed herein may be implemented as a computer program 1020 stored on a transitory or non- transitory computer-usable or computer-readable storage medium (e.g., single medium or multiple media), such as the storage device 1014, which includes complex programming instructions (e.g., complex computer- readable program code) to cause the processing circuitry 1002 to carry out actions described herein.
- the computer-readable program code of the computer program 1020 can comprise software instructions for implementing the functionality of the examples described herein when executed by the processing circuitry 1002.
- the storage device 1014 may be a computer program product (e.g., readable storage medium) storing the computer program 1020 thereon, where at least a portion of a computer program 1020 may be loadable (e.g., into a processor) for implementing the functionality of the examples described herein when executed by the processing circuitry 1002.
- the processing circuitry 1002 may serve as a controller, or control system, for the computer system 1000 that is to implement the functionality described herein.
- the computer system 1000 may include an input device interface 1022 configured to receive input and selections to be communicated to the computer system 1000 when executing instructions, such as from a keyboard, mouse, touch-sensitive surface, etc.
- Such input devices may be connected to the processing circuitry 1002 through the input device interface 1022 coupled to the system bus 1006 but can be connected through other interfaces such as a parallel port, an Institute of Electrical and Electronic Engineers (IEEE) 1394 serial port, a Universal Serial Bus (USB) port, an IR interface, and the like.
- the computer system 1000 may include an output device interface 1024 configured to forward output, such as to a display, a video display unit (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)).
- the computer system 1000 may include a communications interface 1026 suitable for communicating with a network as appropriate or desired.
- Example 1 A device for controlling augmented reality, comprising: a non- transitory computer memory operable to store computer-readable instructions; and a processor operable to read the computer-readable instructions, the processor when executing the computer-readable instructions is configured to: obtain an image of a physical object captured by an image capturer; identify, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtain a virtual spatial representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determine an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
- Example 2 The device of example 1, wherein the processor when executing the computer-readable instructions is further configured to detect a pixel position in the image for each of the plurality of the key features.
- Example 3 The device of example 1, wherein the plurality of key features are projected, using a projection matrix, onto the virtual spatial representation by aligning the image with the virtual spatial representation.
- Example 4 The device of example 1, wherein the processor when executing the computer-readable instructions is further configured to overlay, in the virtual spatial representation, the image relative to the identified respective position of the plurality of key features of the physical object.
- Example 5 The device of example 1, wherein the plurality of key features of the physical object and the respective position for each of the plurality of key features are identified by a remote server, and wherein the processor when executing the computer-readable instructions is further configured to: transmit the image of the physical object to the remote server; and receive the plurality of key features and the respective position from the remote server.
- Example 6 The device of example 1, wherein the processor when executing the computer- readable instructions is further configured to: project the rays through the image onto the virtual spatial representation of the physical object to detect a position intersection for each of the plurality of key features of the physical object between the image and the virtual spatial representation; and determine the orientation of the virtual spatial representation based on the detected position intersection.
- Example 7 The device of example 5, wherein the processor when executing the computer-readable instructions is further configured to: obtain characteristic information of at least one of the plurality of key features of the physical object from the remote server; and determine the orientation of the virtual spatial representation based on the characteristic information and the projection of the respective positions of the plurality of key features.
- Example 8 The device of example 1, wherein the processor when executing the computer- readable instructions is further configured to: obtain information indicating a position or the orientation of the physical object from a sensor mounted on the physical object; and determine the orientation of the virtual spatial representation based on the information received from the sensor and the projection of the respective positions of the plurality of key features.
- Example 9 A moveable device for tracking a physical object, comprising: an image capturer, for capturing an image of the physical object; and the device of example 1.
- Example 10 The moveable device of example 9, wherein the moveable device is an augmented reality headset.
- Example 12 The moveable device of example 9, wherein the physical object is a vehicle, and the plurality of key features are at least one of a headlight, logo, tire and/or chassis.
- Example 13 The moveable device of example 9, wherein the image of the physical object is a two-dimensional image captured by the image capturer when the moveable device is moving relative to the physical object.
- Example 14 The moveable device of example 10, wherein the augmented reality headset includes an AR lens.
- Example 15 The moveable device of example 9, wherein, the image capturer captures a series of images of the physical object; and project the rays through at least one of the captured images to the virtual spatial representation of the physical object; wherein, the respective position for each of the plurality of key features is an average central position identified from the series of images.
- Example 16 The moveable device of example 11, wherein: the transceiver further receives information indicating a position or the orientation of the physical object from a sensor mounted on the physical object; and the processor when executing the computer-readable instructions is configured to, determines the orientation of the virtual spatial representation based on the information received from the sensor and the projection of the respective positions of the plurality of key features.
- Example 17 The moveable device of example 16, wherein the sensor is a gyroscope and/or an accelerometer sensor.
- Example 18 A method for controlling augmented reality, comprising: obtaining, by a processor of an augmented reality device, an image of a physical object captured by an image capturer; identifying, by the processor, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtaining, by the processor, a virtual spatial representation of the physical object captured by a depth sensor; projecting, by the processor, rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determining, by the processor, an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
- Example 19 An augmented reality system, comprising: a remote server for identifying, by using a trained object recognition model, a plurality of key features of a physical object and a respective position for each of the plurality of key features from an image of the physical object; and an augmented reality headset, for transmitting the image of the physical object to the remote server, and for receiving the plurality of key features and the respective positions from the remote server, wherein, the augmented reality headset comprises: an image capturer, for capturing the image of the physical object; a transceiver, for transmitting the image to the remote server and receiving the plurality of key features and the respective positions from the remote server; a non-transitory computer memory operable to store computer-readable instructions; and a processor operable to read the computer-readable instructions, the processor when executing the computer-readable instructions is configured to: obtain a virtual spatial representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the
- Example 20 The augmented reality system of example 19, wherein the image of the physical object is captured by the image capturer when the augmented reality headset is moving relative to the physical object, and the remote server: receives a series of images of the physical object from the augmented reality headset, and identifies, using the trained object recognition model, the plurality of key features of the physical object and the respective position for each of the plurality of key features from the received series of images, wherein the respective position is an average central position identified from the received series of images.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Multimedia (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Image Analysis (AREA)
Abstract
A device and method for controlling augmented reality. The device may comprise a non-transitory computer memory and a processor. The processor is configured to: obtain an image of a physical object captured by an image capturer; identify, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtain a virtual spatial representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determine an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
Description
DEVICE AND METHOD FOR CONTROLLING AUGMENTED REALITY
TECHNICAL FIELD
[0001] The disclosure relates generally to augmented reality (AR). In particular aspects, the disclosure relates to a device, method, AR headset, AR system and non-transitory computer- readable storage medium for controlling AR.
BACKGROUND
[0002] Augmented reality relates to mixed reality or computer-mediated reality. Unlike virtual reality which completely replaces a user's real-world environment with a simulated virtual world, AR enhances a user's ongoing perception of a real world. AR provides perceptually enriched experiences by supplementing an object or an element in a natural and real- world with virtual information. The supplemental information may be computer-generated perceptual information about the environment and objects thereof, like scores and comments over a live video of a sporting event. Rather than simply displaying the virtual information, AR provides an interactive experience and immersive sensation by blending the digital world into the user's perception of the real world. AR’s characteristics may require augmentation techniques to be performed in real time, accurate three-dimensional recognition and manipulation of virtual and real objects.
SUMMARY
[0003] To leverage advantages of AR applications, knowledge of the physical world or any changes thereof may be obtained. However, AR technologies may be complex, cumbersome, and/or require high performance requirements of devices in AR applications. Examples and different aspects of this disclosure may simplify determination of an orientation (e.g., an upright
orientation) of a spatial representation of a physical object in the digital space, such as by use of automatic recognition of the physical object. Accordingly, as compared to existing heavy- computational techniques (e.g., requiring use of a pre-existing computer model), the computations of the examples and different aspects of this disclosure may be faster and/or more efficient, thereby reducing processing requirements of relevant AR devices. In addition, this may allow smaller and/or more lightweight AR head-sets to be employed, improving use and wearability.
[0004] According to a first aspect of the disclosure, a device for controlling augmented reality, comprising: a non-transitory computer memory operable to store computer-readable instructions; and a processor operable to read the computer-readable instructions, the processor when executing the computer-readable instructions is configured to: obtain an image of a physical object captured by an image capturer; identify, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtain a virtual spatial representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determine an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation. The first aspect of the disclosure may seek to employ a simplified way to obtain the orientation information of the physical object in the real world. By extracting several key features from the image and projecting those key features from the image onto the spatial representation of the physical object, the orientation may be determined faster and more efficiently. Such device may
be used in AR applications for positioning supplemental digital information, components and objects overlayed and relative to the physical object.
[0005] In some examples, including in at least one preferred example, optionally the processor when executing the computer-readable instructions is further configured to detect a pixel position in the image for each of the plurality of the key features.
[0006] In some examples, including in at least one preferred example, optionally the plurality of key features are projected, using a projection matrix, onto the virtual spatial representation by aligning the image with the virtual spatial representation.
[0007] In some examples, including in at least one preferred example, optionally the processor when executing the computer-readable instructions is further configured to overlay, in the virtual spatial representation, the image relative to the identified respective position of the plurality of key features of the physical object.
[0008] In some examples, including in at least one preferred example, optionally the plurality of key features of the physical object and the respective position for each of the plurality of key features are identified by a remote server, and wherein the processor when executing the computer-readable instructions is further configured to: transmit the image of the physical object to the remote server; and receive the plurality of key features and the respective position from the remote server. By transferring this task from the device to a remote server, the device may be further off loaded.
[0009] In some examples, including in at least one preferred example, optionally the processor when executing the computer-readable instructions is further configured to: project the rays through the image onto the virtual spatial representation of the physical object to detect a
position intersection for each of the plurality of key features of the physical object between the image and the virtual spatial representation; and determine the orientation of the virtual spatial representation based on the detected position intersection.
[0010] In some examples, including in at least one preferred example, optionally the processor when executing the computer-readable instructions is further configured to: obtain characteristic information of at least one of the plurality of key features of the physical object from the remote server; and determine the orientation of the virtual spatial representation based on the characteristic information and the projection of the respective positions of the plurality of key features.
[0011] In some examples, including in at least one preferred example, optionally the processor when executing the computer-readable instructions is further configured to: obtain information indicating a position or the orientation of the physical object from a sensor mounted on the physical object; and determine the orientation of the virtual spatial representation based on the information received from the sensor and the projection of the respective positions of the plurality of key features. Then a more accurate or precise orientation determination may be obtained with the sensor data.
[0012] In some examples, including in at least one preferred example, optionally the virtual spatial representation is a spatial mesh.
[0013] According to a second aspect of the disclosure, a moveable device for tracking a physical object, comprising: an image capturer, for capturing an image of a physical object; and the device according to the first aspect of the disclosure. Specifically, the moveable device comprising: a non-transitory computer memory operable to store computer-readable instructions; and a processor operable to read the computer-readable instructions, the processor when
executing the computer-readable instructions is configured to: obtain an image of a physical object captured by an image capturer; identify, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtain a virtual spatial representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determine an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
[0014] In some examples, including in at least one preferred example, optionally the moveable device is an augmented reality headset.
[0015] In some examples, including in at least one preferred example, optionally the augmented reality headset includes AR lens.
[0016] In some examples, including in at least one preferred example, optionally the plurality of key features of the physical object and the respective position are identified by a remote server, and the moveable device further comprises, a transceiver, for transmitting the image of the physical object to the remote server and receiving the plurality of key features of the physical object and the respective position from the remote server.
[0017] In some examples, including in at least one preferred example, optionally the physical object is a vehicle, and the plurality of key features are at least one of a headlight, logo, tire and/or chassis.
[0018] In some examples, including in at least one preferred example, optionally the image of the physical object is a two-dimensional image captured by the image capturer when the moveable device is moving relative to the physical object.
[0019] In some examples, including in at least one preferred example, optionally the processor when executing the computer-readable instructions is configured to: project the rays through the image onto the virtual spatial representation of the physical object to detect a position intersection for each of the plurality of key features of the physical object between the image and the virtual spatial representation; and determine the orientation of the virtual spatial representation based on the detected position intersection.
[0020] In some examples, including in at least one preferred example, optionally the image capturer captures a series of images of the physical object; and project the rays through at least one of the captured images to the virtual spatial representation of the physical object; wherein, the respective position for each of the plurality of key features is an average central position identified from the series of images.
[0021] In some examples, including in at least one preferred example, optionally the transceiver further receives information indicating a position or the orientation of the physical object from a sensor mounted on the physical object; and the processor when executing the computer-readable instructions is configured to, determines the orientation of the virtual spatial representation based on the information received from the sensor and the projection of the respective positions of the plurality of key features.
[0022] In some examples, including in at least one preferred example, optionally the sensor is a gyroscope and/or an accelerometer sensor.
[0023] According to a third aspect of the disclosure, a method for controlling augmented reality, comprising: obtaining, by a processor of an augmented reality device, an image of a physical object captured by an image capturer; identifying, by the processor, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtaining, by the processor, a virtual spatial representation of the physical object captured by a depth sensor; projecting, by the processor, rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determining, by the processor, an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
[0024] According to a fourth aspect of the disclosure, an augmented reality system, comprising: a remote server for identifying, by using a trained object recognition model, a plurality of key features of a physical object and a respective position for each of the plurality of key features from an image of the physical object; and an augmented reality headset, for transmitting the image of the physical object to the remote server, and for receiving the plurality of key features and the respective positions from the remote server, wherein, the augmented reality headset comprises: an image capturer, for capturing the image of the physical object; a transceiver, for transmitting the image to the remote server and receiving the plurality of key features and the respective positions from the remote server; a non-transitory computer memory operable to store computer-readable instructions; and a processor operable to read the computer- readable instructions, the processor when executing the computer-readable instructions is configured to: obtain a virtual spatial representation of the physical object captured by a depth
sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determine an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
[0025] In some examples, including in at least one preferred example, optionally the image of the physical object is captured by the image capturer when the augmented reality headset is moving relative to the physical object, and the remote server: receives a series of images of the physical object from the augmented reality headset, and identifies, using the trained object recognition model, the plurality of key features of the physical object and the respective position for each of the plurality of key features from the received series of images, wherein the respective position is an average central position identified from the received series of images.
[0026] According to a fifth aspect of the disclosure, a non-transitory computer-readable storage medium comprising instructions, which when executed by a processor device, cause the processor device to: obtain an image of a physical object captured by an image capturer; identify, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtain a virtual spatial representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determine an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
[0027] According to a sixth aspect of the disclosure, a computer program product comprising program code for performing, when executed by the processor device, the method of the third aspect of the disclosure.
[0028] The disclosed aspects, examples (including any preferred examples), and/or accompanying claims may be suitably combined with each other as would be apparent to anyone of ordinary skill in the art.
[0029] Additional features and advantages are disclosed in the following description, claims, and drawings, and in part will be readily apparent therefrom to those skilled in the art or recognized by practicing the disclosure as described herein.
[0030] There are also disclosed herein computer systems, control units, code modules, computer-implemented methods, computer readable media, and computer program products associated with the above discussed technical benefits.
BRIEF DESCRIPTION OF DRAWINGS
[0031] Examples are described in more detail below with reference to the appended drawings.
[0032] FIG. 1 shows an exemplary AR application system according to an example of the disclosure.
[0033] FIG. 2 shows an exemplary interaction and illustrating diagram of an AR device according to the exemplary AR application system shown in FIG. 1, according to an example of the disclosure.
[0034] FIG. 3A shows an exemplary annotated image of a training data set according to an example of the disclosure.
[0035] FIG. 3B shows exemplary different types of physical objects that may be trained according to an example of the disclosure.
[0036] FIG. 3C shows an exemplary coordinate system of a physical object according to an example of the disclosure.
[0037] FIG. 3D shows an exemplary virtual spatial representation of a physical object according to an example of the disclosure.
[0038] FIG. 4A shows an exemplary image of a physical object and key features thereof according to an example of the disclosure.
[0039] FIG. 4B shows another exemplary image of a physical object and key features thereof according to another example of the disclosure.
[0040] FIG. 5 shows an exemplary projection of a key feature onto a virtual spatial representation of a physical object according to an example of the disclosure.
[0041] FIG. 6A shows an exemplary projection of rays through an image as shown in FIG. 4A onto a virtual spatial representation according to an example of the disclosure.
[0042] FIG. 6B shows an enlarged view of the exemplary projection of rays as shown in FIG. 6A according to an example of the disclosure.
[0043] FIG. 7 shows an exemplary flow chart of an exemplary method for being implemented on the AR device as shown in FIG. 2 according to an example of the disclosure.
[0044] FIG. 8A shows another exemplary AR application system with an illustrated user’s AR visual view according to an example of the disclosure.
[0045] FIG. 8B shows an exemplary AR view displayed to the user according to an example of the disclosure.
[0046] FIG. 9 is a schematic diagram of an exemplary computer system for implementing examples disclosed herein in relation to an AR device according to an example of the disclosure.
DETAILED DESCRIPTION
[0047] The detailed description set forth below provides information and examples of the disclosed technology with sufficient detail to enable those skilled in the art to practice the disclosure.
[0048] In the field of AR applications, various technologies may be used. For example, technologies such as simultaneous localization and mapping (SLAM) and computer-aided design (CAD) modules may be utilized to build a map, localize a physical object, and/or carry out a series of AR tasks. In this way, a change within the user’s view of the real world is tracked (e.g., the latest position of a physical object in the real world). Then the user’s AR visual display or view may be adjusted accordingly, to reflect the change. However, in a CAD based scenario for example, wherein not all exact digital representation of each and every physical object is available to train CAD based technology, the use of AR may be impractical. In addition, when other complicated and heavy computational AR technologies are employed, the use of AR may be limited.
[0049] According to the examples and different aspects of the disclosure, an AR device may capture an image of a physical object with an image capturer. The image capturer may be, for example, an AR camera. A plurality of key features of the physical object from the image
may be identified by using a trained object recognition model. This may be performed at the AR device or at a remote device. Such identification is possibly carried out without using or depending on the existing complicated computational models, such as CAD models. A respective position in the image, for each of the plurality of key features, may also be identified. In addition, the type of object may be identified, such as a particular maker and model of a vehicle. This may be done based on the key features and their locations. The AR device may also obtain a virtual spatial representation of the physical object, which may be captured by a depth sensor of the AR device. The AR device may then project rays through the image onto the virtual spatial representation thereby projecting the plurality of key features onto the virtual spatial representation. Then an orientation of the virtual spatial representation may be determined by the AR device based on the projection.
[0050] Therefore, a faster and more efficient way (e.g., consuming less computing power) to orient the physical object in a digital space may be achieved. It may be more efficient to only identify a few key parts of physical object and project rays through the image to project those key features onto the virtual spatial representation by the AR device. For example, identifying key parts or features of physical object may be more efficient in terms of trained datasets if such key parts are the same or similar across an entire product family of the physical object. Training may last the whole design lifecycle of the physical object when the key parts are the same. The lifecycle may be as long as 2-4 years or 10 years (e.g., for passenger car or truck industry, respectively) or even longer, depending on the type of the physical object. Moreover, the computations may be made lighter, thus faster, since some examples of the invention do not require the use of heavy-computational models, such as CAD/SLAM models. Furthermore, this may also potentially reduce the amount of trained data of the physical object required. In
addition, by reducing the need for complex computations (e.g., world map), the computational burden on the AR device may be alleviated and the identification time may be reduced. This may be especially true for an AR headset or head mounted see through AR glasses, since such AR devices have relatively low computation performance. Moreover, the result (e.g., the determined orientation of the physical object) may be utilized by other AR technologies, such as CAD model and other CAD based technologies (e.g., PTC Vuforia engine). For example, the orientation can be used to secure or improve the quality of the tracking of a physical object.
[0051] In some examples, the key features may be noticeable parts of the physical object in the image. The noticeable parts may be distinguishable, identifiable, or unique parts specific to the physical object, or specific to a particular type of objects to which the physical object belonging. Therefore, such noticeable parts may be used to match and find the position of the physical object in the real world. By aligning the physical world with the digital world displayed to the user in this simplified manner, a faster and less computationally intensive process may be obtained. Accordingly, this may reflect changes in the physical world more fully and timely, thereby a more accurate AR depiction may be provided.
[0052] In some examples, additional information received from a sensor mounted on the physical object may be utilized to increase precision. For example, when the physical object is a vehicle, the sensor may be an onboard gyroscope and/or an accelerometer sensor. As data on the position of the vehicle may be obtained directly from the vehicle mounted sensor, precision of the position of the vehicle may be improved.
[0053] Details of exemplary systems and methods to achieve the aforementioned advantages and benefits are described herein. However, alternatives to the structure, layout, arrangement, etc., are contemplated without departing from the examples and aspects of the
present disclosure. Although the disclosure may be described with respect to a particular vehicle, the disclosure is not restricted to any particular vehicle. The disclosure can be applied to heavy- duty vehicles, such as trucks, buses, and construction equipment, among other vehicle types. Furthermore, in addition to vehicles, the disclosure can also be applied to various other scenarios.
[0054] FIG. 1 shows an exemplary AR system according to an example of the disclosure. In FIG. 1, the aforementioned AR device and the physical object are illustrated as AR headset 10 with AR lens 10a and vehicle 20, respectively. The AR headset 10 and vehicle 20 may communicate with each other via a network 50. The network 50 may be any kind of applicable networks that provide appropriate network connections therebetween. In other examples, such exemplary AR device may be a head mounted see through AR glass and the like. When a user wears the AR headset 10, the AR headset 10 may capture an image of the vehicle 20 which is observed by the user. Then operations described in examples and aspects of this disclosure may be carried out by the AR headset 10, based on the image it captured. In addition, in the situation that the AR headset 10 has limited performance capability, or an off-loading of processing from the AR headset 10 is desired, some operations may be performed by another device, such as a remote server 30. In this case, the AR headset 10 may communicate with the remote server 30 via the network 50, to transmit information, such as the image it captured, to the remote server and receive information from the remote server 30. For example, the remote server 30 may process the image received from the AR headset 10 to identify the plurality of key features of the physical object by using a trained object recognition model. In another example, the remote server 30 may train the object recognition model, which may be maintained by the remote server
30 or loaded into the AR headset 10. The remote server 30 may access database 40 as needed to perform those relevant operations.
[0055] FIG. 2 illustrates an exemplary diagram of the AR headset 10 in FIG. 1. The AR headset 10 may include a system bus 100 to communicatively connect the components included therein. A non-transitory computer memory 110 may store necessary data and information for operation of the AR headset 10. For example, computer-readable instructions 112 may be stored in the non-transitory computer memory 110. Processor 120 may read the computer-readable instructions 112 stored in the non-transitory computer memory 110, for example, via the system bus 100. Upon reading the computer-readable instructions 112, the processor 120 may be configured to perform various operations described herein in relation to the AR headset 10. Furthermore, an image capturer 130 may be provided for capturing an image or photo 202 of the vehicle 20. The image capturer 130 may have a built in AR lens 10a as shown in FIG. 1 for certain camera effects as needed. For example, AR lens 10a may, such as simulate the size, position, and field of view of human eyes, and/or add AR elements to the digital world to be displayed by the headset 10. A transceiver 140 communicates with the remote server 30 when needed. For example, transceiver 140 may transmit the image 202 captured by the AR headset 10 to the remote server 30. Similarly, transceiver 140 may also receive information from the remote server 30, for example, a plurality of key features 204 in the image 202 which are identified by the remote server 30. More detailed information about the key features 204 may be stored in the non-transitory computer memory 110, a memory built in or communicatively connected with the remote server 30, or any other appropriate storage devices. Such detailed information, as will be described below, may indicate, for example, characteristic information of the key features 204, such as what is an identified key feature 204. The detailed information may be retrieved as
needed, when, for example, it may assist with aligning the physical world with the digital world displayed by the AR headset 10, or it may facilitate to retrieve AR content to be added to the displayed digital world, etc.
[0056] Moreover, a depth sensor 150 may be used to conduct measurements during the process of creating a virtual spatial representation 151 of a physical object in the real world. The virtual spatial representation 151 represents what a user sees through the AR headset 10. The depth sensor 150 may acquire information relating to multiple-point distance across, for example, a wide field-of-view (FOV) of surroundings of the AR headset 10 and help to create the virtual spatial representation 151. When the location of the AR headset 10 changes, the depth sensor 150 may update its measurements to capture all the changes in the scene. To discover the depth of the environment surrounding the AR headset 10, different technologies may be used, as in known depth sensors. Therefore, relevant devices alternatively or additionally to the ones shown in FIG. 2 and described herein may be employed according to the particular use. Depending on the functions the AR headset 10 carries and technologies the AR headset 10 employs, other components may also be provided therein.
[0057] Additionally, the AR headset 10 may also maintain or access a trained object recognition model 160, which may be stored in the non-transitory computer memory 110. The trained object recognition model 160 may comprise one or more training data sets 160a. A training set 160a may be a plurality of images (e.g., hundreds or thousands of images). Features of interest in the images may be annotated or labeled. An exemplary annotated image is shown in FIG. 3 A. The image is a street view. In this image, traffic related information is annotated and labeled, including people, traffic lights and vehicles. FIG. 3B shows exemplary types of physical objects that may be used as training in some examples. Each type of physical objects may be
characterized by one or more key parts or features of the physical object. In some examples, the key parts may be headlights, logo, or the like of a vehicle. Detected differences in size, shape, location, etc., may characterize a different key part. In FIG. 3B, five different types of vehicles, which may represent five different truck families including thousands of trucks, which may be used by the trained object recognition model 160. The training data may then cover the full range of trucks within the same family during the lifecycle of the product family, as long as the corresponding key feature(s) has not been changed. In some examples, the exemplary different types of vehicles in FIG. 3B may be considered as specific examples of the vehicle 20.
[0058] FIG. 3C shows an exemplary coordinate system of a physical object according to an example of the disclosure. A coordinate system may be used to position a physical object in the real world, and used to calculate a precise positioning of any digital information to be overlaid on the physical object. FIG. 3C shows a 3D physical object 301 and a coordinate system 302. In this example, the 3D physical object 301 is illustrated as a tractor. The coordinate system 302 is a 3D coordinate system XYZ, wherein the origin is (0,0,0). In some examples, the trained object recognition model 160 may place the 3D physical object 301 relative to the origin of the coordinate system 302. In this way, relative positions of all the components of the 3D physical object 301 learned by the trained object recognition model 160. For example, the trained object recognition model 160 may know where the headlights, the logo, and other key features are located in the coordinate system 302.
[0059] In some examples, the AR headset 10 may recognize key features and their respective positions on the physical object. Those positions may be correlated to the coordinate system 302 in the physical world (e.g., by knowing relative positions of the key features to the origin). The origin of the coordinate system 302 may be used to determine relative positioning of
those key features and thereby positioning the physical object and optionally any digital information on the physical object. By the correlation, the AR headset 10 may display the physical object to the user and display relevant digital information in the proper location. The AR headset 10 may merge the physical world and digital world. For example, the AR headset 10 may display additional information regarding the physical object, as shown in Fig. 8B
[0060] The features of interest may also be the key features 204 herein, such as the lights and logo in figures 4A and 4B. The object recognition model 160 may contain an algorithm for a processing of the stored training data sets 160a. For example, the object recognition model 160 may include weights and relevant calculations for processing an image of the training data sets 160a, in order to determine if the image comprises a key feature 204 and where that key feature is located.
[0061] In this way, the trained object recognition model 160 may support the AR headset 10 to identify the key features 204 from the image 202, and also their respective positions in the image 202. The trained object recognition model 160 may be trained in a plurality of manners. As an example, a first training data set 160a comprising images of the physical vehicle 20 may be obtained. The first training data set 160a may include, for example, images of the different vehicle models available from a particular manufacturer. Key features of the physical vehicle 20 in the first training data set 160a may be identified by processing the images of the physical vehicle 20 using an object detection algorithm. The first training data set 160a may then be reduced into a second training data set 160a based on the identified restricted number of key features of the physical vehicle 20. The trained object recognition model 160 may be obtained by training an object recognition model to recognize the key features of the physical vehicle 20 in real-time image streams based on the reduced second training data set 160a. The exemplary
training may be done in advance, for example by the remote server 30 or other appropriate processing apparatus. The trained object recognition model 160 may also be trained during use. In another example, the trained object recognition model 160 may be updated periodically or as needed when stored in the AR headset 10.
[0062] The trained object recognition model 160 may also be trained with genericized data, such as the genericized key features which will be described in detail in the following. Upon being trained with relevant genericized information for a plurality of similar objects, such as for multiple types of trucks, the trained object recognition model 160 may identify a truck captured by the image capturer without requiring the user to indicate or input the specific type of the truck within his/her FOV. With the genericized data training, computations by the trained object recognition model 160 and therefore the overall processing time by the AR headset 10 may faster, more efficient. Furthermore, this renders the application of the trained object recognition model 160 and accordingly the AR headset possibly to various scenarios to identify an object in the physical world without preexisting knowledge or identification of the specific type of the object in advance. Therefore, in an example, a user will not be interrupted by being required to input details of the object that he/she is observing (e.g., the type of a truck) when looking through the AR headset. In this case, it’s possible that a user may enjoy a real-time AR display of a truck in front of him/her, then when turning around and facing another truck (either the same or different type of truck) nearby, they may continue to enjoy an AR display of that truck without any interruptions.
[0063] In the following, detailed operations performed by the AR headset 10 with its components mentioned above and illustrated in FIG. 2 will be described in conjunction with FIGs. 3D-6B.
[0064] With continued reference to FIG. 2 the image capturer 130 may be triggered to take an image of the vehicle 20, for example when a user facing the vehicle 20 puts the AR headset 10 on or when the user moves relative to the vehicle 20. The image capturer 130 may be an AR camera. In the example of FIG. 2, the image capturer 130 captures a two-dimensional (2D) image 202 of a physical object. In other examples, the image may take many other forms, such as views from multiple AR cameras, multi-dimensional data from a three-dimensional (3D) image capturer, etc. The image capturer 130 may optionally capture live images of vehicle 20 in a real-time manner. Also, the depth sensor 150 mounted on the AR headset 10 may capture and create the virtual spatial representation 151 of vehicle 20. The virtual spatial representation 151 for example, may be a spatial mesh. FIG. 3D shows an exemplary virtual spatial representation of a physical object according to an example of the disclosure, with different views of the physical object. In FIG. 3D, the depth sensor 150 creates a spatial mesh 151a of the physical environment including a vehicle.
[0065] With continued reference to FIG. 2, the AR headset 10 may identify a plurality of key features 204 together with corresponding positions (e.g., pixel positions) from the image 202 captured by the image capturer 130, with the trained object recognition model 160, without using a CAD model. In some examples, a key feature may be a predefined part of the physical object. For example, a key feature may be a noticeable design part of the physical object. In another example, a key feature may represent a unique part of a physical object or a particular type of physical objects, as compared to others. As an example, when the physical object is a vehicle, key features may be one or more of headlights, logo, chassis, tire, etc. FIGs. 4A and 4B show exemplary images 202a and 202b of a physical object and key features 204 thereof respectively, according to examples of the disclosure. Specifically, FIGs. 4A and 4B show images of two
different types of trucks as illustrated exemplary scenarios. Also, different views of the trucks from the user are shown. From the front view of a truck shown in FIG. 4A, the noticeable design parts may be one or more of the left headlight 204a, right headlight 204b, logo 204c, upper rearview mirrors 204d, lower rear-view mirrors 204e, wipers 204f or the like. Additionally, for the side view of another truck shown in FIG. 4B, the noticeable design parts may be one or more of wind deflector 204g, tire 204h, chassis 204i, the door 204j on the driver side, the knob 204k on the door, the tanks 2041or the like. Therefore, the key features 204 may be selected from a plurality of those components.
[0066] In some examples, the key features 204 may be genericized features of the object. Such genericized features may be common features for the object or similar objects. For example, for vehicles being the objects that observed by a user, the genericized features may be any features that a vehicle commonly has, such as a front windshield, a logo, a right headlight, a left headlight, a front bumper, etc. The genericized features may also be features that shared by certain objects, including the inter relationships among those shared features. For example, for trucks that belong to the same series, they may have a logo, a right and left headlights which are located at the same or similar approximate locations of the trucks. Similarly, the distance between the right and left headlight, the distance between the right headlight and the logo, and the distance between the left headlight and the logo, may also be the same or similar for the same series of trucks. In this way, genericizing features may allow an identification of similar physical objects. For example, key features (e.g., headlights and logo) genericized from multiple vehicles may be used for identifying those different vehicles.
[0067] In an example, some different vehicles may have the same genericized features if they belong to a same series, a same brand, or if they are considered as a same classification or
type produced by same/different manufactures, or the like. In an example, for the same series of trucks, the distance between the right and left headlights, the distance from the right headlight to the logo, and the distance from the left headlight to the logo, are approximately the same or within a reasonable difference value. Then such genericized key features (e.g., the right and left headlights and the logo) may be identified by the AR headset 10, by using for example the trained object recognition model 160, for various vehicles that belong to the same series from which those key features are genericized. As another example, an object recognition model may be trained with the same genericized physical features for a number of different series of trucks of the same brand, or different brands.
[0068] Once the key features are identified, they can be matched with a corresponding vehicle. Accordingly, the trained object recognition model 160 may be used for identifying multiple types of vehicles, for example, without preexisting knowledge of the specific vehicle type. Thus, when a user is using the AR application (e.g., wearing the AR headset 10), the user may not be interrupted and required to input further information about the physical object, such as the specific type or model of a truck. Various key features genericized in a similar way may be used for identifying a plurality of different, but similar physical objects.
[0069] For each key feature 204, the AR headset 10 may also identify its respective position in the captured image 202. In the situation shown in FIG. 4B in which they key features are components with a certain size, rather than a single point, the relevant size may be measured by the AR headset 10, such as by an operation conducted by the processor 120. The AR headset 10 may subsequently select the central point of the relevant component as the position of that key feature. 1
[0070] As mentioned above, by using the trained object recognition model 160 without depending on a CAD model, the process of orienting the spatial mesh in the 3D space by identifying the key features 204 from the image 202 and their positions in the image 202 by the AR headset 10 may be less computationally complex than current methods. Thus the identification time may be reduced. Accordingly, less resources may be consumed, and lower performance requirements imposed on the AR headset 10. The trained object recognition model 160 may be developed with appropriate training algorithm(s) with the help of advanced AR technologies, such as computer vision to extract information from an image. Computer vision may be employed for acquiring, processing, analyzing and understanding digital images like image 202, and as well as extracting high-dimensional data (e.g., 3D digital information) from the digital images taken from the real world. Then numerical or symbolic information may be produced accordingly. In this way, a transformation of visual images into descriptions of the physical object being observed in the real world is carried out by using models constructed under the aid of geometry, physics, statistics, leaning theories, etc. In this regard, sub-domains of computer vision may be utilized as needed, for example object detection, object recognition, three-dimensional pose estimation, learning, indexing, motion estimation, three-dimensional scene modeling, image restoration, or the like.
[0071] In some examples, the object recognition model 160 may be trained in advance, and then stored on the AR headset 10. In this scenario, the trained object recognition model 160 on the AR headset 10 may be updated periodically, or as needed. In some examples wherein an off-loading of processing from the AR headset 10 is desired, the trained object recognition model 160 may be developed and maintained at the remote server 30. In this case, it may be stored in the database 40. Then when the AR headset 10 is in use (e.g., powered on, or triggered to work,
etc.), a communicative connection between the remote server 30 and the AR headset 10 may be established via the transceiver 140. The remote server 30 may perform the operations to identify the key features 204 and corresponding positions from the image 202 received from the transceiver 140. Subsequently, the transceiver 140 receives the identified key features 204 together with the respective position of each of the key features 204 in the captured image 202.
[0072] With continued reference to FIG. 2, with the identified key features 204 and their respective positions, the AR headset 10 may project rays through image 202 onto the virtual spatial representation 151 of vehicle 20. In this way, the plurality of key features 204 are projected onto the virtual spatial representation 151. Examples of such projection are shown in FIGs. 5, 6A and 6B respectively. In FIG. 5, an exemplary projection of a key feature onto a virtual spatial representation of a vehicle as shown in FIG. 3D is illustrated. As shown in FIG. 5, the key feature herein is a right headlight 204b of a vehicle, and its location in one of its captured image 202c is indicated as 205. A spatial mesh of the vehicle is shown as a virtual spatial representation 151a, with the key feature 204b located at location 205b. To align the image 202c to the virtual spatial representation 151a, rays 50 are projected through the identified location 205a of image 202c onto the corresponding location 205b of the virtual spatial representation 151a. Similarly, FIG. 6A shows an exemplary projection of rays through a 2D image 202a shown in FIG. 4A onto a vehicle’s virtual spatial representation 151b and the overlaying of the 2D image 202a, and FIG. 6B illustrates an enlarged view thereof. By using the trained object recognition model 160, and without using a CAD model, as shown in those figures 6A and 6B, the AR headset 10 may project rays 50 to align the image 202a to the virtual spatial representation 151b. The key features 204a and 204b are left headlight and right headlight respectively. The rays 50 are projected through the 2D image 202a so that key features 204a and
204b in the image 202a are projected onto the virtual spatial representation 151b. Then upon an alignment, image 202a is overlayed to the virtual spatial representation 151b.
[0073] The AR headset 10 may project the image 202 (e.g., image 202a in FIGs. 6 A and 6B) onto the virtual spatial representation 151 (e.g., virtual spatial representation 151b in FIGs. 6A and 6B) by using a projection matrix. The projection matrix may compensate for visual difference between a visual scene of the physical world from the location of the camera and the AR view displayed to the user by the user looking through AR headset 10. In other words, the visual difference may be due to a location difference of an AR camera and the location of human eyes. Then by, for example, detecting a position intersection for each of the key features 204 between the image 202 and the virtual spatial representation 151, the orientation (e.g., an upright orientation) of vehicle 20 in the virtual spatial representation 151 may be determined. In this manner, the obtained 2D image 202 is aligned to the virtual spatial representation 151. In some examples, the orientation of vehicle 20 is determined without identifying the particular vehicle 20. For example, the type of vehicle 20 may be identified during the process of identification of the key features, projection and alignment mentioned above, but not the specific model. In some examples, more details of vehicle 20 may be identified during the process, such as the specific model, the plate number, etc., if such details are needed later on for such as retrieving relevant AR content or the like.
[0074] In a situation wherein three or more key features are identified by the AR headset 10, the positions of those key features may be used by the AR headset 10 to determine for example an orientation of vehicle 20 in digital space. With three or more key features being identified by the AR headset 10, the respective location for each of those three or more key features in image 202 is identified as well. As three locations may define a plane, the AR headset
10 may determine an upright orientation of vehicle 20 by using those identified locations of the three or more key features. For example, to determine the upright orientation of vehicle 20, the AR headset 10 may triangulate its logo, left and right headlights. By projecting the key features 204 onto the virtual spatial representation 151 and determining the orientation, the AR headset 10 may position the digital image 202 in an AR visual display for the user. The digital image 202 overlays the AR visual display (e.g., the virtual spatial representation 151) relative to the identified positions of the key features 204.
[0075] As mentioned above, in some examples, images of a physical object may be taken in a real-time manner resulting in a real-time image stream including a series of images. In other examples, it may be not a real-time image stream, but any series of images, for example, images captured close in time. In this way, a cluster of positions for a specific key feature may be obtained. Specifically, the trained object recognition model 160 may consider positions extracted from different images as belonging to a cluster of a particular key feature of vehicle 20 if such as the difference(s) in relation to the location of the same particular key feature in those different images is/are within a certain reasonable range. The reasonable range may be predetermined or predefined, it may also be trained along with the training of the trained object recognition model 160. In this case, an average central position calculated based on the positions identified from the series of images and considered as within the same cluster of a particular key feature, is taken as the position of this key feature. Then the projection of rays may be performed through a selected image, some images, or the series of images, onto the virtual spatial representation 151. The introduction of such redundancy may help improve the accuracy of identification of key features and relevant positions thereof. For example, this may help the AR device or the remote server to differentiate and discard extraneous data. Such data may be, for example, information that
relating to other physical objects nearby the physical object that the user is observing. As an example, information about a headlight or other parts of another vehicle close to the vehicle being observed by the user may be captured by the image capturer 130 and thus present in the image 202.
[0076] In the exemplary projections shown in FIGs. 6A and 6B, the identified key features are shown as logo and headlights. However, in other examples, for example, when the user’s view of a vehicle is from its rear or side view, the key features may be other components of the vehicle as mentioned above. Similarly, when the object being observed by the user is not a vehicle, other noticeable components may be used to train an object recognition model.
[0077] With continued reference to FIG. 2, in some other examples, a communicative connection between the AR headset 10 and vehicle 20 may be established via the transceiver 140. Then the transceiver 140 may obtain information from a sensor mounted on the vehicle 20. For example, when a gyroscope or accelerometer sensor is mounted on vehicle 20, optionally, information on position or the orientation of vehicle 20 may be transmitted to the transceiver 140 from the sensor. Then such information may help the determination of the orientation of the virtual spatial representation 151. Such determination may be made based on the information received from the sensor and the projection of the features 204. By obtaining supplemental or additional information from the sensor, the accuracy of orientation determination may be improved.
[0078] FIG. 7 shows an exemplary flow chart of an exemplary method according to an example of the disclosure. This illustrated method may be performed by an AR device, for example, by the AR headset 10 of FIG. 2. An AR device may be triggered to work at power-on, or by other triggering mechanisms such as unfolding the legs of an AR glass, or any feasible or
possible ways. At step 701, the AR headset 10 obtains an image of a physical object (e.g., vehicle 20 of FIG. 2) captured by an image capturer (e.g., the image capturer 130 of FIG. 2). At step 702, the AR headset 10 identifies, using the trained object recognition model 160, a plurality of key features 204 of vehicle 20 in the image 202 and a respective position for each of the key features 204. At step 703, the AR headset 10 obtains a virtual spatial representation 151 of vehicle 20 captured by the depth sensor 150. At step 704, the AR headset 10 projects rays through the image 202 onto the virtual spatial representation 151 of the physical object (i.e., vehicle 20) such that the plurality of key features 204 are projected onto the virtual spatial representation 151. At step 705, the AR headset 10 determines an orientation of the virtual spatial representation 151 based on the projection of the plurality of key features 204 onto the virtual spatial representation 151.
[0079] FIG. 8A shows another exemplary AR system with a user’s AR visual view according to an example of the disclosure. In FIG. 8A, the illustrated AR system is similar as the one shown in FIG. 1, with additional functions performed by the remote server 30. In this example, the remote server 30 performs some of the aforementioned operations conducted by the AR headset 10. For example, the remote server 30 may perform computing vision training 30a for developing models as needed, such as the trained object recognition model 160. The remote server 30 may also have a database for records of physical objects 30b (such as physical vehicles) for training sets. The remote server 30 may maintain other information and functional models or the like as needed in an AR application context. Furthermore, in some examples, the remote server 30 may also provide characteristic information of the identified key features 204. For example, in the example shown in FIG. 4B, the remote server may indicate that the identified features 204 are a logo of a vehicle, a left headlight and a right headlight thereof, respectively.
This information may assist in the determination of the orientation of the vehicle 20. For example, with the characteristic information obtained from the remote server 30, a positional relationship among the three key features 204 may help determine the vehicle’s orientation in digital world. The positional relationship could be the distance between two headlights, the distance between one headlight and the logo, etc.
[0080] Moreover, FIG. 8A also shows an AR visual view displayed to the user. In the AR visual view, an orientation of the virtual spatial representation of a physical object (i.e., a vehicle herein) in the real world is determined and adjusted before being displayed to the user. This means the vehicle is located in the digital world in alignment with where the vehicle actually is in the real world. In other words, the virtual spatial representation is ready for the AR content. In case that characteristic information of the plurality of key features are indicated by the remote server, relevant information to be overlaid on, or additions to the AR visual view may be located more efficiently. For example, when a key feature is a logo, information on this logo or brand may be selected to combine with the surrounding real world of the user. In this manner, the physical and real world is combined with the virtual world, which make the user’s display interactive and digitally manipulated. Because none of those operations and processes require CAD based computations, seamless experiences may be obtained because of less computational tasks are performed. The improvement may be more apparent when the user is moving in relation to an object. The movement may cause the operations to be carried out on a real-time basis.
[0081] FIG. 8B shows an exemplary AR view displayed to the user, according to an example of the disclosure. In FIG. 8B, a user wearing an AR headset 10 may walk around a vehicle. The AR headset 10 may identify the key features of the vehicle, as discussed above.
Then by relating the respective positions of the key features to the coordinate system, the AR headset 10 may recognize that the user is on the left side of the vehicle and looking down at the wheelbase of the vehicle. Then the AR headset 10 may display a digital object 801 in the headset to the user. The digital object 801 may illustrate that the user is on the left of a 4 wheelbase, with additional details of nearby components presented in the appropriate location. In such examples, relevant digital information may be overlaid on the current view of the physical object in the real world. Therefore, the virtual world may be seamlessly interwoven with the physical world, which reflects the latest change in user’s view of the physical world in the AR digital world displayed to the user, may be achieved such that the user may perceive an immersive aspect of the real physical world.
[0082] FIG. 9 is a schematic diagram of an exemplary computer system for implementing examples disclosed herein in relation to an AR device according to an example of the disclosure. The schematic diagram of a computer system 1000 may be utilized for implementing examples disclosed herein. Therefore, the computer system 1000 may be incorporated into an existing AR device, like the AR headset 10 in FIG. 2, to fulfil the functions of processor 120 and non-transitory computer memory 110 thereof. The computer system 1000 is adapted to execute instructions from a computer-readable medium to perform these and/or any of the functions or processing described herein. The computer system 1000 may be connected (e.g., networked) to other machines in a LAN, an intranet, an extranet, or the Internet. While only a single device is illustrated, the computer system 1000 may include any collection of devices that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein. Accordingly, any reference in the disclosure and/or claims to a computer system, computing system, computer device, computing device, control
system, control unit, electronic control unit (ECU), processor device, processing circuitry, etc., includes reference to one or more such devices to individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein. For example, control system may include a single control unit or a plurality of control units connected or otherwise communicatively coupled to each other, such that any performed function may be distributed between the control units as desired. Further, such devices may communicate with each other or other devices by various system architectures, such as directly or via a Controller Area Network (CAN) bus, etc.
[0083] The computer system 1000 may comprise at least one computing device or electronic device capable of including firmware, hardware, and/or executing software instructions to implement the functionality described herein. The computer system 1000 may include processing circuitry 1002 (e.g., processing circuitry including one or more processor devices or control units), a memory 1004, and a system bus 1006. The computer system 1000 may include at least one computing device having the processing circuitry 1002. The system bus 1006 provides an interface for system components including, but not limited to, the memory 1004 and the processing circuitry 1002. The processing circuitry 1002 may include any number of hardware components for conducting data or signal processing or for executing computer code stored in memory 1004. The processing circuitry 1002 may, for example, include a general- purpose processor, an application specific processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), a circuit containing processing components, a group of distributed processing components, a group of distributed computers configured for processing, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to
perform the functions described herein. The processing circuitry 1002 may further include computer executable code that controls operation of the programmable device.
[0084] The system bus 1006 may be any of several types of bus structures that may further interconnect to a memory bus (with or without a memory controller), a peripheral bus, and/or a local bus using any of a variety of bus architectures. The memory 1004 may be one or more devices for storing data and/or computer code for completing or facilitating methods described herein. The memory 1004 may include database components, object code components, script components, or other types of information structure for supporting the various activities herein. Any distributed or local memory device may be utilized with the systems and methods of this description. The memory 1004 may be communicab ly connected to the processing circuitry 1002 (e.g., via a circuit or any other wired, wireless, or network connection) and may include computer code for executing one or more processes described herein. The memory 1004 may include non-volatile memory 1008 (e.g., read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.), and volatile memory 1010 (e.g., random-access memory (RAM)), or any other medium which can be used to carry or store desired program code in the form of machineexecutable instructions or data structures and which can be accessed by a computer or other machine with processing circuitry 1002. A basic input/output system (BIOS) 1012 may be stored in the non-volatile memory 1008 and can include the basic routines that help to transfer information between elements within the computer system 1000.
[0085] The computer system 1000 may further include or be coupled to a non-transitory computer- readable storage medium such as the storage device 1014, which may comprise, for example, an internal or external hard disk drive (HDD) (e.g., enhanced integrated drive
electronics (EIDE) or serial advanced technology attachment (SATA)), HDD (e.g., EIDE or SATA) for storage, flash memory, or the like. The storage device 1014 and other drives associated with computer-readable media and computer-usable media may provide non-volatile storage of data, data structures, computer-executable instructions, and the like.
[0086] Computer-code which is hard or soft coded may be provided in the form of one or more modules. The module(s) can be implemented as software and/or hard-coded in circuitry to implement the functionality described herein in whole or in part. The modules may be stored in the storage device 1014 and/or in the volatile memory 1010, which may include an operating system 1016 and/or one or more program modules 1018. All or a portion of the examples disclosed herein may be implemented as a computer program 1020 stored on a transitory or non- transitory computer-usable or computer-readable storage medium (e.g., single medium or multiple media), such as the storage device 1014, which includes complex programming instructions (e.g., complex computer- readable program code) to cause the processing circuitry 1002 to carry out actions described herein. Thus, the computer-readable program code of the computer program 1020 can comprise software instructions for implementing the functionality of the examples described herein when executed by the processing circuitry 1002. In some examples, the storage device 1014 may be a computer program product (e.g., readable storage medium) storing the computer program 1020 thereon, where at least a portion of a computer program 1020 may be loadable (e.g., into a processor) for implementing the functionality of the examples described herein when executed by the processing circuitry 1002. The processing circuitry 1002 may serve as a controller, or control system, for the computer system 1000 that is to implement the functionality described herein.
[0087] The computer system 1000 may include an input device interface 1022 configured to receive input and selections to be communicated to the computer system 1000 when executing instructions, such as from a keyboard, mouse, touch-sensitive surface, etc. Such input devices may be connected to the processing circuitry 1002 through the input device interface 1022 coupled to the system bus 1006 but can be connected through other interfaces such as a parallel port, an Institute of Electrical and Electronic Engineers (IEEE) 1394 serial port, a Universal Serial Bus (USB) port, an IR interface, and the like. The computer system 1000 may include an output device interface 1024 configured to forward output, such as to a display, a video display unit (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)). The computer system 1000 may include a communications interface 1026 suitable for communicating with a network as appropriate or desired.
[0088] The operational actions described in any of the exemplary aspects herein are described to provide examples and discussion. The actions may be performed by hardware components, may be embodied in machine-executable instructions to cause a processor to perform the actions, or may be performed by a combination of hardware and software. Although a specific order of method actions may be shown or described, the order of the actions may differ. In addition, two or more actions may be performed concurrently or with partial concurrence.
[0089] Example 1: A device for controlling augmented reality, comprising: a non- transitory computer memory operable to store computer-readable instructions; and a processor operable to read the computer-readable instructions, the processor when executing the computer- readable instructions is configured to: obtain an image of a physical object captured by an image capturer; identify, using a trained object recognition model, a plurality of key features of the
physical object in the image and a respective position for each of the plurality of key features in the image; obtain a virtual spatial representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determine an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
[0090] Example 2: The device of example 1, wherein the processor when executing the computer-readable instructions is further configured to detect a pixel position in the image for each of the plurality of the key features.
[0091] Example 3: The device of example 1, wherein the plurality of key features are projected, using a projection matrix, onto the virtual spatial representation by aligning the image with the virtual spatial representation.
[0092] Example 4: The device of example 1, wherein the processor when executing the computer-readable instructions is further configured to overlay, in the virtual spatial representation, the image relative to the identified respective position of the plurality of key features of the physical object.
[0093] Example 5: The device of example 1, wherein the plurality of key features of the physical object and the respective position for each of the plurality of key features are identified by a remote server, and wherein the processor when executing the computer-readable instructions is further configured to: transmit the image of the physical object to the remote server; and receive the plurality of key features and the respective position from the remote server.
[0094] Example 6: The device of example 1, wherein the processor when executing the computer- readable instructions is further configured to: project the rays through the image onto the virtual spatial representation of the physical object to detect a position intersection for each of the plurality of key features of the physical object between the image and the virtual spatial representation; and determine the orientation of the virtual spatial representation based on the detected position intersection.
[0095] Example 7: The device of example 5, wherein the processor when executing the computer-readable instructions is further configured to: obtain characteristic information of at least one of the plurality of key features of the physical object from the remote server; and determine the orientation of the virtual spatial representation based on the characteristic information and the projection of the respective positions of the plurality of key features.
[0096] Example 8: The device of example 1, wherein the processor when executing the computer- readable instructions is further configured to: obtain information indicating a position or the orientation of the physical object from a sensor mounted on the physical object; and determine the orientation of the virtual spatial representation based on the information received from the sensor and the projection of the respective positions of the plurality of key features.
[0097] Example 9: A moveable device for tracking a physical object, comprising: an image capturer, for capturing an image of the physical object; and the device of example 1.
[0098] Example 10: The moveable device of example 9, wherein the moveable device is an augmented reality headset.
[0099] Example 11 : The moveable device of example 9, wherein the plurality of key features of the physical object and the respective position are identified by a remote server, and
the moveable device further comprises, a transceiver, for transmitting the image of the physical object to the remote server and receiving the plurality of key features of the physical object and the respective position from the remote server.
[0100] Example 12: The moveable device of example 9, wherein the physical object is a vehicle, and the plurality of key features are at least one of a headlight, logo, tire and/or chassis.
[0101] Example 13: The moveable device of example 9, wherein the image of the physical object is a two-dimensional image captured by the image capturer when the moveable device is moving relative to the physical object.
[0102] Example 14: The moveable device of example 10, wherein the augmented reality headset includes an AR lens.
[0103] Example 15: The moveable device of example 9, wherein, the image capturer captures a series of images of the physical object; and project the rays through at least one of the captured images to the virtual spatial representation of the physical object; wherein, the respective position for each of the plurality of key features is an average central position identified from the series of images.
[0104] Example 16: The moveable device of example 11, wherein: the transceiver further receives information indicating a position or the orientation of the physical object from a sensor mounted on the physical object; and the processor when executing the computer-readable instructions is configured to, determines the orientation of the virtual spatial representation based on the information received from the sensor and the projection of the respective positions of the plurality of key features.
31
[0105] Example 17: The moveable device of example 16, wherein the sensor is a gyroscope and/or an accelerometer sensor.
[0106] Example 18: A method for controlling augmented reality, comprising: obtaining, by a processor of an augmented reality device, an image of a physical object captured by an image capturer; identifying, by the processor, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtaining, by the processor, a virtual spatial representation of the physical object captured by a depth sensor; projecting, by the processor, rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determining, by the processor, an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
[0107] Example 19: An augmented reality system, comprising: a remote server for identifying, by using a trained object recognition model, a plurality of key features of a physical object and a respective position for each of the plurality of key features from an image of the physical object; and an augmented reality headset, for transmitting the image of the physical object to the remote server, and for receiving the plurality of key features and the respective positions from the remote server, wherein, the augmented reality headset comprises: an image capturer, for capturing the image of the physical object; a transceiver, for transmitting the image to the remote server and receiving the plurality of key features and the respective positions from the remote server; a non-transitory computer memory operable to store computer-readable instructions; and a processor operable to read the computer-readable instructions, the processor when executing the computer-readable instructions is configured to: obtain a virtual spatial
representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determine an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
[0108] Example 20: The augmented reality system of example 19, wherein the image of the physical object is captured by the image capturer when the augmented reality headset is moving relative to the physical object, and the remote server: receives a series of images of the physical object from the augmented reality headset, and identifies, using the trained object recognition model, the plurality of key features of the physical object and the respective position for each of the plurality of key features from the received series of images, wherein the respective position is an average central position identified from the received series of images.
[0109] The terminology used herein is for the purpose of describing particular aspects only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. As used herein, the term "and/or" includes any and all combinations of one or more of the associated listed items. It will be further understood that the terms "comprises," "comprising," "includes," and/or "including" when used herein specify the presence of stated features, integers, actions, steps, operations, elements, and/or components, but do not preclude the presence or addition of one or more other features, integers, actions, steps, operations, elements, components, and/or groups thereof.
[0110] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to
which this disclosure belongs. It will be further understood that terms used herein should be interpreted as having a meaning consistent with their meaning in the context of this specification and the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0111] It is to be understood that the present disclosure is not limited to the aspects described above and illustrated in the drawings; rather, the skilled person will recognize that many changes and modifications may be made within the scope of the present disclosure and appended claims. In the drawings and specification, there have been disclosed aspects for purposes of illustration only and not for purposes of limitation, the scope of the disclosure being set forth in the following claims.
Claims
1. A device for controlling augmented reality, comprising: a non-transitory computer memory operable to store computer-readable instructions; and a processor operable to read the computer-readable instructions, the processor when executing the computer- readable instructions is configured to: obtain an image of a physical object captured by an image capturer; identify, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtain a virtual spatial representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and determine an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
2. The device of claim 1, wherein the processor when executing the computer- readable instructions is further configured to detect a pixel position in the image for each of the plurality of the key features.
3. The device of claim 1, wherein the plurality of key features are projected, using a projection matrix, onto the virtual spatial representation by aligning the image with the virtual spatial representation.
4. The device of claim 1, wherein the processor when executing the computer- readable instructions is further configured to overlay, in the virtual spatial representation, the image relative to the identified respective position of the plurality of key features of the physical object.
5. The device of claim 1, wherein the plurality of key features of the physical object and the respective position for each of the plurality of key features are identified by a remote server, and wherein the processor when executing the computer-readable instructions is further configured to: transmit the image of the physical object to the remote server; and receive the plurality of key features and the respective position from the remote server.
6. The device of claim 1, wherein the processor when executing the computer- readable instructions is further configured to: project the rays through the image onto the virtual spatial representation of the physical object to detect a position intersection for each of the plurality of key features of the physical object between the image and the virtual spatial representation; and determine the orientation of the virtual spatial representation based on the detected position intersection.
7. The device of claim 5, wherein the processor when executing the computer- readable instructions is further configured to: obtain characteristic information of at least one of the plurality of key features of the physical object from the remote server; and determine the orientation of the virtual spatial representation based on the characteristic information and the projection of the respective positions of the plurality of key features.
8. The device of claim 1, wherein the processor when executing the computer- readable instructions is further configured to: obtain information indicating a position or the orientation of the physical object from a sensor mounted on the physical object; and determine the orientation of the virtual spatial representation based on the information received from the sensor and the projection of the respective positions of the plurality of key features.
9. A moveable device for tracking a physical object, comprising: an image capturer, for capturing an image of the physical object; and the device of claim 1.
10. The moveable device of claim 9, wherein the moveable device is an augmented reality headset.
11. The moveable device of claim 9, wherein the plurality of key features of the physical object and the respective position are identified by a remote server, and the moveable device further comprises, a transceiver, for transmitting the image of the physical object to the remote server and receiving the plurality of key features of the physical object and the respective position from the remote server.
12. The moveable device of claim 9, wherein the physical object is a vehicle, and the plurality of key features are at least one of a headlight, logo, tire and/or chassis.
13. The moveable device of claim 9, wherein the image of the physical object is a two-dimensional image captured by the image capturer when the moveable device is moving relative to the physical object.
14. The moveable device of claim 10, wherein the augmented reality headset includes an AR lens.
15. The moveable device of claim 9, wherein, the image capturer captures a series of images of the physical object; and project the rays through at least one of the captured images to the virtual spatial representation of the physical object; wherein, the respective position for each of the plurality of key features is an average central position identified from the series of images.
16. The moveable device of claim 11 , wherein: the transceiver further receives information indicating a position or the orientation of the physical object from a sensor mounted on the physical object; and the processor when executing the computer-readable instructions is configured to, determines the orientation of the virtual spatial representation based on the information received from the sensor and the projection of the respective positions of the plurality of key features.
17. The moveable device of claim 16, wherein the sensor is a gyroscope and/or an accelerometer sensor.
18. A method for controlling augmented reality, comprising: obtaining, by a processor of an augmented reality device, an image of a physical object captured by an image capturer; identifying, by the processor, using a trained object recognition model, a plurality of key features of the physical object in the image and a respective position for each of the plurality of key features in the image; obtaining, by the processor, a virtual spatial representation of the physical object captured by a depth sensor; projecting, by the processor, rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and
determining, by the processor, an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
19. An augmented reality system, comprising: a remote server for identifying, by using a trained object recognition model, a plurality of key features of a physical object and a respective position for each of the plurality of key features from an image of the physical object; and an augmented reality headset, for transmitting the image of the physical object to the remote server, and for receiving the plurality of key features and the respective positions from the remote server, wherein, the augmented reality headset comprises: an image capturer, for capturing the image of the physical object; a transceiver, for transmitting the image to the remote server and receiving the plurality of key features and the respective positions from the remote server; a non-transitory computer memory operable to store computer-readable instructions; and a processor operable to read the computer-readable instructions, the processor when executing the computer-readable instructions is configured to: obtain a virtual spatial representation of the physical object captured by a depth sensor; project rays through the image onto the virtual spatial representation of the physical object such that the plurality of key features are projected onto the virtual spatial representation; and
determine an orientation of the virtual spatial representation based on the projection of the plurality of key features onto the virtual spatial representation.
20. The augmented reality system of claim 19, wherein the image of the physical object is captured by the image capturer when the augmented reality headset is moving relative to the physical object, and the remote server: receives a series of images of the physical object from the augmented reality headset, and identifies, using the trained object recognition model, the plurality of key features of the physical object and the respective position for each of the plurality of key features from the received series of images, wherein the respective position is an average central position identified from the received series of images.
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/IB2023/052172 WO2024184675A1 (en) | 2023-03-07 | 2023-03-07 | Device and method for controlling augmented reality |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4677570A1 true EP4677570A1 (en) | 2026-01-14 |
Family
ID=85724655
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP23712617.2A Pending EP4677570A1 (en) | 2023-03-07 | 2023-03-07 | Device and method for controlling augmented reality |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4677570A1 (en) |
| WO (1) | WO2024184675A1 (en) |
-
2023
- 2023-03-07 WO PCT/IB2023/052172 patent/WO2024184675A1/en not_active Ceased
- 2023-03-07 EP EP23712617.2A patent/EP4677570A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024184675A1 (en) | 2024-09-12 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7854048B2 (en) | Alignment method and alignment device for display devices, in-vehicle display system | |
| CN113808160B (en) | Gaze direction tracking method and device | |
| JP6552729B2 (en) | System and method for fusing the outputs of sensors having different resolutions | |
| US11893707B2 (en) | Vehicle undercarriage imaging | |
| CN110060297B (en) | Information processing apparatus, information processing system, information processing method, and storage medium | |
| CN113661495A (en) | Sight line calibration method, sight line calibration device, sight line calibration equipment, sight line calibration system and sight line calibration vehicle | |
| JP6776440B2 (en) | How to assist the driver of a motor vehicle when driving a motor vehicle, driver assistance system and motor vehicle | |
| US11380111B2 (en) | Image colorization for vehicular camera images | |
| CN114041175A (en) | Neural network for estimating head pose and gaze using photorealistic synthetic data | |
| JP2005268847A (en) | Image generating apparatus, image generating method, and image generating program | |
| EP4214681B1 (en) | Camera placement guidance | |
| CN105522971A (en) | Apparatus and method for controlling outputting of external image of vehicle | |
| CN116311131A (en) | Intelligent driving-up enhancing method, system and device based on multi-view looking around | |
| JP2014165810A (en) | Parameter acquisition device, parameter acquisition method and program | |
| CN115810179B (en) | Human-vehicle visual perception information fusion method and system | |
| US12333754B2 (en) | System and method for capturing a spatial orientation of a wearable device | |
| CN112561952A (en) | Method and system for setting renderable virtual objects for a target | |
| JP2021047024A (en) | Estimator, estimation method and program | |
| CN117441190A (en) | Part positioning method and device | |
| WO2024184675A1 (en) | Device and method for controlling augmented reality | |
| US12499615B2 (en) | Systems and methods for 3D accident reconstruction | |
| CN118505512A (en) | Image super-resolution enhancement method, device, equipment and storage medium | |
| CN117036403A (en) | Image processing methods, devices, electronic equipment and storage media | |
| US20260045027A1 (en) | Projective bisector mirror | |
| CN120707649A (en) | Augmented reality HUD testing method, device, vehicle and computer program product |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251006 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |