EP4659197A1 - Object detection accuracy improvement in wireless communications systems at least using radar sensing - Google Patents

Object detection accuracy improvement in wireless communications systems at least using radar sensing

Info

Publication number
EP4659197A1
EP4659197A1 EP23920245.0A EP23920245A EP4659197A1 EP 4659197 A1 EP4659197 A1 EP 4659197A1 EP 23920245 A EP23920245 A EP 23920245A EP 4659197 A1 EP4659197 A1 EP 4659197A1
Authority
EP
European Patent Office
Prior art keywords
neural network
network
information
object detection
physical area
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP23920245.0A
Other languages
German (de)
French (fr)
Inventor
Saeed Reza KHOSRAVIRAD
Sina SHAHSAVARI
Jakub SAPIS
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Nokia Solutions and Networks Oy
Original Assignee
Nokia Solutions and Networks Oy
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nokia Solutions and Networks Oy filed Critical Nokia Solutions and Networks Oy
Publication of EP4659197A1 publication Critical patent/EP4659197A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00Image analysis
    • G06T7/10Segmentation; Edge detection
    • G06T7/11Region-based segmentation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/047Probabilistic or stochastic networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/088Non-supervised learning, e.g. competitive learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/764Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/50Context or environment of the image
    • G06V20/56Context or environment of the image exterior to a vehicle by using sensors mounted on the vehicle
    • G06V20/58Recognition of moving objects or obstacles, e.g. vehicles or pedestrians; Recognition of traffic objects, e.g. traffic signs, traffic lights or roads
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/10Image acquisition modality
    • G06T2207/10028Range image; Depth image; 3D point clouds
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/20Special algorithmic details
    • G06T2207/20084Artificial neural networks [ANN]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00Indexing scheme for image analysis or image enhancement
    • G06T2207/30Subject of image; Context of image processing
    • G06T2207/30248Vehicle exterior or interior
    • G06T2207/30252Vehicle exterior; Vicinity of vehicle

Definitions

  • Exemplary embodiments herein relate generally to wireless communications systems and, more specifically, relates to object detection using wireless communications systems as radar sensing elements.
  • JCAS joint communication and sensing
  • JCAS is being examined for sensing because a wireless communications system has most or all of the infrastructure is in place, e.g., with transmit/receive (Tx/Rx) nodes with a full area coverage as well as a good interconnection between nodes. This allows the wireless communications system to perform radar sensing of objects that are detected by the system.
  • Tx/Rx transmit/receive
  • a method in an exemplary embodiment, includes converting, using a neural network, sensory information of a physical area into one or more images.
  • the method includes performing, using one or more other neural networks, at least object detection on the one or more images to determine an output indicating at least whether or not one or more objects were detected.
  • An additional exemplary embodiment includes a computer program, comprising instructions for performing the method of the previous paragraph, when the computer program is run on an apparatus.
  • the computer program according to this paragraph wherein the computer program is a computer program product comprising a computer-readable medium bearing the instructions embodied therein for use with the apparatus.
  • Another example is the computer program according to this paragraph, wherein the program is directly loadable into an internal memory of the apparatus.
  • An exemplary apparatus includes one or more processors and one or more memories storing instructions that, when executed by the one or more processors, cause the apparatus at least to perform: converting, using a neural network, sensory information of a physical area into one or more images; and performing, using one or more other neural networks, at least object detection on the one or more images to determine an output indicating at least whether or not one or more objects were detected.
  • An exemplary computer program product includes a computer-readable storage medium bearing instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: converting, using a neural network, sensory information of a physical area into one or more images; and performing, using one or more other neural networks, at least object detection on the one or more images to determine an output indicating at least whether or not one or more objects were detected.
  • an apparatus comprises means for performing: converting, using a neural network, sensory information of a physical area into one or more images; and performing, using one or more other neural networks, at least object detection on the one or more images to determine an output indicating at least whether or not one or more objects were detected.
  • a method includes training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area.
  • the training comprises: creating generated depth maps by the generator neural network based at least on corresponding sensory information of a physical area, the generated depth maps comprising images of the physical area; determining classifications of fake or real by a discriminator neural network based on the generated depth maps and one or more real scene depth maps, the one or more real scene depth maps comprising a multi-dimensional map of the physical area, where the classifications are routed to the generator neural network and the discriminator neural network; and performing the creating the generated depth maps and the determining classifications until one or more criteria are met based on one or more loss functions.
  • the method also comprises outputting information that defines a trained generator neural network, the trained generator neural network able to reconstruct images of the physical area based at least on corresponding sensory information of the physical area.
  • An additional exemplary embodiment includes a computer program, comprising instructions for performing the method of the previous paragraph, when the computer program is run on an apparatus.
  • the computer program according to this paragraph wherein the computer program is a computer program product comprising a computer-readable medium bearing the instructions embodied therein for use with the apparatus.
  • Another example is the computer program according to this paragraph, wherein the program is directly loadable into an internal memory of the apparatus.
  • An exemplary apparatus includes one or more processors and one or more memories storing instructions that, when executed by the one or more processors, cause the apparatus at least to perform: training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area, the training comprising: creating generated depth maps by the generator neural network based at least on corresponding sensory information of a physical area, the generated depth maps comprising images of the physical area; determining classifications of fake or real by a discriminator neural network based on the generated depth maps and one or more real scene depth maps, the one or more real scene depth maps comprising a multidimensional map of the physical area, where the classifications are routed to the generator neural network and the discriminator neural network; and performing the creating the generated depth maps and the determining classifications until one or more criteria are met based on one or more loss functions; and outputting information that defines a trained generator neural network, the trained generator neural network able to reconstruct images of the physical area based at least on corresponding sensory information of the physical
  • An exemplary computer program product includes a computer-readable storage medium bearing instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area, the training comprising: creating generated depth maps by the generator neural network based at least on corresponding sensory information of a physical area, the generated depth maps comprising images of the physical area; determining classifications of fake or real by a discriminator neural network based on the generated depth maps and one or more real scene depth maps, the one or more real scene depth maps comprising a multidimensional map of the physical area, where the classifications are routed to the generator neural network and the discriminator neural network; performing the creating the generated depth maps and the determining classifications until one or more criteria are met based on one or more loss functions; and outputting information that defines a trained generator neural network, the trained generator neural network able to reconstruct images of the physical area based at least on corresponding sensory information of the physical area.
  • an apparatus comprises means for performing: training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area, the training comprising: creating generated depth maps by the generator neural network based at least on corresponding sensory information of a physical area, the generated depth maps comprising images of the physical area; determining classifications of fake or real by a discriminator neural network based on the generated depth maps and one or more real scene depth maps, the one or more real scene depth maps comprising a multi-dimensional map of the physical area, where the classifications are routed to the generator neural network and the discriminator neural network; performing the creating the generated depth maps and the determining classifications until one or more criteria are met based on one or more loss functions; and outputting information that defines a trained generator neural network, the trained generator neural network able to reconstruct images of the physical area based at least on corresponding sensory information of the physical area.
  • FIG. 1A is a block diagram of one possible and non-limiting example system in which example embodiments may be practiced;
  • FIG. IB is a block diagram of an example apparatus suitable for implementing any of the nodes in FIG. 1A;
  • FIG. 2 illustrates using access points as radar points to perform object detection
  • FIG. 3 is a block diagram of radar sensing as a function of the environment;
  • FIG. 4 is a flowchart illustrating an example method for performing object detection;
  • FIG. 5 is a diagram illustrating an example approach of Al (artificial intelligence)-based object detection from radar sensing
  • FIG. 6A is a diagram illustrating an example framework for Al-based object detection
  • FIG. 6B is a diagram illustrating another example framework for Al-based object detection
  • FIG. 7 is a block diagram illustrating a training approach for a generator function
  • FIG. 8A illustrates an example 2-dimensional (2D) heat map
  • FIG. 8B illustrates an example 3-dimensional (3D) binary depth map
  • FIG. 9 illustrates an example loss function used for training a generative adversarial network (GAN).
  • GAN generative adversarial network
  • FIG. 10 illustrates an example for a 3D depth map reconstruction
  • FIG. 11 illustrates an example for a 2D depth map reconstruction
  • FIG. 12 illustrates an example object detection neural network
  • FIG. 13 illustrates an example two-step bounding box and object detection
  • FIG. 14 illustrates an example end-to-end (E2E) network for object detection from a radar heat map.
  • Any flow diagram (such as FIG. 4) or signaling diagram herein is considered to be a logic flow diagram, and illustrates the operation of an exemplary method, results of execution of computer program instructions embodied on a computer readable memory, functions performed by logic implemented in hardware, and/or interconnected means for performing functions in accordance with an exemplary embodiment.
  • Block diagrams (such as FIGS. 6, 7, or 12-14) also illustrate the operation of an exemplary method, results of execution of computer program instructions embodied on a computer readable memory, functions performed by logic implemented in hardware, and/or interconnected means for performing functions in accordance with an exemplary embodiment.
  • the exemplary embodiments herein describe techniques for object detection accuracy improvement in wireless communications systems for radar sensing. Additional description of these techniques is presented after a system into which the exemplary embodiments may be used is described.
  • FIG. 1A illustrates an example system in which one or more example embodiments may be practiced.
  • a number of nodes are shown: a user equipment (UE) 110; a base station 170; and network element(s) 190.
  • UE user equipment
  • a user equipment (UE) 110 as one of the nodes, is in wireless communication via wireless link 111 with a wireless network 100.
  • the UE 110 is a wireless, typically mobile device that can access a wireless network.
  • the UE 110 is illustrated with one or more antennas 128.
  • the ellipses 101 indicate there may be multiple UEs 110.
  • the base station 170 provides access by wireless devices such as the UE 110 to the wireless network 100. It is noted that the base station 170 may also be referred to by other names, such as an access point.
  • the base station 170 is illustrated as having one or more antennas 158. There are many options for the base station
  • the base station 170 may be a RAN (radio access network) node, and in particular may be a gNB, which is the primary term used herein. That is, the base station 170 will be referred to as gNB 170.
  • gNB 170 There are, however, many options including an eNB (evolved node B, e.g., an LTE, long-term evolution, base station) for the base station, as described below, or options other than cellular systems.
  • eNB evolved node B, e.g., an LTE, long-term evolution, base station
  • the base station 170 There are a number of configurations for the base station 170.
  • One such is a “standalone” configuration, which includes all circuity as part of a single unit, and accesses the antennas 158. More commonly today, circuitry is split into one or more remote nodes 150 (accessing antennas 158) and central nodes 160.
  • a gNB might include a distributed unit (DU), or DU and radio unit (RU) as the remote nodes(s), and a central unit (CU) as the central node 160.
  • the base station 170 might include an eNB having a RRH (remote radio head) as remote node 150 and a base band unit (BBU) as a central node.
  • the remote node(s) 150 are coupled to a central node 160 via one or more links
  • the remote nodes 150 may be remote in the sense they are contained in different physical enclosures from a physical enclosure containing a corresponding central node 160.
  • the link(s) 171 may be implemented using fiber optics, wireless techniques, or any other technique for data communications.
  • Two or more base stations 170 communicate using, e.g., link(s) 176.
  • the link(s) 176 may be wired or wireless or both and may implement, e.g., an Xn interface for 5G, an X2 interface for LTE, or other suitable interface for other standards.
  • the wireless network 100 may include a network element or elements 190, as a third illustrated node, that may include core network functionality, and which provides connectivity via a link or links 181 with a data network 191, such as a telephone network and/or a data communications network (e.g., the Internet).
  • a data network 191 such as a telephone network and/or a data communications network (e.g., the Internet).
  • core network functionality for 5G may include access and mobility management function(s) (AMF(s)) and/or user plane functions (UPF(s)) and/or session management function(s) (SMF(s)).
  • AMF(s) access and mobility management function(s)
  • UPF(s) user plane functions
  • SMF(s) session management function
  • Such core network functionality for LTE may include MME (Mobility Management Entity) functionality and/or SGW (Serving Gateway) functionality.
  • MME Mobility Management Entity
  • SGW Serving Gateway
  • the RAN node 170 is coupled via a link 131 to a network element 190.
  • the link 131 may be implemented as, e.g., an NG interface for 5G, or an SI interface for LTE, or other suitable interface for other standards.
  • sensing 106 may be performed by a single base station 170. It is also possible for sensing 106 from multiple base-stations (e.g., BSs 170 and 170-1 or additional BSs) can be combined, e.g., for improved accuracy. In general, the basestation is a place where data is collected.
  • Object detection 107 may happen directly at base-station 170 (or other BSs 170-1 or more), at other network elements 190, and/or in the cloud (e.g., at remote server 108, in a data network 191 that is or has a cloud).
  • the remote server 108 is another possible node in FIG. 1A.
  • Output of the object detection may be used by any algorithm/solution that would use this information, for example to improve network performance by impacting the network scheduler, as an input to a digital twin (i.e., a digital representation of the real world), to assist street lights steering, and many more.
  • a digital twin i.e., a digital representation of the real world
  • the use of the output of object detection is shown occurring (see reference 107) in the BS 170, and/or the network element(s) 190, and/or the remote serverl08, although other locations are possible (e.g., such as a remote server 108 outside the data network 191).
  • This example combines object detection and use (via reference 107), but object detection could be completely separate from use (e.g., perform on two different physical entities such as servers).
  • FIG. 1A shows an offline computer 173, as another possible node, that may be used in certain embodiments.
  • the offline computer 173 in this example is able to communicate through the data network 191 with the wireless network 100.
  • the offline computer 173 may perform some of the operations described herein, as detailed below.
  • FIG. IB illustrates an example apparatus 180 suitable for implementing any of the nodes in FIG. 1A.
  • the apparatus 180 includes circuitry comprising one or more processors 120, one or more memories 125, one or more transceivers 130, one or more network interface(s) 155 and user interface (UI) circuitry and elements 157, interconnected through one or more buses 127. Since this is an example covering all of the nodes in FIG. 1A, some of the nodes may not have all of the circuitry. For example, a base station 170 might not have UI circuitry and elements 157. All of the nodes may have additional circuitry, not described here.
  • FIG. IB is presented merely as an example.
  • Each of the one or more transceivers 130 includes a receiver, Rx, 132 and a transmitter, Tx, 133.
  • the one or more buses 127 may be address, data, and/or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like.
  • the one or more transceivers 130 are connected to one or more antennas 105, which may be one of the antennas 128 (from UE 110) or antennas 158 (from base station 170).
  • the one or more memories 125 include computer program code 123.
  • the apparatus 180 includes a control module 140, comprising one of or both parts 140-1 and/or 140-2, which may be implemented in a number of ways.
  • the control module 140 may be implemented in hardware as control module 140-1, such as being implemented as part of the one or more processors 120.
  • the control module 140-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array.
  • the control module 140 may be implemented as control module 140-2, which is implemented as computer program code (having corresponding instructions) 123 and is executed by the one or more processors 120.
  • the one or more memories 125 store instructions that, when executed by the one or more processors 120, cause the apparatus 180 to perform one or more of the operations as described herein.
  • the one or more processors 120, one or more memories 125, and example algorithms (e.g., as flowcharts and/or signaling diagrams), encoded as instructions, programs, or code, are means for causing performance of the operations described herein.
  • the network interface(s) 155 are wired interfaces communicating using link(s) 156, which may be fiber optic or other wired interfaces.
  • the link(s) 156 may be the link(s) 131 and/or 176 from FIG. 1A.
  • the link(s) 131 and/or 176 from FIG. 1A may also be implemented using transceiver(s) 130 and corresponding wireless link(s) 111.
  • the apparatus 180 may include only wireless transceiver(s) 130, only network interface(s) 155, or both wireless transceiver(s) 130 and network interface/ s) 155.
  • the apparatus 180 may or may not include UI circuitry and elements 157. These may include a display such as a touchscreen, speakers, or interface elements such as for headsets. For instance, a UE 110 of a smartphone would typically include at least a touchscreen and speakers.
  • the UI circuitry and elements 157 may also include circuity to communicate with external UI elements (not shown) such as displays, keyboards, mice, headsets, and the like.
  • the computer readable memories 125 may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, firmware, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory.
  • the computer readable memories 125 may be means for performing storage functions.
  • the processors 120 may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on a multicore processor architecture, as non-limiting examples.
  • the processors 120 may be means for performing functions, such as controlling the apparatus 180, and other functions as described herein.
  • the example embodiments herein relate at least to beyond 5G and 6G (sixth generation) technologies, where it is envisioned to enable joint communications and sensing (JCAS) as a new use case of wireless communications systems. See, e.g., H. Viswanathan and P. E. Mogensen, "Communications in the 6G Era,” in IEEE Access, vol. 8, pp. 57063-57074, 2020, doi: 10.1109/ ACCESS.2020.2981745.
  • JCAS joint communications and sensing
  • a digital twin which is a digital representation of the real world, possibly in real-time or close to real-time.
  • the transmitted radar signal is reflected by the target object.
  • the receiver end can utilize the reflected signal and, through processing, compute the physical attributes of the target object including distance, direction, velocity, and the like.
  • FIG. 2 illustrates using NR (new radio) and/or 6G access points as radar points that can help detect objects in the desired area of sensing.
  • access points 210-1 and 210-2 are utilized as radar points to detect a car 220 crossing a traffic intersection 230.
  • Techniques herein relate to improving the object detection, classification and localization abilities of an, e.g., OFDM (orthogonal frequency division multiplexing)- based cellular JCAS system, such as the network 100 of FIG. 1.
  • OFDM orthogonal frequency division multiplexing
  • the techniques are applicable to any radar or single- or multi-sensory entity that potentially has a distorted and noisy observation of the physical environment, including Lidar (light detection and ranging), backscatter tracking, and the like. That is, the apparatus 180, and some or all of FIG.l, can apply to radar or multi-sensory entities.
  • FIG. 3 shows a block diagram of radar sensing as a complicated function of the environment.
  • Block 310 indicates there is a point cloud, or a 3D or 2D image or map of the sensing environment.
  • Block 320 has the “ideal” radar sensing as the following: one-degree beams in every direction at every time instance; and minimal hardware impact.
  • f 2 (x) a non-ideal radar 6G JCAS (impacted by lower resolution and higher impact from waveform and hardware), is denoted by f 2 (x) as a function of the ideal radar observation.
  • JCAS is impacted by the following: fat beams, one at a time slot; hardware impairment; and poor sampling over time and frequency.
  • the object detection tools are best suited for the point cloud, or 2D/3D images of the environment (LHS, left hand side, block 310 in FIG. 3).
  • the prior knowledge is ‘learned’ by a generative adversarial network (GAN) and is referred to as ‘generative priors’ which is normally not feasible to model analytically, but can be learned by the neural network using a large dataset of observations from the scene of interest.
  • GAN generative adversarial network
  • An E2E neural network design approach is then proposed that takes radar heatmap, converts the heatmap to a 3D point cloud of the scene of interest, and detects objects in the scene.
  • the radar input in examples herein is different from the typical radar input, e.g., in Guan et al., in the sense that the radar inputs has the impact of JCAS signal processing and hardware considered in the signal. That is, the radar input is impacted by signal processing and hardware effects.
  • Guan et al. uses a 2D depth map of the scene as ground truth for training and inference.
  • a 2D map can be used
  • a focus herein is on a 3D point cloud of the environment, which includes all the objects in the scene as well as the static background. Therefore, in the inference stage, example networks herein regenerate the complete scene, including static background, non-target dynamic and static objects, as well as the target objects.
  • an example network herein is capable of detecting multiple ones of the same object in one scene, or differentiating one object of interest from a non-target object in the same scene. This does not happen in Guan et al.
  • FIG. 4 illustrates an example method in accordance with the present disclosure.
  • Blocks 400, 402, and 404 are assumed to be performed offline, that is, not in the wireless network 100 and instead performed by offline computer 173. This is not a limitation, however, and the blocks could be performed by a node in the wireless network 100.
  • Block 406 is assumed to be performed by a node in the wireless network 100.
  • the offline computer 173 rains a generative adversarial network (GAN) for reconstruction of 2D or 3D images of a physical area of interest, using wireless radar sensing of the area. In block 400, this is more likely to use mainly synthetic data, since a large number of data samples are needed. However, in case of availability of data, real world measurements could be used too.
  • the offline computer 173 trains an object detection network to identify a class of objects (e.g., cars, pedestrians, bikes, trucks, and the like) using, e.g., 2D or 3D images of the area of interest. In an example, object detection performed by the offline computer 173 may include classifying the detected objects.
  • the offline computer 173 performs end-to-end (E2E) fine-tuning of a cascade network of the GAN from block 400 and the detection network from block 402.
  • E2E end-to-end
  • block 406 learning is transferred from the cascade network of block 404 using real-word measurement data.
  • block 406 transfers learning into the real-world domain, using real world measurement data collected by the base station 170 (or base stations).
  • the actual algorithm for block 406 may be run in base station 170, if necessary computational power is available, or by other network elements 190, or in the cloud 191 via, e.g., the remote server 108.
  • Real-world measurement data is impacted by hardware imperfections and other effects that are hard or impossible to capture in simulation. In block 406, this is more likely to use real measurements. However, using synthetic data is also applicable.
  • a generator NN is trained, e.g., using a generative adversarial network (GAN) architecture, to reconstruct images from the scene of interest based on a radar observation of the scene.
  • GAN generative adversarial network
  • the training may be targeted to reconstruct high-level visual representation of the scene, including the objects that are intended for detection and corresponding classification.
  • the generator and discriminator NNs may be trained jointly using ground truth data (e.g., a 3D image of the area of interest) and input data (e.g., a 3D radar heatmap of the area of interest).
  • Training data may include data samples that contain one or multiple of the objects of interest in the classifications; training data can be created synthetically using, e.g., ray-tracing based radar simulation tools.
  • input of the generator is radar input data (e.g., a 3D radar heatmap of the area of the interest).
  • the desired output of the generator is the equivalent 3D image of the area of interest.
  • the GAN is trained using multiple radar inputs, e.g., from multiple mono-static or bi-static radars, where the generator is trained to fuse the multiple input sources together to generate the physical reconstruction of the environment.
  • the GAN is trained using a memory-enabled network, e.g., LSTM (long short-term memory), where multiple consecutive radar frames from the same scene can be fed into the generator.
  • the GAN is trained using a fusion of multiple sensory information, e.g., radar + Lidar + localization information from other sources, including GPS (global positioning system), GNSS (global navigation satellite system), and the like.
  • an object detection (which may include classification) network is trained to detect (and classify) objects of interest in the sensing area of interest.
  • input data includes a 3D image of the area of interest, while ground truth includes presence/absence of object(s), and likelihood of a detected object belonging to one of the defined classed of objects.
  • the GAN reconstructed image of the area of interest is fed into the object detection network.
  • Block 404 a cascade network of the generator output at block 400 and the object detection output at block 402 is formed; the cascade network is fine-tuned end to end to improve the object detection.
  • Block 404 also may involve the following: fine-tuning and transfer learning with E2E training; most layers are frozen while only updating the few last layers; and calibrating the cascade network using a small data set.
  • this involves transfer learning of the E2E training from block 404; similarly to block 404, a few of the last layers of the E2E network are fine-tuned using real- world measured data, e.g., from the scene where the proposed application is to be deployed in.
  • transfer learning also referred to as learning transfer
  • NN convolutional neural network
  • GAN convolutional neural network
  • the network can be “pre-trained”, for example using synthetic data until achieving satisfactory accuracy. This is what is happening in blocks 400, 402, and 404. This pre-trained network performs well but can do better in specific environment. That’s where the second stage of training comes (the transfer learning of block 406), where pre-trained network is being trained again (but not fully, only last few layers) using data collected in the environment. Because only a few layers are being trained (and not from the scratch, but using weights from initial training), the dataset required to achieve a high accuracy is much smaller.
  • the example embodiments may include lower complexity since the training may be performed part-by-part.
  • the example embodiments may provide an improved detection rate as compared to systems that use processed radar data.
  • the example embodiments may provide an improved occupancy grid reconstruction from multiple radar and other sensory information.
  • Th example embodiments may provide an easier fusion of multiple sources of sensory data, e.g., multiple mono-static or bi-static radar.
  • the example embodiments may provide efficient usage of static background information (through developing generative priors), to improve efficiency of radar sensing towards object detection.
  • FIG. 5 illustrates an example approach in Al-based object detection from radar sensing.
  • the approach 500 generally includes training a neural network to detect objects directly from a raw radar image, or from a pre-processed radar image.
  • the approach 500 performs radar processing 510 on raw radar information (info) 505.
  • the radar signal processing includes expert feature extraction, MUSIC (multiple signal classification), and the like.
  • the output is processed radar information 520, and a radar object detection/classification model 525 performs detection and corresponding classification to output objects and classes 535.
  • the block 525 uses model-based algorithms and/or data-based machine learning detection schemes, designed for radar sensing.
  • the advanced Al-based object detection techniques described in reference to FIG. 5 may be developed for visual images and, therefore, may not be readily useful in object detection.
  • the neural network in block 525) may require large datasets to be trained.
  • FIG. 6A illustrates an example framework 600 for Al-based object detection.
  • the framework 600 is configured to receive raw radar information 505.
  • the framework 600 is configured to perform radar inversion 610 on the raw radar information.
  • a radar inversion module e.g., using GANs
  • the radar inversion module generates images 620 from the scene.
  • the images 620 may be processed in block 625.
  • the framework 600 may use (optional) image processing techniques including, e.g., expert feature extraction, transform domains, and the like.
  • the image processing at block 625 outputs processed images 635 to an (image) object detection/classification module 640 for further processing.
  • the object detection/classification module 640 performs model-based algorithms and/or data-based machine learning detection schemes, e.g., bounding box detection, image segmentation, and the like.
  • the object detection/classification module 640 generates one or more outputs, such as, one or more objects and associated one or more classes of the one or more objects.
  • FIGS. 7, 8A, 8B, and 9 related to training for forming a NN able to perform the radar inversion 610
  • FIGS. 10 and 11 illustrate how images 620 could be formed
  • FIGS. 12 and 13 illustrate examples of processing 635 and object detection/classification module 640
  • FIG. 14 illustrates an example of end-to-end processing.
  • FIG. 6B is a diagram illustrating another example framework for Al-based object detection.
  • raw sensory information 505-1 instead of (e.g., only) raw radar information 505.
  • sensory inversion 610-1 that operates on the raw sensory information 505-1 to produce images 620-1 from the scene.
  • the sensory inversion module 610-1 610 e.g., using GAN(s) exploits data-driven priors to reconstruct the scene from the sensory information.
  • the raw sensory information 505-1 is formed by a single sensory entity with one of the following or a multi- sensory entity with multiple ones of the following types of sensory information: Radar, Lidar, backscatter tracking, localization information from other sources, including GPS, GNSS, and the like. Training would also have to account for the selected single type of sensory information or multiple types of sensory information.
  • FIG. 7 illustrates an example training of a generator network for radar inversion, which could be implemented in block 400 of FIG. 4.
  • a GAN architecture 700 has a 3D radar heat map 710 that is sent to a generator model 720, which produces a generated depth map 725.
  • a heat map is a data visualization technique that shows magnitude of a phenomenon using, e.g., color, in two or three dimensions, and a radar heat map is a heat map formed by radar.
  • the depth map 725 and a real scene depth map 715 are input to a discriminator model 730, which provides an output 745 to a classification module 740 that outputs fake/real outputs. These outputs are sent via update 750 to the generator model 720 and via update 755 to the discriminator model 730.
  • a generator neural network (NN) 720 is trained e.g., using a generative adversarial network (GAN) architecture 700, to reconstruct images from the scene of interest from a radar observation of the scene.
  • GAN generative adversarial network
  • Radar heatmap 710 Normalized received power may be used for each point of the coordinate system of interest.
  • the heatmap 710 may be created per beam and then combined together into one radar heatmap using weighted summation of the heatmap 710 per beam.
  • the weights are designed, e.g., depending on the segment of the coordinate system that each beam covers, where weight 1 (one) may be given to the covered area and weight 0 (zero) to the non-covered parts.
  • the gain of the beam at each coordinate point can be chosen as the soft weighting parameter.
  • 3D depth map (ground truth for training) 715 a real scene depth map may be generated by assigning 0 (zero) for every point of the coordinate system where there is free space, and 1 (one) to the rest of the points where objects are present. This is an example of an occupancy grid.
  • FIG. 8A which illustrates a 2D heat map
  • FIG. 8B which illustrates a 3D binary depth map.
  • Each point of the 2D map corresponds to a vertical and azimuth angle from the point of view of the measurement device (e.g., Radar).
  • the value of each point may correspond to the depth (e.g., distance) of the first observable object in the direction of the angles, and the value may further contain information about material, color, conductivity, and the like, although these are not needed.
  • the points of the map are multi-layer, where one layer is the binary map described above, and one layer includes the physical characteristics of the object (e.g., conductivity of the surface or material index out of a preassigned set of material indexes).
  • an occupancy grid usually refers to a grid-based matrix, where each entry corresponds to one point in a 3D grid of interest.
  • the value in each entry can be binary (e.g., object present or not) or non-binary, corresponding to classes of objects or to material and color characteristics of the objects, or the like.
  • a 3D point cloud usually refers to a fairly similar definition, where instead of a grid-based matrix, the data format includes all the points from the 3D space where the observer has collected “information” about them. This means the coordinate of the point is corresponded to the value of the measurement.
  • the discriminator 730 may look at one specific realization of reconstruction at the time and comparing this with the real data.
  • the training process for the GAN architecture 700 could be performed for individual ones of multiple real scene depth maps 715.
  • the GAN architecture 700 can be trained for all the multitude of depth maps at the same time.
  • the expected output 725 of the trained generator NN 720 is a high-level visual representation of the scene, which resembles the 3D ground truth depth map (as the real scene depth map 715).
  • training data may include data samples that contain one or multiple of the objects of interest in the classification.
  • Training data can be created synthetically using, e.g., ray-tracing based radar simulation tools, or using measurements of a radar system operation.
  • the training is targeted to reconstruct high-level visual representation of the scene, including objects that are intended for detection and classification.
  • the generator NN 720 and discriminator NN 730 are trained jointly, where generator creates 3D depth maps 725 from the 3D radar heatmaps 710, while the discriminator tries to distinguish the generated output 725 from the ground truth (the real scene depth map 715).
  • a loss function may be used to measure the loss between the ground truth and generated map.
  • the loss function in FIG. 9 can be used for the GAN training, where 1 values represent the weight of each loss term in the overall loss function.
  • the training may be performed using ground truth data (e.g., 2D or 3D image of the area of interest) and input data (2D or 3D radar heatmap of the area of interest).
  • input of the generator NN 720 is radar input data (e.g., 2D or 3D radar heatmap 710 of the area of the interest), while the desired output 725 of the generator NN 720 is the equivalent 2D or 3D image of the area of interest.
  • radar input data e.g., 2D or 3D radar heatmap 710 of the area of the interest
  • desired output 725 of the generator NN 720 is the equivalent 2D or 3D image of the area of interest.
  • the GAN 700 is trained using multiple radar inputs, e.g., from multiple mono-static or bi-static radars.
  • the multiple inputs may be fed, such that, in one example, each mono static radar or bi-static radar input is fed in as a separate input layer and, in another example, a fusion function is defined to fuse the multiple input sources together, e.g., for each point of the coordinate system of interest, the measurements of all radar input layers are combined with weighted combining.
  • the GAN 700 is trained using a memory-enabled neural network, e.g., LSTM, where multiple consecutive frames of radar heatmap from the same scene are fed into the generator NN 720.
  • the output is the average high-level representation of the physical scene from the multiple frames.
  • the GAN 700 is trained using multiple sensory information, e.g., a radar heatmap, a Lidar heatmap, localization information collected at the 5G LMF (location management function), GPS, GNSS, and the like, and these are fused together and fed into the generator NN 720, or fed in as separate layers to the generator NN 720.
  • multiple sensory information e.g., a radar heatmap, a Lidar heatmap, localization information collected at the 5G LMF (location management function), GPS, GNSS, and the like.
  • Example expected outputs include the following.
  • An example in FIG. 10 demonstrates the case of 3D depth map reconstruction.
  • the radar heat map 710-1 is an input, as is the ground truth 715-1, which are both processed by a generator NN 720 to create an output 725-1.
  • the output 725-1 generated by the trained generator NN 720 resembles the ground truth 715-1 much better visually, compared to the radar heatmap 710-1, which is difficult to characterize.
  • the plots in FIG. 11 show the 2D version of the example in FIG. 10.
  • a 2D version of the ground truth 715-1 and a 2D version of the output 725-1 generated by the network are shown.
  • the depth map is a 2D matrix containing the value of the depth distance to the first object from the point of view of the radar, as described with reference to FIG. 8.
  • Training of an object detection network is now described. This is block 402 of FIG. 4, and may be used to train (image) object detection/classification module 640 of FIG. 6A. An object detection/classification neural network is trained to detect objects of interest in the sensing area of interest.
  • the input includes the 3D or 2D map of the environment, created by the trained generator NN 720, as such training was described above.
  • the trained generator NN 720 is now performing the radar inversion 610 of FIG. 6A, and produces images 620 that are input to the (image) object detection/classification module 640.
  • the input can be a fusion of the map from the generator together with other sources of information, including, e.g., Lidar heatmap and/or localization information.
  • the map from the generator NN 720 can be accompanied by an additional input layer which includes the static background of the scene of interest. This enables the network to distinguish between objects and background. Another option is to subtract the background information from the input map. [00108] Two different example options are assumed for the output, although the techniques herein are not limited to these. These two options are described below.
  • Option 1 the network is trained to estimate likelihood of presence of an object from a set of pre-defined object categories (which may be thought of as classes for this example).
  • object categories which may be thought of as classes for this example.
  • 10 likelihoods may be output, one for each category.
  • FIG. 12 which is an example illustrating an object detection neural network. Reference may also be made to FIG. 6A, as the reference numbers from there are also used in FIG. 12.
  • the input is a 3D point cloud 620-1 with the background subtracted. Background may be subtracted using multiple techniques. For instance, a 3D point cloud could be produced, then the background would be subtracted.
  • the subtraction may be performed at the at the same time of the radar inversion, so the output has the 3D point cloud with background subtracted.
  • the object detection neural network 640-1 uses the grid 635-1 to output a likelihood 650-1 of presence of an object that belongs to (e.g., known) object categories.
  • Option 2 The network comprises a two-step detection.
  • a neural network is trained to detect bounding boxes for objects in the scene.
  • a separate network is then trained (step b) to estimate likelihood of objects in each bounding box separately. This is especially useful for scenes with multiple objects.
  • Option 2 is described through reference to FIG. 13, which is an example illustrating a two-step bounding box and object detection. There are two steps: step a, 1305-1; and step b, 1305-2, which is repeated for all the detected bounding boxes.
  • step a, 1305-1 a 3D point cloud 620-1 with background subtracted is input to a first object detection/classification NN 640-2, which outputs bounding box detection 650-2 for objects in the scene.
  • Step b, 1305-2 uses an input of one bounding box segment 1320 of the 3D point cloud, and step b may be repeated for all the detected bounding boxes.
  • a second object detection/classification NN 640-3 operates on each bounding box segment to produce a likelihood 620-1 of presence of an object belonging to (e.g., known) object categories within the corresponding box segment.
  • training is performed using a 2D or a 3D point cloud map of the area of interest instead of 2D or 3D maps created by the generator NN 720.
  • ground truth for Option 1 may include a vector of binary values for each category of objects (e.g., 0, zero, if object is not present and 1, one, if object is present).
  • the ground truth in step a, 1305-1 may be the exact bounding box of the training set, which for step b, 1305-2, the ground truth includes the binary values for an object bounding box for each category of objects.
  • the GAN reconstructed image of the area of interest is fed into the object detection network.
  • FIG. 14 is an example illustrating an E2E network for object detection from a radar heat map.
  • a radar heat map 505-1 is input to a generator network 610-1 (e.g., a trained generator network 720- 1), which produces a reconstructed 3D point cloud 620-2.
  • the cloud 620-2 is input to the object detection network 640-1, which outputs a likelihood 650-1 of presence of an object that belongs to (e.g., known) object categories.
  • the cascade network is fine-tuned (see block 404 of FIG. 4) end-to-end to improve the object detection.
  • a small set of layers in the NN are selected as tunable layers 1410.
  • This example shows a layer 1410-1 in the generator network 610-1 and a layer 1410-2 in the object detection network 640-1. It is noted that multiple layers may be used, and the networks 610-1 and 640-1 need not have the same number of layers implemented as tunable layers.
  • the tuning is typically performed via tuning the weights that are multiplied by the feedforward values exchanged between layers of a NN.
  • the remaining layers are frozen while only updating the few last tunable layers. This way, one can calibrate the cascaded E2E network using a small data set.
  • a trained NN is used with real-world measurement data, and only a few of the layers of the trained NN are updated. This means a small data set - as compared to the data set used for initially training the NN - is used.
  • the input data may be the input as described above (e.g., radar heatmap and 3D depth map (ground truth for training)), and the output data may be the output from Options 1 or 2 described above. (In case of Option 2, either the step a network is only cascaded with the generator, or the cascaded network has an adaptive number of equivalent step b networks in parallel at the output of step a network.)
  • a few of the last layers of the E2E network may be fine-tuned using real- world measured data defined the same way as the input and output above, but collected, e.g., from the scene in which the proposed object detection application is to be deployed.
  • FIG. 14 may apply to both blocks 404 (e.g., E2E fine- tuning of a cascade network) and 406 (e.g., learning transfer of the system).
  • the data can be also real, e.g., training the network with synthetic data of the scene and fine-tuning the cascade with real data from the same scene.
  • transfer learning (as in block 406) can be performed with synthetic data too. That is, training the network for a certain scene may be performed, and then the learning of the cascade network is transferred to a different scene by fine tuning the learning using, e.g., a set of synthetic data from the new scene.
  • a technical effect and advantage of one or more of the example embodiments disclosed herein is overcoming hardware impairments and other communication network related imperfection. Another technical effect and advantage of one or more of the example embodiments disclosed herein is improving object detection and classification accuracy. Another technical effect and advantage of one or more of the example embodiments disclosed herein is improved performance of any other algorithm that would benefit from higher resolution of radar image.
  • Example 1 A method, comprising:
  • Example 2 The method according to example 1, wherein the performing at least object detection also performs classification of detected objects so the output indicates whether or not one or more objects were detected using likelihood of presence of the one or more objects belonging to an object category.
  • Example 3 The method according to example 2, wherein there are multiple categories, and classification outputs multiple likelihoods corresponding to the multiple categories.
  • Example 4 The method according to any one of examples 2 or 3, wherein the images comprise three-dimensional point clouds where backgrounds have been subtracted.
  • Example 5 The method according to example 4, further comprising converting the three-dimensional point clouds where backgrounds have been subtracted to three-dimensional occupancy grids, wherein the performing at least object detection performs at least object detection on the three-dimensional occupancy grids.
  • Example 6 The method according to any one of examples 2 or 3, wherein the one or more images comprise three-dimensional point clouds with background subtracted and wherein the performing, using the one or more other neural networks, at least object detection on the one or more images to determine an output comprises the following:
  • Example 7 The method according to any one of examples 1 to 6, wherein the converting sensory information and performing at least object detection are performed to train the neural network and the one of more other neural networks using the sensory information to adjust one or more tunable layers in one or both of the neural network or the one or more other neural networks.
  • Example 8 The method according to example any one of examples 1 to 7, wherein the sensory information comprises radar information developed from one or more base stations in a wireless system.
  • Example 9 A method, comprising:
  • training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area, the training comprising:
  • Example 10 The method according to example 9, wherein the real scene depth map comprises an occupancy grid, or one or multiple two-dimensional depth maps, or both the occupancy grid and the one or multiple two-dimensional depth maps.
  • Example 11 The method according to example 9, wherein the sensory information comprises radar information, wherein the real scene depth map comprises a two- dimensional map based on azimuth and elevation, where points in the map correspond to vertical and azimuth angles from a point of view of a radar measurement device generating the radar information, and values of the points correspond to distance of a first observable object in a direction of the angles corresponding to the points.
  • the sensory information comprises radar information
  • the real scene depth map comprises a two- dimensional map based on azimuth and elevation, where points in the map correspond to vertical and azimuth angles from a point of view of a radar measurement device generating the radar information, and values of the points correspond to distance of a first observable object in a direction of the angles corresponding to the points.
  • Example 12 The method according to example 9, wherein the real scene depth map comprises a three-dimensional map generated at least by assigning a first value to points of a coordinate system, corresponding to the physical area, where there is free space, and assigning a second value to points where objects are present in the coordinate system, where a function is used to determine whether an object does or does not exist at individual directions based on azimuth, elevation, and range.
  • Example 13 The method according to example 9, wherein the generative adversarial network is trained using sensory information comprising one or more of the following: radar information, Lidar information, backscatter tracking information, or localization information.
  • Example 14 The method according to example 13, wherein the generative adversarial network is trained using sensory information comprising multiple ones of the following: the radar information, the Lidar information, the backscatter tracking information, or the localization information, and the multiple sensory information are fused together and fed into the generator neural network, or fed in as separate layers to the generator neural network.
  • Example 15 The method according to any one of examples 9 to 14, further comprising training an object detection neural network on identifying classes of objects using images from the physical area.
  • Example 16 The method according to example 15, further comprising:
  • Example 17 The method according to example 16, wherein the performing training the cascade network creates a trained cascade network, and the method further comprises performing another training of the cascade network using measurement data from one or more sensing entities of the physical area, the other training performed to adjust one or more tunable layers in one or both of the trained generator neural network or the object detection neural network.
  • Example 18 A computer program, comprising instructions for performing the methods of any of examples 1 to 17, when the computer program is run on an apparatus.
  • Example 19 The computer program according to example 18, wherein the computer program is a computer program product comprising a computer-readable medium bearing instructions embodied therein for use with the apparatus.
  • Example 20 The computer program according to example 18, wherein the computer program is directly loadable into an internal memory of the apparatus.
  • Example 21 An apparatus, comprising means for performing:
  • Example 22 The apparatus according to example 18, wherein the performing at least object detection also performs classification of detected objects so the output indicates whether or not one or more objects were detected using likelihood of presence of the one or more objects belonging to an object category.
  • Example 23 The apparatus according to example 19, wherein there are multiple categories, and classification outputs multiple likelihoods corresponding to the multiple categories.
  • Example 24 The apparatus according to any one of examples 19 or 20, wherein the images comprise three-dimensional point clouds where backgrounds have been subtracted.
  • Example 25 The apparatus according to example 21, wherein the means are further configured for performing converting the three-dimensional point clouds where backgrounds have been subtracted to three-dimensional occupancy grids, wherein the performing at least object detection performs at least object detection on the three- dimensional occupancy grids.
  • Example 26 The apparatus according to any one of examples 19 or 20, wherein the one or more images comprise three-dimensional point clouds with background subtracted and wherein the performing, using the one or more other neural networks, at least object detection on the one or more images to determine an output comprises the following: [00163] performing a first detection using a first of the one or more other neural networks to produce first output comprising boundary box detection for objects in a scene, wherein there are multiple bounding box segments for the 3D point clouds; and
  • Example 27 The apparatus according to any one of examples 18 to 23, wherein the converting sensory information and performing at least object detection are performed to train the neural network and the one of more other neural networks using the sensory information to adjust one or more tunable layers in one or both of the neural network or the one or more other neural networks.
  • Example 28 The apparatus according to example any one of examples 18 to 24, wherein the sensory information comprises radar information developed from one or more base stations in a wireless system.
  • Example 29 An apparatus, comprising means for performing:
  • training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area, the training comprising:
  • [00170] determining classifications of fake or real by a discriminator neural network based on the generated depth maps and one or more real scene depth maps, the one or more real scene depth maps comprising a multi-dimensional map of the physical area, where the classifications are routed to the generator neural network and the discriminator neural network;
  • Example 30 The apparatus according to example 26, wherein the real scene depth map comprises an occupancy grid, or one or multiple two-dimensional depth maps, or both the occupancy grid and the one or multiple two-dimensional depth maps.
  • Example 31 The apparatus according to example 26, wherein the sensory information comprises radar information, wherein the real scene depth map comprises a two- dimensional map based on azimuth and elevation, where points in the map correspond to vertical and azimuth angles from a point of view of a radar measurement device generating the radar information, and values of the points correspond to distance of a first observable object in a direction of the angles corresponding to the points.
  • the sensory information comprises radar information
  • the real scene depth map comprises a two- dimensional map based on azimuth and elevation, where points in the map correspond to vertical and azimuth angles from a point of view of a radar measurement device generating the radar information, and values of the points correspond to distance of a first observable object in a direction of the angles corresponding to the points.
  • Example 32 The apparatus according to example 26, wherein the real scene depth map comprises a three-dimensional map generated at least by assigning a first value to points of a coordinate system, corresponding to the physical area, where there is free space, and assigning a second value to points where objects are present in the coordinate system, where a function is used to determine whether an object does or does not exist at individual directions based on azimuth, elevation, and range.
  • Example 33 The apparatus according to example 26, wherein the generative adversarial network is trained using sensory information comprising one or more of the following: radar information, Lidar information, backscatter tracking information, or localization information.
  • Example 34 The apparatus according to example 30, wherein the generative adversarial network is trained using sensory information comprising multiple ones of the following: the radar information, the Lidar information, the backscatter tracking information, or the localization information, and the multiple sensory information are fused together and fed into the generator neural network, or fed in as separate layers to the generator neural network.
  • the generative adversarial network is trained using sensory information comprising multiple ones of the following: the radar information, the Lidar information, the backscatter tracking information, or the localization information, and the multiple sensory information are fused together and fed into the generator neural network, or fed in as separate layers to the generator neural network.
  • Example 35 The apparatus according to any one of examples 26 to 31, wherein the means are further configured for performing training an object detection neural network on identifying classes of objects using images from the physical area.
  • Example 36 The apparatus according to example 32, wherein the means are further configured for performing:
  • Example 37 The apparatus according to example 33, wherein the performing training the cascade network creates a trained cascade network, and wherein the means are further configured for performing another training of the cascade network using measurement data from one or more sensing entities of the physical area, the other training performed to adjust one or more tunable layers in one or both of the trained generator neural network or the object detection neural network.
  • Example 38 The apparatus of any preceding apparatus example, wherein the means comprises:
  • At least one processor at least one processor
  • At least one memory storing instructions that, when executed by at least one processor, cause the performance of the apparatus.
  • circuitry may refer to one or more or all of the following:
  • software e.g., firmware
  • circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and/or firmware.
  • circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
  • Embodiments herein may be implemented in software (executed by one or more processors), hardware (e.g., an application specific integrated circuit), or a combination of software and hardware.
  • the software e.g., application logic, an instruction set
  • a “computer-readable medium” may be any media or means that can contain, store, communicate, propagate or transport the instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer, with one example of a computer described and depicted, e.g., in FIG. 1A.
  • a computer-readable medium may comprise a computer-readable storage medium (e.g., memories 125 or other device) that may be any media or means that can contain, store, and/or transport the instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer.
  • a computer-readable storage medium does not comprise propagating signals, and therefore may be considered to be non-transitory.
  • the term “non-transitory”, as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM, random access memory, versus ROM, read-only memory).

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Evolutionary Computation (AREA)
  • General Health & Medical Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Biophysics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Data Mining & Analysis (AREA)
  • General Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Mathematical Physics (AREA)
  • Biomedical Technology (AREA)
  • Molecular Biology (AREA)
  • Multimedia (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Databases & Information Systems (AREA)
  • Medical Informatics (AREA)
  • Probability & Statistics with Applications (AREA)
  • Image Analysis (AREA)

Abstract

A NN converts sensory information of a physical area into image(s). One or more other NNs perform at least object detection on the image(s) to determine an output indicating at least whether or not object(s) were detected. A GAN is trained for reconstruction of images of a physical area by creating generated depth maps by a generator NN based at least on corresponding sensory information of a physical area, the generated depth maps including images of the physical area. Classifications of fake or real are determined by a discriminator NN based on the generated depth maps and real scene depth map(s). The creating and classification are performed until criteria are met. Information is output that defines a trained generator NN able to reconstruct images of the physical area based at least on corresponding sensory information of the physical area.

Description

Object Detection Accuracy Improvement in Wireless Communications Systems at Least using Radar Sensing
TECHNICAL FIELD
[0001] Exemplary embodiments herein relate generally to wireless communications systems and, more specifically, relates to object detection using wireless communications systems as radar sensing elements.
BACKGROUND
[0002] Recently, there has been interest in using wireless communications networks to perform sensing functions. In particular, spatial sensing is possible using the wireless communications networks. One usage of sensing capabilities is referred to a joint communication and sensing (JCAS).
[0003] JCAS is being examined for sensing because a wireless communications system has most or all of the infrastructure is in place, e.g., with transmit/receive (Tx/Rx) nodes with a full area coverage as well as a good interconnection between nodes. This allows the wireless communications system to perform radar sensing of objects that are detected by the system.
[0004] As with many radar systems, accuracy can be improved when using a wireless communications system for radar sensing.
BRIEF SUMMARY
[0005] This section is intended to include examples and is not intended to be limiting.
[0006] In an exemplary embodiment, a method is disclosed that includes converting, using a neural network, sensory information of a physical area into one or more images. The method includes performing, using one or more other neural networks, at least object detection on the one or more images to determine an output indicating at least whether or not one or more objects were detected.
[0007] An additional exemplary embodiment includes a computer program, comprising instructions for performing the method of the previous paragraph, when the computer program is run on an apparatus. The computer program according to this paragraph, wherein the computer program is a computer program product comprising a computer-readable medium bearing the instructions embodied therein for use with the apparatus. Another example is the computer program according to this paragraph, wherein the program is directly loadable into an internal memory of the apparatus.
[0008] An exemplary apparatus includes one or more processors and one or more memories storing instructions that, when executed by the one or more processors, cause the apparatus at least to perform: converting, using a neural network, sensory information of a physical area into one or more images; and performing, using one or more other neural networks, at least object detection on the one or more images to determine an output indicating at least whether or not one or more objects were detected.
[0009] An exemplary computer program product includes a computer-readable storage medium bearing instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: converting, using a neural network, sensory information of a physical area into one or more images; and performing, using one or more other neural networks, at least object detection on the one or more images to determine an output indicating at least whether or not one or more objects were detected.
[0010] In another exemplary embodiment, an apparatus comprises means for performing: converting, using a neural network, sensory information of a physical area into one or more images; and performing, using one or more other neural networks, at least object detection on the one or more images to determine an output indicating at least whether or not one or more objects were detected.
[0011] In an exemplary embodiment, a method is disclosed that includes training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area. The training comprises: creating generated depth maps by the generator neural network based at least on corresponding sensory information of a physical area, the generated depth maps comprising images of the physical area; determining classifications of fake or real by a discriminator neural network based on the generated depth maps and one or more real scene depth maps, the one or more real scene depth maps comprising a multi-dimensional map of the physical area, where the classifications are routed to the generator neural network and the discriminator neural network; and performing the creating the generated depth maps and the determining classifications until one or more criteria are met based on one or more loss functions. The method also comprises outputting information that defines a trained generator neural network, the trained generator neural network able to reconstruct images of the physical area based at least on corresponding sensory information of the physical area.
[0012] An additional exemplary embodiment includes a computer program, comprising instructions for performing the method of the previous paragraph, when the computer program is run on an apparatus. The computer program according to this paragraph, wherein the computer program is a computer program product comprising a computer-readable medium bearing the instructions embodied therein for use with the apparatus. Another example is the computer program according to this paragraph, wherein the program is directly loadable into an internal memory of the apparatus.
[0013] An exemplary apparatus includes one or more processors and one or more memories storing instructions that, when executed by the one or more processors, cause the apparatus at least to perform: training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area, the training comprising: creating generated depth maps by the generator neural network based at least on corresponding sensory information of a physical area, the generated depth maps comprising images of the physical area; determining classifications of fake or real by a discriminator neural network based on the generated depth maps and one or more real scene depth maps, the one or more real scene depth maps comprising a multidimensional map of the physical area, where the classifications are routed to the generator neural network and the discriminator neural network; and performing the creating the generated depth maps and the determining classifications until one or more criteria are met based on one or more loss functions; and outputting information that defines a trained generator neural network, the trained generator neural network able to reconstruct images of the physical area based at least on corresponding sensory information of the physical area.
[0014] An exemplary computer program product includes a computer-readable storage medium bearing instructions that, when executed by an apparatus, cause the apparatus to perform at least the following: training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area, the training comprising: creating generated depth maps by the generator neural network based at least on corresponding sensory information of a physical area, the generated depth maps comprising images of the physical area; determining classifications of fake or real by a discriminator neural network based on the generated depth maps and one or more real scene depth maps, the one or more real scene depth maps comprising a multidimensional map of the physical area, where the classifications are routed to the generator neural network and the discriminator neural network; performing the creating the generated depth maps and the determining classifications until one or more criteria are met based on one or more loss functions; and outputting information that defines a trained generator neural network, the trained generator neural network able to reconstruct images of the physical area based at least on corresponding sensory information of the physical area.
[0015] In another exemplary embodiment, an apparatus comprises means for performing: training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area, the training comprising: creating generated depth maps by the generator neural network based at least on corresponding sensory information of a physical area, the generated depth maps comprising images of the physical area; determining classifications of fake or real by a discriminator neural network based on the generated depth maps and one or more real scene depth maps, the one or more real scene depth maps comprising a multi-dimensional map of the physical area, where the classifications are routed to the generator neural network and the discriminator neural network; performing the creating the generated depth maps and the determining classifications until one or more criteria are met based on one or more loss functions; and outputting information that defines a trained generator neural network, the trained generator neural network able to reconstruct images of the physical area based at least on corresponding sensory information of the physical area.
BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In the drawings:
[0017] FIG. 1A is a block diagram of one possible and non-limiting example system in which example embodiments may be practiced;
[0018] FIG. IB is a block diagram of an example apparatus suitable for implementing any of the nodes in FIG. 1A;
[0019] FIG. 2 illustrates using access points as radar points to perform object detection;
[0020] FIG. 3 is a block diagram of radar sensing as a function of the environment; [0021] FIG. 4 is a flowchart illustrating an example method for performing object detection;
[0022] FIG. 5 is a diagram illustrating an example approach of Al (artificial intelligence)-based object detection from radar sensing;
[0023] FIG. 6A is a diagram illustrating an example framework for Al-based object detection;
[0024] FIG. 6B is a diagram illustrating another example framework for Al-based object detection;
[0025] FIG. 7 is a block diagram illustrating a training approach for a generator function;
[0026] FIG. 8A illustrates an example 2-dimensional (2D) heat map;
[0027] FIG. 8B illustrates an example 3-dimensional (3D) binary depth map;
[0028] FIG. 9 illustrates an example loss function used for training a generative adversarial network (GAN);
[0029] FIG. 10 illustrates an example for a 3D depth map reconstruction;
[0030] FIG. 11 illustrates an example for a 2D depth map reconstruction;
[0031] FIG. 12 illustrates an example object detection neural network;
[0032] FIG. 13 illustrates an example two-step bounding box and object detection; and
[0033] FIG. 14 illustrates an example end-to-end (E2E) network for object detection from a radar heat map.
DETAILED DESCRIPTION OF THE DRAWINGS
[0034] Abbreviations that may be found in the specification and/or the drawing figures are defined below, at the end of the detailed description section.
[0035] The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments. All of the embodiments described in this Detailed Description are exemplary embodiments provided to enable persons skilled in the art to make or use the invention and not to limit the scope of the invention which is defined by the claims. [0036] When more than one drawing reference numeral, word, or acronym is used within this description with and in general as used within this description, the may be interpreted as “or”, “and”, or “both”.
[0037] As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises”, “comprising”, “has”, “having”, “includes” and/or “including”, when used herein, specify the presence of stated features, elements, and/or components etc., but do not preclude the presence or addition of one or more other features, elements, components and/ or combinations thereof.
[0038] Any flow diagram (such as FIG. 4) or signaling diagram herein is considered to be a logic flow diagram, and illustrates the operation of an exemplary method, results of execution of computer program instructions embodied on a computer readable memory, functions performed by logic implemented in hardware, and/or interconnected means for performing functions in accordance with an exemplary embodiment. Block diagrams (such as FIGS. 6, 7, or 12-14) also illustrate the operation of an exemplary method, results of execution of computer program instructions embodied on a computer readable memory, functions performed by logic implemented in hardware, and/or interconnected means for performing functions in accordance with an exemplary embodiment.
[0039] The exemplary embodiments herein describe techniques for object detection accuracy improvement in wireless communications systems for radar sensing. Additional description of these techniques is presented after a system into which the exemplary embodiments may be used is described.
[0040] FIG. 1A illustrates an example system in which one or more example embodiments may be practiced. A number of nodes are shown: a user equipment (UE) 110; a base station 170; and network element(s) 190.
[0041] In FIG. 1A, a user equipment (UE) 110, as one of the nodes, is in wireless communication via wireless link 111 with a wireless network 100. The UE 110 is a wireless, typically mobile device that can access a wireless network. The UE 110 is illustrated with one or more antennas 128. The ellipses 101 indicate there may be multiple UEs 110.
[0042] The base station 170, as another of the nodes, provides access by wireless devices such as the UE 110 to the wireless network 100. It is noted that the base station 170 may also be referred to by other names, such as an access point. The base station 170 is illustrated as having one or more antennas 158. There are many options for the base station
170. In general, the base station 170 may be a RAN (radio access network) node, and in particular may be a gNB, which is the primary term used herein. That is, the base station 170 will be referred to as gNB 170. There are, however, many options including an eNB (evolved node B, e.g., an LTE, long-term evolution, base station) for the base station, as described below, or options other than cellular systems.
[0043] There are a number of configurations for the base station 170. One such is a “standalone” configuration, which includes all circuity as part of a single unit, and accesses the antennas 158. More commonly today, circuitry is split into one or more remote nodes 150 (accessing antennas 158) and central nodes 160. For instance, for 5G (fifth generation), a gNB might include a distributed unit (DU), or DU and radio unit (RU) as the remote nodes(s), and a central unit (CU) as the central node 160. For LTE, the base station 170 might include an eNB having a RRH (remote radio head) as remote node 150 and a base band unit (BBU) as a central node. The remote node(s) 150 are coupled to a central node 160 via one or more links
171. There may be multiple remote nodes 150 for a single central node 160, and this is indicated by ellipses 102, indicating multiple remote nodes, and ellipses 103, indicating additional links 171. The remote nodes 150 are remote in the sense they are contained in different physical enclosures from a physical enclosure containing a corresponding central node 160. The link(s) 171 may be implemented using fiber optics, wireless techniques, or any other technique for data communications.
[0044] Two or more base stations 170 communicate using, e.g., link(s) 176. The link(s) 176 may be wired or wireless or both and may implement, e.g., an Xn interface for 5G, an X2 interface for LTE, or other suitable interface for other standards. In some examples, there is a second BS (base station) 170-1, and the ellipses 104 indicate there could be additional base stations 170.
[0045] The wireless network 100 may include a network element or elements 190, as a third illustrated node, that may include core network functionality, and which provides connectivity via a link or links 181 with a data network 191, such as a telephone network and/or a data communications network (e.g., the Internet). Such core network functionality for 5G may include access and mobility management function(s) (AMF(s)) and/or user plane functions (UPF(s)) and/or session management function(s) (SMF(s)). Such core network functionality for LTE may include MME (Mobility Management Entity) functionality and/or SGW (Serving Gateway) functionality. These are merely exemplary functions that may be supported by the network element(s) 190, and note that both 5G and LTE functions might be supported. The RAN node 170 is coupled via a link 131 to a network element 190. The link 131 may be implemented as, e.g., an NG interface for 5G, or an SI interface for LTE, or other suitable interface for other standards.
[0046] For the examples below, there are some general functions that are performed. In FIG. 1A, sensing 106 (e.g., of radar signals) may be performed by a single base station 170. It is also possible for sensing 106 from multiple base-stations (e.g., BSs 170 and 170-1 or additional BSs) can be combined, e.g., for improved accuracy. In general, the basestation is a place where data is collected.
[0047] Object detection 107 may happen directly at base-station 170 (or other BSs 170-1 or more), at other network elements 190, and/or in the cloud (e.g., at remote server 108, in a data network 191 that is or has a cloud). The remote server 108 is another possible node in FIG. 1A. Output of the object detection may be used by any algorithm/solution that would use this information, for example to improve network performance by impacting the network scheduler, as an input to a digital twin (i.e., a digital representation of the real world), to assist street lights steering, and many more. For simplicity, the use of the output of object detection is shown occurring (see reference 107) in the BS 170, and/or the network element(s) 190, and/or the remote serverl08, although other locations are possible (e.g., such as a remote server 108 outside the data network 191). This example combines object detection and use (via reference 107), but object detection could be completely separate from use (e.g., perform on two different physical entities such as servers).
[0048] Furthermore, FIG. 1A shows an offline computer 173, as another possible node, that may be used in certain embodiments. The offline computer 173 in this example is able to communicate through the data network 191 with the wireless network 100. The offline computer 173 may perform some of the operations described herein, as detailed below.
[0049] FIG. IB illustrates an example apparatus 180 suitable for implementing any of the nodes in FIG. 1A. The apparatus 180 includes circuitry comprising one or more processors 120, one or more memories 125, one or more transceivers 130, one or more network interface(s) 155 and user interface (UI) circuitry and elements 157, interconnected through one or more buses 127. Since this is an example covering all of the nodes in FIG. 1A, some of the nodes may not have all of the circuitry. For example, a base station 170 might not have UI circuitry and elements 157. All of the nodes may have additional circuitry, not described here. FIG. IB is presented merely as an example.
[0050] Each of the one or more transceivers 130 includes a receiver, Rx, 132 and a transmitter, Tx, 133. The one or more buses 127 may be address, data, and/or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment, and the like. The one or more transceivers 130 are connected to one or more antennas 105, which may be one of the antennas 128 (from UE 110) or antennas 158 (from base station 170).
[0051] The one or more memories 125 include computer program code 123. The apparatus 180 includes a control module 140, comprising one of or both parts 140-1 and/or 140-2, which may be implemented in a number of ways. The control module 140 may be implemented in hardware as control module 140-1, such as being implemented as part of the one or more processors 120. The control module 140-1 may be implemented also as an integrated circuit or through other hardware such as a programmable gate array. In another example, the control module 140 may be implemented as control module 140-2, which is implemented as computer program code (having corresponding instructions) 123 and is executed by the one or more processors 120. For instance, the one or more memories 125 store instructions that, when executed by the one or more processors 120, cause the apparatus 180 to perform one or more of the operations as described herein. Furthermore, the one or more processors 120, one or more memories 125, and example algorithms (e.g., as flowcharts and/or signaling diagrams), encoded as instructions, programs, or code, are means for causing performance of the operations described herein.
[0052] The network interface(s) 155 are wired interfaces communicating using link(s) 156, which may be fiber optic or other wired interfaces. The link(s) 156 may be the link(s) 131 and/or 176 from FIG. 1A. The link(s) 131 and/or 176 from FIG. 1A may also be implemented using transceiver(s) 130 and corresponding wireless link(s) 111. The apparatus 180 may include only wireless transceiver(s) 130, only network interface(s) 155, or both wireless transceiver(s) 130 and network interface/ s) 155.
[0053] The apparatus 180 may or may not include UI circuitry and elements 157. These may include a display such as a touchscreen, speakers, or interface elements such as for headsets. For instance, a UE 110 of a smartphone would typically include at least a touchscreen and speakers. The UI circuitry and elements 157 may also include circuity to communicate with external UI elements (not shown) such as displays, keyboards, mice, headsets, and the like.
[0054] The computer readable memories 125 may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, firmware, magnetic memory devices and systems, optical memory devices and systems, fixed memory and removable memory. The computer readable memories 125 may be means for performing storage functions. The processors 120 may be of any type suitable to the local technical environment, and may include one or more of general-purpose computers, special purpose computers, microprocessors, digital signal processors (DSPs) and processors based on a multicore processor architecture, as non-limiting examples. The processors 120 may be means for performing functions, such as controlling the apparatus 180, and other functions as described herein.
[0055] Having thus introduced one suitable but non-limiting technical context for the practice of the exemplary embodiments, the exemplary embodiments will now be described with greater specificity.
[0056] The example embodiments herein relate at least to beyond 5G and 6G (sixth generation) technologies, where it is envisioned to enable joint communications and sensing (JCAS) as a new use case of wireless communications systems. See, e.g., H. Viswanathan and P. E. Mogensen, "Communications in the 6G Era," in IEEE Access, vol. 8, pp. 57063-57074, 2020, doi: 10.1109/ ACCESS.2020.2981745. The abundance of existing wireless communications infrastructure, as well as the inevitable growth in deployment of new hardware, provides a massive opportunity to utilize those cellular access points and devices for large scale sensing of the whereabouts towards enabling the digital twin of the world. By large scale sensing, it is understood that there are multiple sensing base stations in a given area, that could combine their sensing information to create a digital twin, which is a digital representation of the real world, possibly in real-time or close to real-time. In radar sensing, the transmitted radar signal is reflected by the target object. The receiver end can utilize the reflected signal and, through processing, compute the physical attributes of the target object including distance, direction, velocity, and the like. This is depicted in FIG. 2, which illustrates using NR (new radio) and/or 6G access points as radar points that can help detect objects in the desired area of sensing. Here, access points 210-1 and 210-2 (e.g., versions of base stations 170) are utilized as radar points to detect a car 220 crossing a traffic intersection 230.
[0057] Techniques herein relate to improving the object detection, classification and localization abilities of an, e.g., OFDM (orthogonal frequency division multiplexing)- based cellular JCAS system, such as the network 100 of FIG. 1. However, the techniques are applicable to any radar or single- or multi-sensory entity that potentially has a distorted and noisy observation of the physical environment, including Lidar (light detection and ranging), backscatter tracking, and the like. That is, the apparatus 180, and some or all of FIG.l, can apply to radar or multi-sensory entities.
[0058] Further details are presented after an overview of technology in this area is presented. Object detection and classification is extensively investigated under image processing, i.e., localizing and classifying an object from a 2D or 3D image of the environment and distinguishing the object from the surroundings using techniques such as image segmentation, object recognition, and the like.
[0059] The advances in those techniques have proven to be not directly scalable to radar images. In other words, an object detection technique developed and tested with visual images generally is not expected to perform well on radar observations of the same environment. This is mainly due to the fact that such techniques are developed based on the features from the visual world that can be captured from the 2D or 3D images of the environment, but not directly from the radar images of the same environment.
[0060] Additionally, radar observation of an environment is in most cases a reduced version of the visual observation of the environment. This is depicted as an example in FIG. 3, which shows a block diagram of radar sensing as a complicated function of the environment. Block 310 indicates there is a point cloud, or a 3D or 2D image or map of the sensing environment.
[0061] Depict the function of an “ideal” radar sensing (which can be created by a perfect radar hardware with low impact of noise and interference, as well as a fine angular resolution) as i(x). This is applied to block 310. Block 320 has the “ideal” radar sensing as the following: one-degree beams in every direction at every time instance; and minimal hardware impact. Similarly, a non-ideal radar 6G JCAS (impacted by lower resolution and higher impact from waveform and hardware), is denoted by f2 (x) as a function of the ideal radar observation. As indicated by block 330, JCAS is impacted by the following: fat beams, one at a time slot; hardware impairment; and poor sampling over time and frequency. An example problem at hand is then as follows:
[0062] 1) The object detection tools are best suited for the point cloud, or 2D/3D images of the environment (LHS, left hand side, block 310 in FIG. 3).
[0063] 2) What is available as observation of the environment in JCAS is the output i.e., RHS, right hand side, block 330 of FIG. 3.
[0064] The inventors have asked whether it is possible to take the 6G JCAS radar observation and process the observation back into the 2D or 3D image of the physical area in order to improve the object detection and classification?
[0065] Herein, an ML (machine leaming)-based approach is proposed to ‘invert’ the radar observation back to the 2D or 3D image of the scene of interest. This inversion is illustrated by the arrows 301 and 302. This inversion process is normally not feasible. However, using the prior knowledge about the environment as well as the form and structure of objects that are to be detected, it has been shown by the inventors that it is possible to train a neural network to efficiently perform such inversion in real-time.
[0066] The prior knowledge is ‘learned’ by a generative adversarial network (GAN) and is referred to as ‘generative priors’ which is normally not feasible to model analytically, but can be learned by the neural network using a large dataset of observations from the scene of interest. An E2E neural network design approach is then proposed that takes radar heatmap, converts the heatmap to a 3D point cloud of the scene of interest, and detects objects in the scene.
[0067] The matter of an ‘inverse problem’ in general has been studied in the literature. The following have studied super-resolution radar, using Al (artificial intelligence) (GAN based and otherwise): K. Armanious, S. Abdulatif, F. Aziz, U. Schneider and B. Yang, "An Adversarial Super-Resolution Remedy for Radar Design Trade-offs," 201927th European Signal Processing Conference (EUSIPCO), 2019, pp. 1-5, doi: 10.23919/EUSIPCO.2019.8902510; Fang, Shiwei, and Shahriar Nirjon, "Al-enhanced 3D RF representation using low-cost mmwave radar”, Proceedings of the 16th ACM Conference on Embedded Networked Sensor Systems, 2018; and J. Guan, S. Madani, S. Jog, S. Gupta and H. Hassanieh, "Through Fog High-Resolution Imaging Using Millimeter Wave Radar," 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 11461-11470, doi: 10.1109/CVPR42600.2020.01148. The work in the following uses ‘image segmentation’ methods to detect object in radar sensing, thus it fits the conventional approach of object detection described in the below-described FIG. 5: A. Danzer, T. Griebel, M. Bach and K. Dietmayer, "2D Car Detection in Radar Data with PointNets," 2019 IEEE Intelligent Transportation Systems Conference (ITSC), 2019, pp. 61-66, doi: 10.1109/ITSC.2019.8917000.
[0068] In case of Guan et al., that work studies the use of generative priors to create high-resolution image of an object by training a GAN. The focus is on using the generative priors related to the object’s form and structure. In contrast, an approach herein is to invert a radar effect, which includes inverting the effect of multi-path and specularity. One aim herein is object detection. Therefore, instead of focusing on high-resolution representation of objects (as in Guan et al.), we focus on recreating high-level features of the objects, which can be utilized in the E2E network for object detection. As a result, our proposal requires a smaller set of data samples for different objects, and enables the E2E network to actually distinguish between objects, even though the 3D map may be the same between legacy techniques and the techniques herein. The work in Guan et al. also differentiates with the techniques herein at least in terms of the following.
[0069] 1) Proposed herein is an E2E approach to object detection, e.g., with a small dataset. This is achieved by separating the object detection network from the radar inversion network and training each separately, but, at the end, tuning the parameters E2E. By contrast, Guan et al. focuses on high resolution representation of a specific object, which requires large datasets.
[0070] 2) In terms of context, the work herein addresses object detection for, e.g.,
JCAS. The work in Guan et al. looks into super-resolution radar context. This implies differences in methods of radar signal collection (waveform, frequency, hardware effects, and resource sharing with communications function). Thus, the radar input in examples herein is different from the typical radar input, e.g., in Guan et al., in the sense that the radar inputs has the impact of JCAS signal processing and hardware considered in the signal. That is, the radar input is impacted by signal processing and hardware effects.
[0071] 3) Additionally, Guan et al. uses a 2D depth map of the scene as ground truth for training and inference. In contrast, while a 2D map can be used, a focus herein is on a 3D point cloud of the environment, which includes all the objects in the scene as well as the static background. Therefore, in the inference stage, example networks herein regenerate the complete scene, including static background, non-target dynamic and static objects, as well as the target objects. As a result, an example network herein is capable of detecting multiple ones of the same object in one scene, or differentiating one object of interest from a non-target object in the same scene. This does not happen in Guan et al.
[0072] FIG. 4 illustrates an example method in accordance with the present disclosure. Blocks 400, 402, and 404 are assumed to be performed offline, that is, not in the wireless network 100 and instead performed by offline computer 173. This is not a limitation, however, and the blocks could be performed by a node in the wireless network 100. Block 406 is assumed to be performed by a node in the wireless network 100.
[0073] At block 400, the offline computer 173 rains a generative adversarial network (GAN) for reconstruction of 2D or 3D images of a physical area of interest, using wireless radar sensing of the area. In block 400, this is more likely to use mainly synthetic data, since a large number of data samples are needed. However, in case of availability of data, real world measurements could be used too. At block 402, the offline computer 173 trains an object detection network to identify a class of objects (e.g., cars, pedestrians, bikes, trucks, and the like) using, e.g., 2D or 3D images of the area of interest. In an example, object detection performed by the offline computer 173 may include classifying the detected objects. At block 404, the offline computer 173 performs end-to-end (E2E) fine-tuning of a cascade network of the GAN from block 400 and the detection network from block 402.
[0074] In block 406, learning is transferred from the cascade network of block 404 using real-word measurement data. In further detail, block 406 transfers learning into the real-world domain, using real world measurement data collected by the base station 170 (or base stations). The actual algorithm for block 406 may be run in base station 170, if necessary computational power is available, or by other network elements 190, or in the cloud 191 via, e.g., the remote server 108. Real-world measurement data is impacted by hardware imperfections and other effects that are hard or impossible to capture in simulation. In block 406, this is more likely to use real measurements. However, using synthetic data is also applicable.
[0075] Further details are provided now. It should be noted, for these details, that the generator and discriminator may be referred to as generator and discriminator, generator model and discriminator model, or generator neural network (NN) and discriminator neural network (NN), respectively. [0076] At block 400, a generator NN is trained, e.g., using a generative adversarial network (GAN) architecture, to reconstruct images from the scene of interest based on a radar observation of the scene. In an example, the training may be targeted to reconstruct high-level visual representation of the scene, including the objects that are intended for detection and corresponding classification.
[0077] During training, the generator and discriminator NNs may be trained jointly using ground truth data (e.g., a 3D image of the area of interest) and input data (e.g., a 3D radar heatmap of the area of interest). Training data may include data samples that contain one or multiple of the objects of interest in the classifications; training data can be created synthetically using, e.g., ray-tracing based radar simulation tools. During inference, input of the generator is radar input data (e.g., a 3D radar heatmap of the area of the interest). The desired output of the generator is the equivalent 3D image of the area of interest.
[0078] In one embodiment of block 400, the GAN is trained using multiple radar inputs, e.g., from multiple mono-static or bi-static radars, where the generator is trained to fuse the multiple input sources together to generate the physical reconstruction of the environment. In another embodiment of block 400, the GAN is trained using a memory-enabled network, e.g., LSTM (long short-term memory), where multiple consecutive radar frames from the same scene can be fed into the generator. In still another embodiment of block 400, the GAN is trained using a fusion of multiple sensory information, e.g., radar + Lidar + localization information from other sources, including GPS (global positioning system), GNSS (global navigation satellite system), and the like.
[0079] At block 402, an object detection (which may include classification) network is trained to detect (and classify) objects of interest in the sensing area of interest. During training, input data includes a 3D image of the area of interest, while ground truth includes presence/absence of object(s), and likelihood of a detected object belonging to one of the defined classed of objects. During inference, the GAN reconstructed image of the area of interest (from step 1) is fed into the object detection network.
[0080] At block 404, a cascade network of the generator output at block 400 and the object detection output at block 402 is formed; the cascade network is fine-tuned end to end to improve the object detection. Block 404 also may involve the following: fine-tuning and transfer learning with E2E training; most layers are frozen while only updating the few last layers; and calibrating the cascade network using a small data set. [0081] In block 406, this involves transfer learning of the E2E training from block 404; similarly to block 404, a few of the last layers of the E2E network are fine-tuned using real- world measured data, e.g., from the scene where the proposed application is to be deployed in. For transfer learning (also referred to as learning transfer), consider the following. Training of any NN network (CNN, convolutional neural network, GAN, or other) from scratch requires a huge dataset and is time consuming and computationally burdensome. To overcome this problem and make a NN more generic, the network can be “pre-trained”, for example using synthetic data until achieving satisfactory accuracy. This is what is happening in blocks 400, 402, and 404. This pre-trained network performs well but can do better in specific environment. That’s where the second stage of training comes (the transfer learning of block 406), where pre-trained network is being trained again (but not fully, only last few layers) using data collected in the environment. Because only a few layers are being trained (and not from the scratch, but using weights from initial training), the dataset required to achieve a high accuracy is much smaller.
[0082] The example embodiments may include lower complexity since the training may be performed part-by-part. The example embodiments may provide an improved detection rate as compared to systems that use processed radar data. The example embodiments may provide an improved occupancy grid reconstruction from multiple radar and other sensory information. Th example embodiments may provide an easier fusion of multiple sources of sensory data, e.g., multiple mono-static or bi-static radar. The example embodiments may provide efficient usage of static background information (through developing generative priors), to improve efficiency of radar sensing towards object detection.
[0083] FIG. 5 illustrates an example approach in Al-based object detection from radar sensing. The approach 500 generally includes training a neural network to detect objects directly from a raw radar image, or from a pre-processed radar image. In FIG. 5, the approach 500 performs radar processing 510 on raw radar information (info) 505. As indicated by block 515, the radar signal processing includes expert feature extraction, MUSIC (multiple signal classification), and the like. The output is processed radar information 520, and a radar object detection/classification model 525 performs detection and corresponding classification to output objects and classes 535. As indicated in block 530, the block 525 uses model-based algorithms and/or data-based machine learning detection schemes, designed for radar sensing. [0084] In some instances, the advanced Al-based object detection techniques described in reference to FIG. 5 may be developed for visual images and, therefore, may not be readily useful in object detection. As another example, the neural network (in block 525) may require large datasets to be trained.
[0085] FIG. 6A illustrates an example framework 600 for Al-based object detection. The framework 600 is configured to receive raw radar information 505. The framework 600 is configured to perform radar inversion 610 on the raw radar information. As indicated by block 615, a radar inversion module (e.g., using GANs) is configured to exploit data-driven generative priors to reconstruct the scene from radar information. The radar inversion module generates images 620 from the scene. The images 620 may be processed in block 625. The framework 600 may use (optional) image processing techniques including, e.g., expert feature extraction, transform domains, and the like.
[0086] The image processing at block 625, outputs processed images 635 to an (image) object detection/classification module 640 for further processing. As illustrated in block 645, the object detection/classification module 640 performs model-based algorithms and/or data-based machine learning detection schemes, e.g., bounding box detection, image segmentation, and the like. The object detection/classification module 640 generates one or more outputs, such as, one or more objects and associated one or more classes of the one or more objects.
[0087] In terms of how FIG. 6A applies to the rest of the figures, this is a broad overview: FIGS. 7, 8A, 8B, and 9 related to training for forming a NN able to perform the radar inversion 610; FIGS. 10 and 11 illustrate how images 620 could be formed; FIGS. 12 and 13 illustrate examples of processing 635 and object detection/classification module 640; and FIG. 14 illustrates an example of end-to-end processing.
[0088] It is noted that radar information and processing is primarily described herein. However, other sensory information can be used instead of or in addition to the radar. See FIG. 6B, which is a diagram illustrating another example framework for Al-based object detection. In this example, there is raw sensory information 505-1 instead of (e.g., only) raw radar information 505. There is sensory inversion 610-1 that operates on the raw sensory information 505-1 to produce images 620-1 from the scene. As block 615-1 indicates, the sensory inversion module 610-1 610 (e.g., using GAN(s)) exploits data-driven priors to reconstruct the scene from the sensory information. As indicated by block 690, the raw sensory information 505-1 is formed by a single sensory entity with one of the following or a multi- sensory entity with multiple ones of the following types of sensory information: Radar, Lidar, backscatter tracking, localization information from other sources, including GPS, GNSS, and the like. Training would also have to account for the selected single type of sensory information or multiple types of sensory information.
[0089] FIG. 7 illustrates an example training of a generator network for radar inversion, which could be implemented in block 400 of FIG. 4. In this example, a GAN architecture 700 has a 3D radar heat map 710 that is sent to a generator model 720, which produces a generated depth map 725. A heat map is a data visualization technique that shows magnitude of a phenomenon using, e.g., color, in two or three dimensions, and a radar heat map is a heat map formed by radar. The depth map 725 and a real scene depth map 715 are input to a discriminator model 730, which provides an output 745 to a classification module 740 that outputs fake/real outputs. These outputs are sent via update 750 to the generator model 720 and via update 755 to the discriminator model 730.
[0090] In further detail, a generator neural network (NN) 720 is trained e.g., using a generative adversarial network (GAN) architecture 700, to reconstruct images from the scene of interest from a radar observation of the scene. The training is illustrated in FIG. 7 and works as follows.
[0091] Radar heatmap 710: Normalized received power may be used for each point of the coordinate system of interest. In JCAS systems, where a limited number of analog beams are available at the transmitter and receiver, the heatmap 710 may be created per beam and then combined together into one radar heatmap using weighted summation of the heatmap 710 per beam. The weights are designed, e.g., depending on the segment of the coordinate system that each beam covers, where weight 1 (one) may be given to the covered area and weight 0 (zero) to the non-covered parts. In other examples, the gain of the beam at each coordinate point can be chosen as the soft weighting parameter.
[0092] 3D depth map (ground truth for training) 715: a real scene depth map may be generated by assigning 0 (zero) for every point of the coordinate system where there is free space, and 1 (one) to the rest of the points where objects are present. This is an example of an occupancy grid. As another example, see FIG. 8A, which illustrates a 2D heat map, and FIG. 8B, which illustrates a 3D binary depth map. For the 2D scene heat map of FIG. 8A, each direction (azimuth, elevation) represents the distance to the closest object. That is, r = f(x,y) (where x = azimuth and y = elevation) corresponds to the distance to the closest object. Each point of the 2D map corresponds to a vertical and azimuth angle from the point of view of the measurement device (e.g., Radar). In further detail, the value of each point may correspond to the depth (e.g., distance) of the first observable object in the direction of the angles, and the value may further contain information about material, color, conductivity, and the like, although these are not needed.
[0093] In the 3D binary depth map of FIG. 8B, for each direction and range (azimuth, elevation, and range), may return 1 (one) if an object exists, but return 0 (zero) otherwise. See the following:
In a more general example, the points of the map are multi-layer, where one layer is the binary map described above, and one layer includes the physical characteristics of the object (e.g., conductivity of the surface or material index out of a preassigned set of material indexes).
[0094] Other possibilities for the real scene depth map 715 include one or both of an occupancy grid or one or multiple two dimensional depth maps (e.g., from different angles). An (e.g., a binary) occupancy grid usually refers to a grid-based matrix, where each entry corresponds to one point in a 3D grid of interest. The value in each entry can be binary (e.g., object present or not) or non-binary, corresponding to classes of objects or to material and color characteristics of the objects, or the like. Meanwhile, a 3D point cloud usually refers to a fairly similar definition, where instead of a grid-based matrix, the data format includes all the points from the 3D space where the observer has collected “information” about them. This means the coordinate of the point is corresponded to the value of the measurement.
[0095] Other options include heat maps and even a multitude of depth maps could be taken, e.g., from different angles and or from different sources, e.g., radar, Lidar, and others. That is, there could be multiple implementations of the real scene depth map 715, as the environment is constantly changing. However, the discriminator 730 may look at one specific realization of reconstruction at the time and comparing this with the real data. In other words, the training process for the GAN architecture 700 could be performed for individual ones of multiple real scene depth maps 715. Alternatively, the GAN architecture 700 can be trained for all the multitude of depth maps at the same time. [0096] For FIG. 7, the expected output 725 of the trained generator NN 720 is a high-level visual representation of the scene, which resembles the 3D ground truth depth map (as the real scene depth map 715).
[0097] For the training data, these may include data samples that contain one or multiple of the objects of interest in the classification. Training data can be created synthetically using, e.g., ray-tracing based radar simulation tools, or using measurements of a radar system operation.
[0098] With respect to the training, the training is targeted to reconstruct high-level visual representation of the scene, including objects that are intended for detection and classification. The generator NN 720 and discriminator NN 730 are trained jointly, where generator creates 3D depth maps 725 from the 3D radar heatmaps 710, while the discriminator tries to distinguish the generated output 725 from the ground truth (the real scene depth map 715). In the latter, a loss function may be used to measure the loss between the ground truth and generated map. For example, the loss function in FIG. 9 can be used for the GAN training, where 1 values represent the weight of each loss term in the overall loss function. The training may be performed using ground truth data (e.g., 2D or 3D image of the area of interest) and input data (2D or 3D radar heatmap of the area of interest).
[0099] For inferencing, input of the generator NN 720 is radar input data (e.g., 2D or 3D radar heatmap 710 of the area of the interest), while the desired output 725 of the generator NN 720 is the equivalent 2D or 3D image of the area of interest.
[00100] In one embodiment, the GAN 700 is trained using multiple radar inputs, e.g., from multiple mono-static or bi-static radars. In such cases, the multiple inputs may be fed, such that, in one example, each mono static radar or bi-static radar input is fed in as a separate input layer and, in another example, a fusion function is defined to fuse the multiple input sources together, e.g., for each point of the coordinate system of interest, the measurements of all radar input layers are combined with weighted combining.
[00101] In one embodiment, the GAN 700 is trained using a memory-enabled neural network, e.g., LSTM, where multiple consecutive frames of radar heatmap from the same scene are fed into the generator NN 720. In such case, the output is the average high-level representation of the physical scene from the multiple frames. This is a very useful embodiment for JCAS systems, where the time resolution of radar heatmap frames is very small compared to the variation speed in the scene (e.g., movement of a car). Therefore, the generator can leverage multiple frames to perform the 3D map reconstruction, which improves the accuracy of object detection in the next stage too.
[00102] In one embodiment, the GAN 700 is trained using multiple sensory information, e.g., a radar heatmap, a Lidar heatmap, localization information collected at the 5G LMF (location management function), GPS, GNSS, and the like, and these are fused together and fed into the generator NN 720, or fed in as separate layers to the generator NN 720.
[00103] Example expected outputs include the following. An example in FIG. 10 demonstrates the case of 3D depth map reconstruction. The radar heat map 710-1 is an input, as is the ground truth 715-1, which are both processed by a generator NN 720 to create an output 725-1. As seen in FIG. 10, the output 725-1 generated by the trained generator NN 720 resembles the ground truth 715-1 much better visually, compared to the radar heatmap 710-1, which is difficult to characterize.
[00104] The plots in FIG. 11 show the 2D version of the example in FIG. 10. A 2D version of the ground truth 715-1 and a 2D version of the output 725-1 generated by the network are shown. In case of 2D, the depth map is a 2D matrix containing the value of the depth distance to the first object from the point of view of the radar, as described with reference to FIG. 8.
[00105] Training of an object detection network is now described. This is block 402 of FIG. 4, and may be used to train (image) object detection/classification module 640 of FIG. 6A. An object detection/classification neural network is trained to detect objects of interest in the sensing area of interest.
[00106] The input includes the 3D or 2D map of the environment, created by the trained generator NN 720, as such training was described above. Consider that the trained generator NN 720 is now performing the radar inversion 610 of FIG. 6A, and produces images 620 that are input to the (image) object detection/classification module 640. In one embodiment, the input can be a fusion of the map from the generator together with other sources of information, including, e.g., Lidar heatmap and/or localization information.
[00107] The map from the generator NN 720 can be accompanied by an additional input layer which includes the static background of the scene of interest. This enables the network to distinguish between objects and background. Another option is to subtract the background information from the input map. [00108] Two different example options are assumed for the output, although the techniques herein are not limited to these. These two options are described below.
[00109] Option 1 : the network is trained to estimate likelihood of presence of an object from a set of pre-defined object categories (which may be thought of as classes for this example). One example of this is, if there are 10 object categories, 10 likelihoods may be output, one for each category. Turn to FIG. 12, which is an example illustrating an object detection neural network. Reference may also be made to FIG. 6A, as the reference numbers from there are also used in FIG. 12. In the example of FIG. 12, the input is a 3D point cloud 620-1 with the background subtracted. Background may be subtracted using multiple techniques. For instance, a 3D point cloud could be produced, then the background would be subtracted. Alternatively, the subtraction may be performed at the at the same time of the radar inversion, so the output has the 3D point cloud with background subtracted. There is an optional conversion 625-1 (e.g., an image processing) to a 3D occupancy grid 635-1. The object detection neural network 640-1 uses the grid 635-1 to output a likelihood 650-1 of presence of an object that belongs to (e.g., known) object categories.
[00110] Option 2 : The network comprises a two-step detection. In step a, a neural network is trained to detect bounding boxes for objects in the scene. A separate network is then trained (step b) to estimate likelihood of objects in each bounding box separately. This is especially useful for scenes with multiple objects.
[00111] Option 2 is described through reference to FIG. 13, which is an example illustrating a two-step bounding box and object detection. There are two steps: step a, 1305-1; and step b, 1305-2, which is repeated for all the detected bounding boxes. For step a, 1305-1, a 3D point cloud 620-1 with background subtracted is input to a first object detection/classification NN 640-2, which outputs bounding box detection 650-2 for objects in the scene.
[00112] Step b, 1305-2, uses an input of one bounding box segment 1320 of the 3D point cloud, and step b may be repeated for all the detected bounding boxes. A second object detection/classification NN 640-3 operates on each bounding box segment to produce a likelihood 620-1 of presence of an object belonging to (e.g., known) object categories within the corresponding box segment. [00113] In one embodiment, which applies to both options 1 and 2, training is performed using a 2D or a 3D point cloud map of the area of interest instead of 2D or 3D maps created by the generator NN 720.
[00114] With respect to training, ground truth for Option 1 may include a vector of binary values for each category of objects (e.g., 0, zero, if object is not present and 1, one, if object is present).
[00115] For Option 2 above, the ground truth in step a, 1305-1, may be the exact bounding box of the training set, which for step b, 1305-2, the ground truth includes the binary values for an object bounding box for each category of objects.
[00116] During inference, the GAN reconstructed image of the area of interest is fed into the object detection network.
[00117] Concerning end-to-end fine tuning and transfer learning (blocks 404 and 406 of FIG. 4), consider the following. A cascade network 600-1 of the generator NN 720 and the object detection from FIG. 12 is formed as seen on FIG. 14. FIG. 14 is an example illustrating an E2E network for object detection from a radar heat map. In this example, a radar heat map 505-1 is input to a generator network 610-1 (e.g., a trained generator network 720- 1), which produces a reconstructed 3D point cloud 620-2. The cloud 620-2 is input to the object detection network 640-1, which outputs a likelihood 650-1 of presence of an object that belongs to (e.g., known) object categories.
[00118] The cascade network is fine-tuned (see block 404 of FIG. 4) end-to-end to improve the object detection. For fine-tuning, a small set of layers in the NN are selected as tunable layers 1410. This example shows a layer 1410-1 in the generator network 610-1 and a layer 1410-2 in the object detection network 640-1. It is noted that multiple layers may be used, and the networks 610-1 and 640-1 need not have the same number of layers implemented as tunable layers. Furthermore, the tuning is typically performed via tuning the weights that are multiplied by the feedforward values exchanged between layers of a NN.
[00119] The remaining layers are frozen while only updating the few last tunable layers. This way, one can calibrate the cascaded E2E network using a small data set. As previously described with reference to block 406 of FIG. 4, as part of learning transfer, a trained NN is used with real-world measurement data, and only a few of the layers of the trained NN are updated. This means a small data set - as compared to the data set used for initially training the NN - is used. [00120] The input data may be the input as described above (e.g., radar heatmap and 3D depth map (ground truth for training)), and the output data may be the output from Options 1 or 2 described above. (In case of Option 2, either the step a network is only cascaded with the generator, or the cascaded network has an adaptive number of equivalent step b networks in parallel at the output of step a network.)
[00121] With respect to transfer learning (see step 406 of FIG. 4), a few of the last layers of the E2E network may be fine-tuned using real- world measured data defined the same way as the input and output above, but collected, e.g., from the scene in which the proposed object detection application is to be deployed.
[00122] It is noted that FIG. 14 may apply to both blocks 404 (e.g., E2E fine- tuning of a cascade network) and 406 (e.g., learning transfer of the system). It is further noted that, for the cascade network in block 404, the data can be also real, e.g., training the network with synthetic data of the scene and fine-tuning the cascade with real data from the same scene. Furthermore, transfer learning (as in block 406) can be performed with synthetic data too. That is, training the network for a certain scene may be performed, and then the learning of the cascade network is transferred to a different scene by fine tuning the learning using, e.g., a set of synthetic data from the new scene.
[00123] Without in any way limiting the scope, interpretation, or application of the claims appearing below, a technical effect and advantage of one or more of the example embodiments disclosed herein is overcoming hardware impairments and other communication network related imperfection. Another technical effect and advantage of one or more of the example embodiments disclosed herein is improving object detection and classification accuracy. Another technical effect and advantage of one or more of the example embodiments disclosed herein is improved performance of any other algorithm that would benefit from higher resolution of radar image.
[00124] The following are additional exemplary embodiments.
[00125] Example 1. A method, comprising:
[00126] converting, using a neural network, sensory information of a physical area into one or more images; and
[00127] performing, using one or more other neural networks, at least object detection on the one or more images to determine an output indicating at least whether or not one or more objects were detected. [00128] Example 2. The method according to example 1, wherein the performing at least object detection also performs classification of detected objects so the output indicates whether or not one or more objects were detected using likelihood of presence of the one or more objects belonging to an object category.
[00129] Example 3. The method according to example 2, wherein there are multiple categories, and classification outputs multiple likelihoods corresponding to the multiple categories.
[00130] Example 4. The method according to any one of examples 2 or 3, wherein the images comprise three-dimensional point clouds where backgrounds have been subtracted.
[00131] Example 5. The method according to example 4, further comprising converting the three-dimensional point clouds where backgrounds have been subtracted to three-dimensional occupancy grids, wherein the performing at least object detection performs at least object detection on the three-dimensional occupancy grids.
[00132] Example 6. The method according to any one of examples 2 or 3, wherein the one or more images comprise three-dimensional point clouds with background subtracted and wherein the performing, using the one or more other neural networks, at least object detection on the one or more images to determine an output comprises the following:
[00133] performing a first detection using a first of the one or more other neural networks to produce first output comprising boundary box detection for objects in a scene, wherein there are multiple bounding box segments for the 3D point clouds; and
[00134] performing a second detection on the first output using a second of the one or more other neural networks to produce second output comprising the likelihoods of presence of objects belonging to corresponding object categories in individual bounding box segments.
[00135] Example 7. The method according to any one of examples 1 to 6, wherein the converting sensory information and performing at least object detection are performed to train the neural network and the one of more other neural networks using the sensory information to adjust one or more tunable layers in one or both of the neural network or the one or more other neural networks. [00136] Example 8. The method according to example any one of examples 1 to 7, wherein the sensory information comprises radar information developed from one or more base stations in a wireless system.
[00137] Example 9. A method, comprising:
[00138] training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area, the training comprising:
[00139] creating generated depth maps by the generator neural network based at least on corresponding sensory information of a physical area, the generated depth maps comprising images of the physical area;
[00140] determining classifications of fake or real by a discriminator neural network based on the generated depth maps and one or more real scene depth maps, the one or more real scene depth maps comprising a multi-dimensional map of the physical area, where the classifications are routed to the generator neural network and the discriminator neural network; and
[00141] performing the creating the generated depth maps and the determining classifications until one or more criteria are met based on one or more loss functions; and
[00142] outputting information that defines a trained generator neural network, the trained generator neural network able to reconstruct images of the physical area based at least on corresponding sensory information of the physical area.
[00143] Example 10. The method according to example 9, wherein the real scene depth map comprises an occupancy grid, or one or multiple two-dimensional depth maps, or both the occupancy grid and the one or multiple two-dimensional depth maps.
[00144] Example 11. The method according to example 9, wherein the sensory information comprises radar information, wherein the real scene depth map comprises a two- dimensional map based on azimuth and elevation, where points in the map correspond to vertical and azimuth angles from a point of view of a radar measurement device generating the radar information, and values of the points correspond to distance of a first observable object in a direction of the angles corresponding to the points.
[00145] Example 12. The method according to example 9, wherein the real scene depth map comprises a three-dimensional map generated at least by assigning a first value to points of a coordinate system, corresponding to the physical area, where there is free space, and assigning a second value to points where objects are present in the coordinate system, where a function is used to determine whether an object does or does not exist at individual directions based on azimuth, elevation, and range.
[00146] Example 13. The method according to example 9, wherein the generative adversarial network is trained using sensory information comprising one or more of the following: radar information, Lidar information, backscatter tracking information, or localization information.
[00147] Example 14. The method according to example 13, wherein the generative adversarial network is trained using sensory information comprising multiple ones of the following: the radar information, the Lidar information, the backscatter tracking information, or the localization information, and the multiple sensory information are fused together and fed into the generator neural network, or fed in as separate layers to the generator neural network.
[00148] Example 15. The method according to any one of examples 9 to 14, further comprising training an object detection neural network on identifying classes of objects using images from the physical area.
[00149] Example 16. The method according to example 15, further comprising:
[00150] performing training a cascade network of the trained generator neural network and the object detection neural network, wherein output of the trained generator neural network is input to the object detection neural network, the training the cascade network of the trained generator neural network performed to adjust one or more tunable layers in one or both of the trained generator neural network or the object detection neural network.
[00151] Example 17. The method according to example 16, wherein the performing training the cascade network creates a trained cascade network, and the method further comprises performing another training of the cascade network using measurement data from one or more sensing entities of the physical area, the other training performed to adjust one or more tunable layers in one or both of the trained generator neural network or the object detection neural network.
[00152] Example 18. A computer program, comprising instructions for performing the methods of any of examples 1 to 17, when the computer program is run on an apparatus. [00153] Example 19. The computer program according to example 18, wherein the computer program is a computer program product comprising a computer-readable medium bearing instructions embodied therein for use with the apparatus.
[00154] Example 20. The computer program according to example 18, wherein the computer program is directly loadable into an internal memory of the apparatus.
[00155] Example 21. An apparatus, comprising means for performing:
[00156] converting, using a neural network, sensory information of a physical area into one or more images; and
[00157] performing, using one or more other neural networks, at least object detection on the one or more images to determine an output indicating at least whether or not one or more objects were detected.
[00158] Example 22. The apparatus according to example 18, wherein the performing at least object detection also performs classification of detected objects so the output indicates whether or not one or more objects were detected using likelihood of presence of the one or more objects belonging to an object category.
[00159] Example 23. The apparatus according to example 19, wherein there are multiple categories, and classification outputs multiple likelihoods corresponding to the multiple categories.
[00160] Example 24. The apparatus according to any one of examples 19 or 20, wherein the images comprise three-dimensional point clouds where backgrounds have been subtracted.
[00161] Example 25. The apparatus according to example 21, wherein the means are further configured for performing converting the three-dimensional point clouds where backgrounds have been subtracted to three-dimensional occupancy grids, wherein the performing at least object detection performs at least object detection on the three- dimensional occupancy grids.
[00162] Example 26. The apparatus according to any one of examples 19 or 20, wherein the one or more images comprise three-dimensional point clouds with background subtracted and wherein the performing, using the one or more other neural networks, at least object detection on the one or more images to determine an output comprises the following: [00163] performing a first detection using a first of the one or more other neural networks to produce first output comprising boundary box detection for objects in a scene, wherein there are multiple bounding box segments for the 3D point clouds; and
[00164] performing a second detection on the first output using a second of the one or more other neural networks to produce second output comprising the likelihoods of presence of objects belonging to corresponding object categories in individual bounding box segments.
[00165] Example 27. The apparatus according to any one of examples 18 to 23, wherein the converting sensory information and performing at least object detection are performed to train the neural network and the one of more other neural networks using the sensory information to adjust one or more tunable layers in one or both of the neural network or the one or more other neural networks.
[00166] Example 28. The apparatus according to example any one of examples 18 to 24, wherein the sensory information comprises radar information developed from one or more base stations in a wireless system.
[00167] Example 29. An apparatus, comprising means for performing:
[00168] training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area, the training comprising:
[00169] creating generated depth maps by the generator neural network based at least on corresponding sensory information of a physical area, the generated depth maps comprising images of the physical area;
[00170] determining classifications of fake or real by a discriminator neural network based on the generated depth maps and one or more real scene depth maps, the one or more real scene depth maps comprising a multi-dimensional map of the physical area, where the classifications are routed to the generator neural network and the discriminator neural network; and
[00171] performing the creating the generated depth maps and the determining classifications until one or more criteria are met based on one or more loss functions; and
[00172] outputting information that defines a trained generator neural network, the trained generator neural network able to reconstruct images of the physical area based at least on corresponding sensory information of the physical area. [00173] Example 30. The apparatus according to example 26, wherein the real scene depth map comprises an occupancy grid, or one or multiple two-dimensional depth maps, or both the occupancy grid and the one or multiple two-dimensional depth maps.
[00174] Example 31. The apparatus according to example 26, wherein the sensory information comprises radar information, wherein the real scene depth map comprises a two- dimensional map based on azimuth and elevation, where points in the map correspond to vertical and azimuth angles from a point of view of a radar measurement device generating the radar information, and values of the points correspond to distance of a first observable object in a direction of the angles corresponding to the points.
[00175] Example 32. The apparatus according to example 26, wherein the real scene depth map comprises a three-dimensional map generated at least by assigning a first value to points of a coordinate system, corresponding to the physical area, where there is free space, and assigning a second value to points where objects are present in the coordinate system, where a function is used to determine whether an object does or does not exist at individual directions based on azimuth, elevation, and range.
[00176] Example 33. The apparatus according to example 26, wherein the generative adversarial network is trained using sensory information comprising one or more of the following: radar information, Lidar information, backscatter tracking information, or localization information.
[00177] Example 34. The apparatus according to example 30, wherein the generative adversarial network is trained using sensory information comprising multiple ones of the following: the radar information, the Lidar information, the backscatter tracking information, or the localization information, and the multiple sensory information are fused together and fed into the generator neural network, or fed in as separate layers to the generator neural network.
[00178] Example 35. The apparatus according to any one of examples 26 to 31, wherein the means are further configured for performing training an object detection neural network on identifying classes of objects using images from the physical area.
[00179] Example 36. The apparatus according to example 32, wherein the means are further configured for performing:
[00180] performing training a cascade network of the trained generator neural network and the object detection neural network, wherein output of the trained generator neural network is input to the object detection neural network, the training the cascade network of the trained generator neural network performed to adjust one or more tunable layers in one or both of the trained generator neural network or the object detection neural network.
[00181] Example 37. The apparatus according to example 33, wherein the performing training the cascade network creates a trained cascade network, and wherein the means are further configured for performing another training of the cascade network using measurement data from one or more sensing entities of the physical area, the other training performed to adjust one or more tunable layers in one or both of the trained generator neural network or the object detection neural network.
[00182] Example 38. The apparatus of any preceding apparatus example, wherein the means comprises:
[00183] at least one processor; and
[00184] at least one memory storing instructions that, when executed by at least one processor, cause the performance of the apparatus.
[00185] As used in this application, the term “circuitry” may refer to one or more or all of the following:
[00186] (a) hardware-only circuit implementations (such as implementations in only analog and/or digital circuitry) and
[00187] (b) combinations of hardware circuits and software, such as (as applicable):
(i) a combination of analog and/or digital hardware circuit(s) with software/firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions) and
[00188] (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.
[00189] This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and/or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[00190] Embodiments herein may be implemented in software (executed by one or more processors), hardware (e.g., an application specific integrated circuit), or a combination of software and hardware. In an example embodiment, the software (e.g., application logic, an instruction set) is maintained on any one of various conventional computer-readable media. In the context of this document, a “computer-readable medium” may be any media or means that can contain, store, communicate, propagate or transport the instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer, with one example of a computer described and depicted, e.g., in FIG. 1A. A computer-readable medium may comprise a computer-readable storage medium (e.g., memories 125 or other device) that may be any media or means that can contain, store, and/or transport the instructions for use by or in connection with an instruction execution system, apparatus, or device, such as a computer. A computer-readable storage medium does not comprise propagating signals, and therefore may be considered to be non-transitory. The term “non-transitory”, as used herein, is a limitation of the medium itself (i.e., tangible, not a signal) as opposed to a limitation on data storage persistency (e.g., RAM, random access memory, versus ROM, read-only memory).
[00191] If desired, the different functions discussed herein may be performed in a different order and/or concurrently with each other. Furthermore, if desired, one or more of the above-described functions may be optional or may be combined.
[00192] Although various aspects of the invention are set out in the independent claims, other aspects of the invention comprise other combinations of features from the described embodiments and/or the dependent claims with the features of the independent claims, and not solely the combinations explicitly set out in the claims.
[00193] It is also noted herein that while the above describes example embodiments of the invention, these descriptions should not be viewed in a limiting sense. Rather, there are several variations and modifications which may be made without departing from the scope of the present invention as defined in the appended claims.

Claims

What is claimed is:
1. A method, comprising: converting, using a neural network, sensory information of a physical area into one or more images; and performing, using one or more other neural networks, at least object detection on the one or more images to determine an output indicating at least whether or not one or more objects were detected.
2. The method according to claim 1, wherein the performing at least object detection also performs classification of detected objects so the output indicates whether or not one or more objects were detected using likelihood of presence of the one or more objects belonging to an object category.
3. The method according to claim 2, wherein there are multiple categories, and classification outputs multiple likelihoods corresponding to the multiple categories.
4. The method according to any one of claims 2 or 3, wherein the images comprise three- dimensional point clouds where backgrounds have been subtracted.
5. The method according to claim 4, further comprising converting the three-dimensional point clouds where backgrounds have been subtracted to three-dimensional occupancy grids, wherein the performing at least object detection performs at least object detection on the three-dimensional occupancy grids.
6. The method according to any one of claims 2 or 3, wherein the one or more images comprise three-dimensional point clouds with background subtracted and wherein the performing, using the one or more other neural networks, at least object detection on the one or more images to determine an output comprises the following: performing a first detection using a first of the one or more other neural networks to produce first output comprising boundary box detection for objects in a scene, wherein there are multiple bounding box segments for the 3D point clouds; and performing a second detection on the first output using a second of the one or more other neural networks to produce second output comprising the likelihoods of presence of objects belonging to corresponding object categories in individual bounding box segments.
7. The method according to any one of claims 1 to 6, wherein the converting sensory information and performing at least object detection are performed to train the neural network and the one of more other neural networks using the sensory information to adjust one or more tunable layers in one or both of the neural network or the one or more other neural networks.
8. The method according to claim any one of claims 1 to 7, wherein the sensory information comprises radar information developed from one or more base stations in a wireless system.
9. A method, comprising: training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area, the training comprising: creating generated depth maps by the generator neural network based at least on corresponding sensory information of a physical area, the generated depth maps comprising images of the physical area; determining classifications of fake or real by a discriminator neural network based on the generated depth maps and one or more real scene depth maps, the one or more real scene depth maps comprising a multidimensional map of the physical area, where the classifications are routed to the generator neural network and the discriminator neural network; and performing the creating the generated depth maps and the determining classifications until one or more criteria are met based on one or more loss functions; and outputting information that defines a trained generator neural network, the trained generator neural network able to reconstruct images of the physical area based at least on corresponding sensory information of the physical area.
10. The method according to claim 9, wherein the real scene depth map comprises an occupancy grid, or one or multiple two-dimensional depth maps, or both the occupancy grid and the one or multiple two-dimensional depth maps.
11. The method according to claim 9, wherein the sensory information comprises radar information, wherein the real scene depth map comprises a two-dimensional map based on azimuth and elevation, where points in the map correspond to vertical and azimuth angles from a point of view of a radar measurement device generating the radar information, and values of the points correspond to distance of a first observable object in a direction of the angles corresponding to the points.
12. The method according to claim 9, wherein the real scene depth map comprises a three- dimensional map generated at least by assigning a first value to points of a coordinate system, corresponding to the physical area, where there is free space, and assigning a second value to points where objects are present in the coordinate system, where a function is used to determine whether an object does or does not exist at individual directions based on azimuth, elevation, and range.
13. The method according to claim 9, wherein the generative adversarial network is trained using sensory information comprising one or more of the following: radar information, Lidar information, backscatter tracking information, or localization information.
14. The method according to claim 13, wherein the generative adversarial network is trained using sensory information comprising multiple ones of the following: the radar information, the Lidar information, the backscatter tracking information, or the localization information, and the multiple sensory information are fused together and fed into the generator neural network, or fed in as separate layers to the generator neural network.
15. The method according to any one of claims 9 to 14, further comprising training an object detection neural network on identifying classes of objects using images from the physical area.
16. The method according to claim 15, further comprising: performing training a cascade network of the trained generator neural network and the object detection neural network, wherein output of the trained generator neural network is input to the object detection neural network, the training the cascade network of the trained generator neural network performed to adjust one or more tunable layers in one or both of the trained generator neural network or the object detection neural network.
17. The method according to claim 16, wherein the performing training the cascade network creates a trained cascade network, and the method further comprises performing another training of the cascade network using measurement data from one or more sensing entities of the physical area, the other training performed to adjust one or more tunable layers in one or both of the trained generator neural network or the object detection neural network.
18. A computer program, comprising instructions for performing the methods of any of claims 1 to 17, when the computer program is run on an apparatus.
19. The computer program according to claim 18, wherein the computer program is a computer program product comprising a computer-readable medium bearing instructions embodied therein for use with the apparatus.
20. The computer program according to claim 18, wherein the computer program is directly loadable into an internal memory of the apparatus.
21. An apparatus, comprising means for performing: converting, using a neural network, sensory information of a physical area into one or more images; and performing, using one or more other neural networks, at least object detection on the one or more images to determine an output indicating at least whether or not one or more objects were detected.
22. The apparatus according to claim 21, wherein the performing at least object detection also performs classification of detected objects so the output indicates whether or not one or more objects were detected using likelihood of presence of the one or more objects belonging to an object category.
23. The apparatus according to claim 22, wherein there are multiple categories, and classification outputs multiple likelihoods corresponding to the multiple categories.
24. The apparatus according to any one of claims 22 or 23, wherein the images comprise three-dimensional point clouds where backgrounds have been subtracted.
25. The apparatus according to claim 24, wherein the means are further configured for performing converting the three-dimensional point clouds where backgrounds have been subtracted to three-dimensional occupancy grids, wherein the performing at least object detection performs at least object detection on the three-dimensional occupancy grids.
26. The apparatus according to any one of claims 22 or 23, wherein the one or more images comprise three-dimensional point clouds with background subtracted and wherein the performing, using the one or more other neural networks, at least object detection on the one or more images to determine an output comprises the following: performing a first detection using a first of the one or more other neural networks to produce first output comprising boundary box detection for objects in a scene, wherein there are multiple bounding box segments for the 3D point clouds; and performing a second detection on the first output using a second of the one or more other neural networks to produce second output comprising the likelihoods of presence of objects belonging to corresponding object categories in individual bounding box segments.
27. The apparatus according to any one of claims 21 to 26, wherein the converting sensory information and performing at least object detection are performed to train the neural network and the one of more other neural networks using the sensory information to adjust one or more tunable layers in one or both of the neural network or the one or more other neural networks.
28. The apparatus according to claim any one of claims 21 to 27, wherein the sensory information comprises radar information developed from one or more base stations in a wireless system.
29. An apparatus, comprising means for performing: training a generative adversarial network comprising a generator neural network and a discriminator neural network for reconstruction of images of a physical area, the training comprising: creating generated depth maps by the generator neural network based at least on corresponding sensory information of a physical area, the generated depth maps comprising images of the physical area; determining classifications of fake or real by a discriminator neural network based on the generated depth maps and one or more real scene depth maps, the one or more real scene depth maps comprising a multidimensional map of the physical area, where the classifications are routed to the generator neural network and the discriminator neural network; and performing the creating the generated depth maps and the determining classifications until one or more criteria are met based on one or more loss functions; and outputting information that defines a trained generator neural network, the trained generator neural network able to reconstruct images of the physical area based at least on corresponding sensory information of the physical area.
30. The apparatus according to claim 29, wherein the real scene depth map comprises an occupancy grid, or one or multiple two-dimensional depth maps, or both the occupancy grid and the one or multiple two-dimensional depth maps.
31. The apparatus according to claim 29, wherein the sensory information comprises radar information, wherein the real scene depth map comprises a two-dimensional map based on azimuth and elevation, where points in the map correspond to vertical and azimuth angles from a point of view of a radar measurement device generating the radar information, and values of the points correspond to distance of a first observable object in a direction of the angles corresponding to the points.
32. The apparatus according to claim 29, wherein the real scene depth map comprises a three-dimensional map generated at least by assigning a first value to points of a coordinate system, corresponding to the physical area, where there is free space, and assigning a second value to points where objects are present in the coordinate system, where a function is used to determine whether an object does or does not exist at individual directions based on azimuth, elevation, and range.
33. The apparatus according to claim 29, wherein the generative adversarial network is trained using sensory information comprising one or more of the following: radar information, Lidar information, backscatter tracking information, or localization information.
34. The apparatus according to claim 33, wherein the generative adversarial network is trained using sensory information comprising multiple ones of the following: the radar information, the Lidar information, the backscatter tracking information, or the localization information, and the multiple sensory information are fused together and fed into the generator neural network, or fed in as separate layers to the generator neural network.
35. The apparatus according to any one of claims 29 to 34, wherein the means are further configured for performing training an object detection neural network on identifying classes of objects using images from the physical area.
36. The apparatus according to claim 35, wherein the means are further configured for performing: performing training a cascade network of the trained generator neural network and the object detection neural network, wherein output of the trained generator neural network is input to the object detection neural network, the training the cascade network of the trained generator neural network performed to adjust one or more tunable layers in one or both of the trained generator neural network or the object detection neural network.
37. The apparatus according to claim 36, wherein the performing training the cascade network creates a trained cascade network, and wherein the means are further configured for performing another training of the cascade network using measurement data from one or more sensing entities of the physical area, the other training performed to adjust one or more tunable layers in one or both of the trained generator neural network or the object detection neural network.
38. The apparatus of any preceding apparatus claim, wherein the means comprises: at least one processor; and at least one memory storing instructions that, when executed by at least one processor, cause the performance of the apparatus.
EP23920245.0A 2023-01-30 2023-01-30 Object detection accuracy improvement in wireless communications systems at least using radar sensing Pending EP4659197A1 (en)

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/US2023/061569 WO2024163007A1 (en) 2023-01-30 2023-01-30 Object detection accuracy improvement in wireless communications systems at least using radar sensing

Publications (1)

Publication Number Publication Date
EP4659197A1 true EP4659197A1 (en) 2025-12-10

Family

ID=92147340

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23920245.0A Pending EP4659197A1 (en) 2023-01-30 2023-01-30 Object detection accuracy improvement in wireless communications systems at least using radar sensing

Country Status (2)

Country Link
EP (1) EP4659197A1 (en)
WO (1) WO2024163007A1 (en)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP4712514A1 (en) * 2024-09-16 2026-03-18 Deutsche Telekom AG Method for sensing in a mobile communication network

Family Cites Families (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP3709216B1 (en) * 2018-02-09 2023-08-02 Bayerische Motoren Werke Aktiengesellschaft Methods and apparatuses for object detection in a scene represented by depth data of a range detection sensor and image data of a camera
EP3525000B1 (en) * 2018-02-09 2021-07-21 Bayerische Motoren Werke Aktiengesellschaft Methods and apparatuses for object detection in a scene based on lidar data and radar data of the scene
WO2021016596A1 (en) * 2019-07-25 2021-01-28 Nvidia Corporation Deep neural network for segmentation of road scenes and animate object instances for autonomous driving applications
US11532168B2 (en) * 2019-11-15 2022-12-20 Nvidia Corporation Multi-view deep neural network for LiDAR perception

Also Published As

Publication number Publication date
WO2024163007A1 (en) 2024-08-08

Similar Documents

Publication Publication Date Title
Cui et al. Sensing-assisted high reliable communication: A transformer-based beamforming approach
CN110675418B (en) Target track optimization method based on DS evidence theory
Shen et al. ELLK-Net: An efficient lightweight large kernel network for SAR ship detection
CN110689562A (en) Trajectory loop detection optimization method based on generation of countermeasure network
CN114898355A (en) Method and system for self-supervised learning of body-to-body movements for autonomous driving
Ohta et al. Point cloud-based proactive link quality prediction for millimeter-wave communications
Theagarajan et al. Integrating deep learning-based data driven and model-based approaches for inverse synthetic aperture radar target recognition
Liao et al. VI-NeRF-SLAM: A real-time visual–inertial SLAM with NeRF mapping
Wang et al. Fast detection and obstacle avoidance on UAVs using lightweight convolutional neural network based on the fusion of radar and camera: X. Wang et al.
WO2024163007A1 (en) Object detection accuracy improvement in wireless communications systems at least using radar sensing
Jin et al. Radar and lidar deep fusion: Providing doppler contexts to time-of-flight lidar
Chaturvedi et al. Small object detection using retinanet with hybrid anchor box hyper tuning using interface of Bayesian mathematics
Nie et al. An efficient nocturnal scenarios beamforming based on multi-modal enhanced by object detection
Wang et al. Multi-modal environmental information sensing based path loss prediction for V2I communications
Ding et al. DCR-YOLO: An enhanced anti-UAV detection method based on triple collaborative optimization strategy
CN121305452A (en) A method and system for fusing and sensing multi-source navigation data for ship pilotage in low visibility environments
Tang et al. SAR-ShipSwin: enhancing SAR ship detection with robustness in complex environment: J. Tang et al.
CN120475399A (en) A millimeter-wave beam tracking method integrating vision and 3D point cloud perception
Kim et al. PillarGen: Enhancing radar point cloud density and quality via pillar-based point generation network
Park et al. Cross-modal knowledge distillation for efficient radar-only beam prediction in mmwave communications
Lu et al. Path loss prediction for vehicle-to-infrastructure communications via synesthesia of machines (SoM)
Chaabane Multi-source Data Fusion to Enhance Wireless Communication Beyond 5G for Smart City Transformation
Kwon et al. A data augmentation approach to 28GHz path loss modeling using CNNs
Costa et al. PerceptNet-V2X: Perception Network for Vehicle to Everything Scenarios in Autonomous Driving
Ren et al. T-UNet: A novel TC-based point cloud super-resolution model for mechanical lidar

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250901

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR