EP4675590A1 - Method for estimating a collision hazard level in a traffic environment - Google Patents

Method for estimating a collision hazard level in a traffic environment

Info

Publication number
EP4675590A1
EP4675590A1 EP24185927.1A EP24185927A EP4675590A1 EP 4675590 A1 EP4675590 A1 EP 4675590A1 EP 24185927 A EP24185927 A EP 24185927A EP 4675590 A1 EP4675590 A1 EP 4675590A1
Authority
EP
European Patent Office
Prior art keywords
traffic environment
road user
detected road
view
point
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24185927.1A
Other languages
German (de)
French (fr)
Inventor
Hazem ABDELKAWY
Luke PALMER
Petar PALASEK
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Glimpse Technology Ltd
Toyota Motor Corp
Original Assignee
Glimpse Technology Ltd
Toyota Motor Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Glimpse Technology Ltd, Toyota Motor Corp filed Critical Glimpse Technology Ltd
Priority to EP24185927.1A priority Critical patent/EP4675590A1/en
Publication of EP4675590A1 publication Critical patent/EP4675590A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G08SIGNALLING
    • G08GTRAFFIC CONTROL SYSTEMS
    • G08G1/00Traffic control systems for road vehicles
    • G08G1/16Anti-collision systems
    • G08G1/164Centralised systems, e.g. external to vehicles
    • GPHYSICS
    • G08SIGNALLING
    • G08GTRAFFIC CONTROL SYSTEMS
    • G08G1/00Traffic control systems for road vehicles
    • G08G1/01Detecting movement of traffic to be counted or controlled
    • G08G1/0104Measuring and analyzing of parameters relative to traffic conditions
    • G08G1/0108Measuring and analyzing of parameters relative to traffic conditions based on the source of data
    • G08G1/0116Measuring and analyzing of parameters relative to traffic conditions based on the source of data from roadside infrastructure, e.g. beacons
    • GPHYSICS
    • G08SIGNALLING
    • G08GTRAFFIC CONTROL SYSTEMS
    • G08G1/00Traffic control systems for road vehicles
    • G08G1/01Detecting movement of traffic to be counted or controlled
    • G08G1/04Detecting movement of traffic to be counted or controlled using optical or ultrasonic detectors
    • GPHYSICS
    • G08SIGNALLING
    • G08GTRAFFIC CONTROL SYSTEMS
    • G08G1/00Traffic control systems for road vehicles
    • G08G1/16Anti-collision systems
    • G08G1/166Anti-collision systems for active traffic, e.g. moving vehicles, pedestrians, bikes

Definitions

  • the present invention relates to the field of traffic monitoring, and more particularly, to a method and an apparatus for estimating a collision hazard level in a traffic environment.
  • ADAS Advanced Driver Assistance Systems
  • driver assistance systems have been designed to guide the driver and assist him in risk anticipation by automating, improving, or adapting some or all of the tasks related to driving a vehicle.
  • driver assistance systems rely generally on sensors to assist the driver throughout their journey and often communicate with a server to receive real-time information along the vehicle's route.
  • Modern driver assistance systems also provide additional information related to the route, such as weather conditions, traffic congestion, duration, and disruptions on the road (diversions or lane closures, for example).
  • the driver assistance systems can also be configured to provide information, including regulatory indications, such as speed limits or the presence of checkpoints.
  • the driver assistance systems have been designed to enhance the driver's ability to perceive and react to important traffic participants by processing information from a range of sensors mounted on the vehicle and by providing real-time feedback and, in some cases, automating vehicle responses to potential hazards.
  • sensors typically include radar, LiDAR, ultrasonic sensors, and cameras that capture visual data, identifying traffic signs, lane markings, pedestrians, and other vehicles.
  • ego-centric driving data As more vehicles come equipped with these cameras, the amount of ego-centric driving data, which is footage captured from the driver's perspective, is rapidly growing. This data captures a wide range of traffic scenarios, presenting an opportunity to enhance driving assistance systems. In particular, utilizing machine learning and computer vision techniques on ego-centric driving data can significantly improve the ability to detect and respond to traffic hazard.
  • ego-centric data collected from the driver's viewpoint, provides valuable but limited information about the surrounding environment. While it offers insights into immediate hazards within the driver's field of vision, it may lack context about events occurring outside this narrow perspective. As a result, hazard detection systems relying solely on ego-centric data may overlook potential risks that are not directly observable from the driver's viewpoint, such as obstacles obscured by other vehicles or pedestrians outside the immediate field of vision.
  • the object of the present invention is to at least substantially remedy the above-mentioned drawbacks.
  • the method is referred to hereinafter as the estimating method.
  • “a”, “an”, and “the” are intended to refer to “at least one” or “each” and to include plural forms as well. Conversely, generic plural forms include the singular.
  • the invention focuses on enhancing safety by providing an accurate assessment of collision risks in a traffic environment.
  • This first representation is then used to generate a second representation from the perspective of each detected road user, providing a different view of the traffic environment than the first representation.
  • the first representation may thus be considered as a viewpoint that is invariant or equivariant, ensuring stability and consistency across different perspectives within the traffic environment.
  • the first representation serves then as a reliable base, unaffected by varying viewpoints initially.
  • Processing these second representations for example, using a model trained on readily accessible egocentric viewpoint data, enables the assessment of risk levels associated with other road users as seen from the viewpoint of a selected road user.
  • This approach leverages valuable information captured by cameras, such as infrastructure-mounted cameras, to enhance the understanding of the traffic environment.
  • constructing the first representation of the traffic environment comprises generating a bird's eye view of the traffic environment which comprises the predetermined feature of each detected road user, and processing the bird's eye view to construct a scene graph as the first representation of the traffic environment.
  • a bird's eye view is a perspective that shows the traffic environment as if seen from above.
  • This top-down view provides a comprehensive overview of the entire traffic environment, especially capturing spatial relationships and the arrangement of the objects.
  • the scene graph constructed based on this perspective, organizes the detected road users and other predetermined feature(s) of the environment into a graphical format. This structured representation facilitates further processing and analysis, enabling to comprehend and interact with the traffic environment more effectively.
  • generating the bird's eye view of the traffic environment comprises establishing a matrix correspondence between pixels of the traffic environment captured from the first point of view and a ground plane onto which each detected road user in the traffic environment is projected from top-down perspective.
  • the predetermined feature of each detected road user is its position, and/or its class label, and/or its pose, and/or its speed, and/or its acceleration, and/or its trajectory.
  • the trajectory of each detected road user is estimated by processing successive images of the traffic environment captured from the first point of view in real-time.
  • the second representation of the traffic environment is an egocentric scene graph corresponding to a specific viewing point, viewing direction, and field of view.
  • an egocentric scene graph refers to a structured representation of the surrounding environment from the perspective of the road user (or "ego" position).
  • the detected road user is a vehicle or a pedestrian.
  • the traffic environment is an intersection with vehicles and pedestrians crossing, and/or a highway interchange with multiple lanes of traffic, and/or a parking lot entrance.
  • the present invention further relates to a computer program set including instructions for executing the steps of the above-described method when said program set is executed by at least one computer.
  • This program set can use any programming language and take the form of source code, object code, or a code intermediate between source code and object code, such as a partially compiled form or any other desirable form.
  • the present invention further relates to a recording medium readable by at least one computer and having recorded thereon at least one computer program including instructions for executing the steps of the above-described method.
  • the present invention further relates to an apparatus for estimating a collision hazard level in a traffic environment captured from a first point of view, comprising:
  • the apparatus referred to hereinafter as the estimating apparatus may be configured to carry out the above-mentioned estimating method and may have part or all of the above-described features.
  • the estimating apparatus may have the hardware structure of a computer.
  • the first point view is that of an infrastructure camera.
  • Figure 1 shows a block diagram of an estimating apparatus 10 comprising a constructing module 11 and a processing module 12 according to embodiments of the present invention.
  • the estimating apparatus 10 may be configured to estimate a collision hazard level in a traffic environment captured from a first point of view, for example by using an infrastructure-mounted camera, to provide a more comprehensive and accurate assessment of traffic hazards.
  • the estimating apparatus 10 may comprise an electronic circuit, a processor (shared, dedicated, or group), a combinational logic circuit, a memory that executes one or more software programs, and/or other suitable components that provide the described functionality.
  • the estimating apparatus 10 may be a computing device.
  • the estimating apparatus 10 may be connected to a memory, which may store data, e.g., at least one computer program, which, when executed, carries out the estimating method according to the present invention.
  • the memory can be a ROM (for "Read Only Memory”), a CD-ROM, a microelectronic circuit ROM, or in the form of magnetic storage means, for example, a diskette (floppy disk), a flash disk, a SSD (for "Solid State Drive”) or a hard disk.
  • ROM Read Only Memory
  • CD-ROM Compact Disc-ROM
  • microelectronic circuit ROM or in the form of magnetic storage means, for example, a diskette (floppy disk), a flash disk, a SSD (for "Solid State Drive”) or a hard disk.
  • the memory can be an integrated circuit in which the program is incorporated, the circuit being adapted to execute the estimating method or to be used in its execution.
  • the constructing module 11 is a module configured to construct a first representation of the traffic environment that is captured from the first point of view.
  • the captured traffic environment can be in form of raw RGB image continuous stream: I t : RGB image at time t , where I t indicates the image corresponding to a given time point t.
  • Track() such as DeepSORT (for "Deep Simple Online and Realtime Tracking” algorithm), or ByteTrack algorithm, or StrongSORT algorithm
  • the trajectory of each detected road user is thus estimated by the tracker mechanism that is configured to process successive images of the traffic environment from the first point of view in real-time.
  • the constructing module 11 is then configured to generate a bird's eye view of the traffic environment which comprises the predetermined feature(s) of each detected road user, such as its position, and/or its class label, and/or its pose, and/or its speed, and/or its acceleration, and/or its trajectory.
  • the constructing module 11 is configured to establish a matrix correspondence between pixels of the traffic environment captured from the first point of view and a ground plane onto which each detected road user in the traffic environment is projected from top-down perspective.
  • R and t define the camera extrinsic parameters
  • K represents the camera intrinsic parameters
  • the calculation involves projecting a defined set of three points p 0 w , p 1 w , p 2 w , onto a plane in the world coordinate system, where the z-coordinate is set to 0, using the extrinsic camera parameters.
  • the constructing module 11 is further configured to process the bird's eye view to construct a scene graph as the first representation of the traffic environment, as observed by the camera (the infrastructure camera here).
  • the constructing module 11 is further configured to generate a second representation of the traffic environment from a second point view of each detected road user, based on the first representation of the traffic environment.
  • the second representation of the traffic environment is an egocentric scene graph corresponding to a specific viewing point, viewing direction, and field of view. Accordingly, the egocentric scene graph refers to a structured representation of the surrounding environment from the perspective of the road user (or "ego" position).
  • the resulting egocentric scene graph ICSG′′′ t is the output of the Prune function as it was viewed from the point p v BEV facing towards the point p d BEV , with the field of view of ⁇ FOV degrees.
  • SG t p ⁇ BEV p d BEV ⁇ FOV ICSG t ′′′ ,
  • the constructing module 11 may be implemented as software running on the estimating apparatus 10 or may be implemented partially or fully as a hardware element of the estimating apparatus 10.
  • the processing module 12 is configured to process all the second representations of the traffic environment to estimate the collision hazard level between each detected road user and every other detected road user. More specifically, the processing module 12 is configured to send all the second representations to a typical model trained on data from a perspective of the road user.
  • I SG t EstimateHazard SG t , wherein I SGt corresponds to the hazard scores for the detected objects included in the egocentric scene graph SG t , noting that each of its nodes corresponds to a detected object. Accordingly, this results in thorough understanding of the traffic environment by accurately assessing the risk levels associated with each road user in traffic. Such comprehension enables to identify potential hazards effectively, as it considers the positions, and movements of all individuals and vehicles involved.
  • FIG. 2 is a flowchart of the estimating method 20 according to an embodiment of the present invention.
  • the estimating method is performed for estimating a collision hazard level in the traffic environment captured from the first point of view.
  • the estimating method 20 comprises a step 21 performed by the constructing module 11 to construct the first representation of the traffic environment comprising the predetermined feature(s) of each detected road user in the traffic environment.
  • the step 21 comprises a sub-step 211 performed by the constructing module 11 to generate the bird's eye view of the traffic environment by establishing a matrix correspondence between pixels of the traffic environment captured from the first point of view and the ground plane onto which each detected road user in the traffic environment is projected from top-down perspective, as explained above.
  • Step 21 further comprises a sub-step 212 performed by the constructing module 11 to construct the scene graph as the first representation of the traffic environment by processing the bird's eye view generated during sub-step 211.
  • the estimating method 20 comprises a step 22 performed by the constructing module 11 to generate the second representation of the traffic environment from the second point of view of each detected road user as explained above.
  • the estimating method 20 comprises a step 23 performed by the processing module 12 to process all the second representations of the traffic environment to estimate the collision hazard level between each detected road user and every other detected road user.

Landscapes

  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Chemical & Material Sciences (AREA)
  • Analytical Chemistry (AREA)
  • Traffic Control Systems (AREA)

Abstract

A method (20) for estimating a collision hazard level in a traffic environment captured from a first point of view, comprising:- constructing (21) a first representation of the traffic environment comprising at least one predetermined feature of each detected road user in the traffic environment to generate a second representation of the traffic environment from a second point of view of each detected road user; and- processing (23) all the second representations of the traffic environment to estimate a collision hazard level between each detected road user and every other detected road user.

Description

    BACKGROUND OF THE INVENTION 1. Field of the invention
  • The present invention relates to the field of traffic monitoring, and more particularly, to a method and an apparatus for estimating a collision hazard level in a traffic environment.
  • 2. Description of Related Art
  • The automation of vehicles represents a major challenge for automotive safety and driving optimization. In an effort to eliminate human error while driving a vehicle, Advanced Driver Assistance Systems (ADAS), also known as "driver assistance systems", have been designed to guide the driver and assist him in risk anticipation by automating, improving, or adapting some or all of the tasks related to driving a vehicle.
  • These driver assistance systems rely generally on sensors to assist the driver throughout their journey and often communicate with a server to receive real-time information along the vehicle's route. Modern driver assistance systems also provide additional information related to the route, such as weather conditions, traffic congestion, duration, and disruptions on the road (diversions or lane closures, for example). The driver assistance systems can also be configured to provide information, including regulatory indications, such as speed limits or the presence of checkpoints.
  • Recently, the driver assistance systems have been designed to enhance the driver's ability to perceive and react to important traffic participants by processing information from a range of sensors mounted on the vehicle and by providing real-time feedback and, in some cases, automating vehicle responses to potential hazards. These sensors typically include radar, LiDAR, ultrasonic sensors, and cameras that capture visual data, identifying traffic signs, lane markings, pedestrians, and other vehicles.
  • However, detecting potential hazards has become increasingly challenging due to the growing complexity of traffic patterns and environments as roadways become more congested and diverse, with a multitude of vehicles, pedestrians, cyclists, and various infrastructure elements. In addition, vehicles may interact with each other in unpredictable ways, making it difficult to anticipate potential collisions or dangerous situations, and the presence of vulnerable road users adds another layer of complexity, as their behavior can be highly variable and difficult to predict.
  • As more vehicles come equipped with these cameras, the amount of ego-centric driving data, which is footage captured from the driver's perspective, is rapidly growing. This data captures a wide range of traffic scenarios, presenting an opportunity to enhance driving assistance systems. In particular, utilizing machine learning and computer vision techniques on ego-centric driving data can significantly improve the ability to detect and respond to traffic hazard.
  • However, relying solely on this data restricts the scope and effectiveness of hazard detection for several reasons. In fact, ego-centric data, collected from the driver's viewpoint, provides valuable but limited information about the surrounding environment. While it offers insights into immediate hazards within the driver's field of vision, it may lack context about events occurring outside this narrow perspective. As a result, hazard detection systems relying solely on ego-centric data may overlook potential risks that are not directly observable from the driver's viewpoint, such as obstacles obscured by other vehicles or pedestrians outside the immediate field of vision.
  • To address these limitations and enhance the effectiveness of hazard detection, it is advantageous to integrate multiple data sources and viewpoints, including by infrastructure-mounted cameras, to provide a more comprehensive and accurate assessment of traffic hazards.
  • A major problem arises when trying to apply these models trained on ego-centric data to these different viewpoints, such as those captured by infrastructure-mounted cameras. These models often rely on visual features that are unique to the driver's perspective, which do not translate well to other angles. This discrepancy reduces the model's effectiveness when used with cameras installed at different locations, such as traffic lights or road signs.
  • Accordingly, the existing solutions fall short in addressing the need for traffic monitoring and hazard detection across diverse complex environments.
  • SUMMARY
  • The object of the present invention is to at least substantially remedy the above-mentioned drawbacks.
  • In this respect, the present invention relates to a method for estimating a collision hazard level in a traffic environment captured from a first point of view, comprising:
    • constructing a first representation of the traffic environment comprising at least one predetermined feature of each detected road user in the traffic environment to generate a second representation of the traffic environment from a second point of view of each detected road user; and
    • processing all the second representations of the traffic environment to estimate a collision hazard level between each detected road user and every other detected road user.
  • For conciseness, the method is referred to hereinafter as the estimating method. Further, as used herein, for conciseness and unless the context indicates otherwise, "a", "an", and "the" are intended to refer to "at least one" or "each" and to include plural forms as well. Conversely, generic plural forms include the singular.
  • The invention focuses on enhancing safety by providing an accurate assessment of collision risks in a traffic environment.
  • To this end, the first representation of the traffic environment is constructed from a first point of view to get a comprehensive understanding of the traffic situation, for example from an infrastructure-mounted camera which is typically a stationary camera installed on poles, buildings, or other structures in the traffic environment. The images captured by such camera can be used to monitor and analyze traffic flow, detect accidents, and gather data on vehicle movements.
  • This first representation is then used to generate a second representation from the perspective of each detected road user, providing a different view of the traffic environment than the first representation. The first representation may thus be considered as a viewpoint that is invariant or equivariant, ensuring stability and consistency across different perspectives within the traffic environment. The first representation serves then as a reliable base, unaffected by varying viewpoints initially.
  • Processing these second representations, for example, using a model trained on readily accessible egocentric viewpoint data, enables the assessment of risk levels associated with other road users as seen from the viewpoint of a selected road user. This approach leverages valuable information captured by cameras, such as infrastructure-mounted cameras, to enhance the understanding of the traffic environment.
  • Optionally, constructing the first representation of the traffic environment comprises generating a bird's eye view of the traffic environment which comprises the predetermined feature of each detected road user, and processing the bird's eye view to construct a scene graph as the first representation of the traffic environment.
  • A bird's eye view is a perspective that shows the traffic environment as if seen from above. This top-down view provides a comprehensive overview of the entire traffic environment, especially capturing spatial relationships and the arrangement of the objects. The scene graph, constructed based on this perspective, organizes the detected road users and other predetermined feature(s) of the environment into a graphical format. This structured representation facilitates further processing and analysis, enabling to comprehend and interact with the traffic environment more effectively.
  • Optionally, generating the bird's eye view of the traffic environment comprises establishing a matrix correspondence between pixels of the traffic environment captured from the first point of view and a ground plane onto which each detected road user in the traffic environment is projected from top-down perspective.
  • Optionally, the predetermined feature of each detected road user is its position, and/or its class label, and/or its pose, and/or its speed, and/or its acceleration, and/or its trajectory.
  • Optionally, the trajectory of each detected road user is estimated by processing successive images of the traffic environment captured from the first point of view in real-time.
  • Optionally, the second representation of the traffic environment is an egocentric scene graph corresponding to a specific viewing point, viewing direction, and field of view.
  • It should be noted that an egocentric scene graph refers to a structured representation of the surrounding environment from the perspective of the road user (or "ego" position).
  • Optionally, the detected road user is a vehicle or a pedestrian.
  • Optionally, processing all the second representations of the traffic environment comprises sending them to a model trained on data from a perspective of the road user in order to estimate said collision hazard level.
  • Optionally, the traffic environment is an intersection with vehicles and pedestrians crossing, and/or a highway interchange with multiple lanes of traffic, and/or a parking lot entrance.
  • The present invention further relates to a computer program set including instructions for executing the steps of the above-described method when said program set is executed by at least one computer.
  • This program set can use any programming language and take the form of source code, object code, or a code intermediate between source code and object code, such as a partially compiled form or any other desirable form.
  • The present invention further relates to a recording medium readable by at least one computer and having recorded thereon at least one computer program including instructions for executing the steps of the above-described method.
  • The present invention further relates to an apparatus for estimating a collision hazard level in a traffic environment captured from a first point of view, comprising:
    • a constructing module configured to construct a first representation of the traffic environment comprising at least one predetermined feature of each detected road user in the traffic environment to generate a second representation of the traffic environment from a second point of view of each detected road user; and
    • a processing module configured to process all the second representations of the traffic environment to estimate a collision hazard level between each detected road user and every other detected road user.
  • The apparatus referred to hereinafter as the estimating apparatus, may be configured to carry out the above-mentioned estimating method and may have part or all of the above-described features. The estimating apparatus may have the hardware structure of a computer.
  • Optionally, the first point view is that of an infrastructure camera.
  • BRIEF DESCRIPTION OF THE DRAWINGS
  • Features, advantages, and technical and industrial significance of exemplary embodiments of the invention will be described below with reference to the accompanying drawings, in which like signs denote like elements, and wherein:
    • FIG. 1 is a block diagram of an estimating apparatus according to an embodiment of the present invention; and
    • FIG.2 is a flowchart of an estimating method according to an embodiment of the present invention.
    DETAILED DESCRIPTION OF EMBODIMENTS
  • Figure 1 shows a block diagram of an estimating apparatus 10 comprising a constructing module 11 and a processing module 12 according to embodiments of the present invention.
  • The estimating apparatus 10 may be configured to estimate a collision hazard level in a traffic environment captured from a first point of view, for example by using an infrastructure-mounted camera, to provide a more comprehensive and accurate assessment of traffic hazards.
  • The estimating apparatus 10 may comprise an electronic circuit, a processor (shared, dedicated, or group), a combinational logic circuit, a memory that executes one or more software programs, and/or other suitable components that provide the described functionality. In other words, the estimating apparatus 10 may be a computing device. The estimating apparatus 10 may be connected to a memory, which may store data, e.g., at least one computer program, which, when executed, carries out the estimating method according to the present invention.
  • For example, the memory can be a ROM (for "Read Only Memory"), a CD-ROM, a microelectronic circuit ROM, or in the form of magnetic storage means, for example, a diskette (floppy disk), a flash disk, a SSD (for "Solid State Drive") or a hard disk.
  • Alternatively, the memory can be an integrated circuit in which the program is incorporated, the circuit being adapted to execute the estimating method or to be used in its execution.
  • The constructing module 11 is a module configured to construct a first representation of the traffic environment that is captured from the first point of view. The captured traffic environment can be in form of raw RGB image continuous stream: I t : RGB image at time t , where I t indicates the image corresponding to a given time point t.
  • Then, within each image I t, road users, such as pedestrians and vehicles are detected and identified by the constructing module 11. For this purpose, well-known object detection algorithms, symbolized by the function "Detect()", such as YOLO (for "You Only Look Once"), R-CNN (for "Region-based Convolutional Neural Network") and DETR (for "Detection Transformer"), can be applied to each image as follows: O t = Detect I t . Here, Ot corresponds to a set of detect objects in the image I t.
  • The constructing module 11 is then configured to apply a well-known tracking mechanism, symbolized by the function "Track()", such as DeepSORT (for "Deep Simple Online and Realtime Tracking" algorithm), or ByteTrack algorithm, or StrongSORT algorithm, in order to maintain a consistent identity for each detected object across consecutive frames: T t = Track O t T t 1 , wherein Tt represents the tracked objects at time t, given the previously tracked objects Tt-1.
  • In other words, the trajectory of each detected road user is thus estimated by the tracker mechanism that is configured to process successive images of the traffic environment from the first point of view in real-time.
  • The constructing module 11 is then configured to generate a bird's eye view of the traffic environment which comprises the predetermined feature(s) of each detected road user, such as its position, and/or its class label, and/or its pose, and/or its speed, and/or its acceleration, and/or its trajectory.
  • To this end, the constructing module 11 is configured to establish a matrix correspondence between pixels of the traffic environment captured from the first point of view and a ground plane onto which each detected road user in the traffic environment is projected from top-down perspective.
  • For example, a transformation matrix T that maps a set of N points P' in the image coordinate system onto the ground plane n▪ p = d is defined as: T 4,3 = R 3,3 t 3,1 0 1,3 1 4,4 1 d 0 0 0 d 0 0 0 d n x n y n z 4,3 K 3,3 1 ,
  • In this context, R and t define the camera extrinsic parameters, K represents the camera intrinsic parameters, and n denotes the normal vector of the ground plane. This normal vector is computed as the cross product of two vectors defined by points p 0 c , p 1 c , p 2 c , which are located on the plane in the camera coordinate systems: n = n x n y n z = p 1 c p 0 c × p 2 c p 0 c .
  • The calculation involves projecting a defined set of three points p 0 w , p 1 w , p 2 w , onto a plane in the world coordinate system, where the z-coordinate is set to 0, using the extrinsic camera parameters. This projection yields the matrix PW given by: P w = p 0 w , p 1 w , p 2 w = 0 1 0 0 0 1 0 0 0 1 1 1 , P 3,3 c = p 0 c , p 1 c , p 2 c = R 3,3 t 3,1 0 1,3 1 4,4 P w 4,3 : 3 , : .
  • The calculation involves computing the dot product between the normal vector n and one of the points situated on the plane in the camera system, denoted as: d = n p 0 c .
  • As the z-coordinate of each point will be set to 0, the row responsible for its computation in the transformation matrix T can be omitted. This results in the transformation matrix TBEV, where the element at position [3;3] becomes: T 3,3 BEV = 1 0 0 0 0 1 0 0 0 0 0 1 T 4,3 .
  • Each column of the transformation P 3 N BEVh of the Pi coordinate system onto the ground plane is obtained by multiplying the element at position [3;3] of the transformation matrix TBEV with the corresponding column P 3 N i , as expressed by equation (10): P 3 N BEVh = T 3,3 BEV P 3 N i .
  • After projecting the coordinate system onto the ground plane in homogeneous coordinates, the resulting matrix P 3 N BEVh needs to be unhomogenized. This is achieved by performing element-wise division of the first two rows by the third row as follows: P 2 N BEV = p 0 x p 1 x p N 1 x p 0 y p 1 x p N 1 y = P 3 N BEVh : 2 , : / P 3 N BEVh 3 , : : 2 , : .
  • With this transformation, it is possible to project the positions of the detected objects onto the ground plane (bird's eye view). This allows estimating their x and y coordinates in absolute units, specifically in meters.
  • The constructing module 11 is further configured to process the bird's eye view to construct a scene graph as the first representation of the traffic environment, as observed by the camera (the infrastructure camera here).
  • More specifically, the constructing module 11 is configured to generate the scene graph at a timestamp t as a function of the set of detected objects O t and the estimated transformation matrix TBEV as follows: ICSG t = ConstructGraph O t T BEV .
  • Each detected object of the scene graph ICSGt is thus represented as a node encompassing a set of feature(s) from the detection, including for example the object's ground-plane position P i BEV = P i x , P i y T measured in meters, noting that the weight of an edge between nodes corresponds to the Euclidean distance between the respective objects they represent.
  • The constructing module 11 is further configured to generate a second representation of the traffic environment from a second point view of each detected road user, based on the first representation of the traffic environment. More specifically, the second representation of the traffic environment is an egocentric scene graph corresponding to a specific viewing point, viewing direction, and field of view. Accordingly, the egocentric scene graph refers to a structured representation of the surrounding environment from the perspective of the road user (or "ego" position).
  • To this end, the constructing module 11 is configured to transform the scene graph ICSG into an egocentric scene graph viewed from a location p v BEV on the ground plane, facing in the direction defined by another point on the ground plane, p d BEV , and a field of view angle αFOV as follows: SG t p υ BEV p d BEV α FOV = Prune ICSG t p υ BEV p d BEV α FOV .
  • The successive steps of the function "Prune()", processed by the constructing module 11, include:
    1. (1) defining a vector v between the points p v BEV (the viewing point) and p d BEV (the viewing direction);
    2. (2) constructing a viewing cone defined by two vectors, v1 and vr, where v1 is the vector v rotated around the point p v BEV by FOV 2 degrees anticlockwise, and vr is the vector v rotated around the point p v BEV by FOV 2 degrees clockwise.
    3. (3) translating all the nodes in the scene graph ICSGt, i.e. the features describing the positions of the nodes, and the viewing cone (the vectors v, v1 and vr) by p v BEV , i.e. moving p v BEV to the origin; p v BEV = 0,0 , noting that this step also transforms all the features with respect to the translation: ICSG t , v , v l , v r = Translate ICSG t v v l v r , p υ BEV .
    4. (4) rotating all the nodes in ISCG't and the viewing cone (the vectors v', v'1 and v'r) around the origin so that the rotated version of v', i.e. v", points in the same direction as the vertical axis, i.e. in the direction of (0, 1), noting that this step should also transform all the features in respect to the rotation: ICSG t " , v " , v l " , v r " = Rotate ICSG t , v , v l , v r , π 2 atan 2 υ y υ x
    5. (5) removing all the nodes (and edges involving those nodes) in ICSG" that do not lie inside the translated and rotated viewing cone, i.e. the nodes that do not lie in between the vectors v"1 and v"r: ICSG t = RemoveNodes ICSG t " v l " v r " .
  • The resulting egocentric scene graph ICSG‴t is the output of the Prune function as it was viewed from the point p v BEV facing towards the point p d BEV , with the field of view of αFOV degrees. SG t p υ BEV p d BEV α FOV = ICSG t ,
  • The constructing module 11 may be implemented as software running on the estimating apparatus 10 or may be implemented partially or fully as a hardware element of the estimating apparatus 10.
  • The processing module 12 is configured to process all the second representations of the traffic environment to estimate the collision hazard level between each detected road user and every other detected road user. More specifically, the processing module 12 is configured to send all the second representations to a typical model trained on data from a perspective of the road user.
  • Such models, particularly Time to Collision-based models, are well-known. The person skilled in the art can refer to the following documentation for example to implement said model:
    • Li, C., Chan, S. H., & Chen, Y. T. (2020, October). Who make drivers stop? towards driver-centric risk assessment: Risk object identification via causal inference. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 10711-10718). IEEE;
    • Gupta, P., Biswas, A., Admoni, H., & Held, D. (2024). Object Importance Estimation using Counterfactual Reasoning for Intelligent Driving. IEEE Robotics and Automation Letters; and
    • Li, C., Chan, S. H., & Chen, Y. T. (2023). Droid: Driver-centric risk object identification. IEEE transactions on pattern analysis and machine intelligence.
  • The second representations are then processed as follows: I SG t = EstimateHazard SG t , wherein I SGt corresponds to the hazard scores for the detected objects included in the egocentric scene graph SGt, noting that each of its nodes corresponds to a detected object. Accordingly, this results in thorough understanding of the traffic environment by accurately assessing the risk levels associated with each road user in traffic. Such comprehension enables to identify potential hazards effectively, as it considers the positions, and movements of all individuals and vehicles involved.
  • The tasks carried out by the estimating apparatus 10 are detailed hereinafter with respect to the corresponding estimating method, an embodiment of which is illustrated in Figure 2.
  • Figure 2 is a flowchart of the estimating method 20 according to an embodiment of the present invention. The estimating method is performed for estimating a collision hazard level in the traffic environment captured from the first point of view.
  • To this end, the estimating method 20 comprises a step 21 performed by the constructing module 11 to construct the first representation of the traffic environment comprising the predetermined feature(s) of each detected road user in the traffic environment. For this purpose, the step 21 comprises a sub-step 211 performed by the constructing module 11 to generate the bird's eye view of the traffic environment by establishing a matrix correspondence between pixels of the traffic environment captured from the first point of view and the ground plane onto which each detected road user in the traffic environment is projected from top-down perspective, as explained above.
  • Step 21 further comprises a sub-step 212 performed by the constructing module 11 to construct the scene graph as the first representation of the traffic environment by processing the bird's eye view generated during sub-step 211.
  • Then, the estimating method 20 comprises a step 22 performed by the constructing module 11 to generate the second representation of the traffic environment from the second point of view of each detected road user as explained above.
  • Finally, the estimating method 20 comprises a step 23 performed by the processing module 12 to process all the second representations of the traffic environment to estimate the collision hazard level between each detected road user and every other detected road user.
  • Although the present disclosure refers to specific exemplary embodiments, modifications may be provided to these examples without departing from the general scope of the invention as defined by the claims. In particular, individual characteristics of the different illustrated/mentioned embodiments may be combined in additional embodiments. Therefore, the description and the drawings should be considered in an illustrative rather than in a restrictive sense.

Claims (13)

  1. A method (20) for estimating a collision hazard level in a traffic environment captured from a first point of view, comprising:
    - constructing (21) a first representation of the traffic environment comprising at least one predetermined feature of each detected road user in the traffic environment to generate a second representation of the traffic environment from a second point of view of each detected road user; and
    - processing (23) all the second representations of the traffic environment to estimate a collision hazard level between each detected road user and every other detected road user.
  2. The method (20) according to claim 1, wherein constructing the first representation of the traffic environment (21) comprises generating (211) a bird's eye view of the traffic environment which comprises the predetermined feature of each detected road user, and processing (212) the bird's eye view to construct a scene graph as the first representation of the traffic environment.
  3. The method (20) according to claim 2, wherein generating (211) the bird's eye view of the traffic environment comprises establishing a matrix correspondence between pixels of the traffic environment captured from the first point of view and a ground plane onto which each detected road user in the traffic environment is projected from top-down perspective.
  4. The method (20) according to any one of claims 1 to 3, wherein the predetermined feature of each detected road user is its position, and/or its class label, and/or its pose, and/or its speed, and/or its acceleration, and/or its trajectory.
  5. The method (20) according to claim 4, wherein the trajectory of each detected road user is estimated by processing successive images of the traffic environment captured from the first point of view in real-time.
  6. The method (20) according to any one of claims 1 to 5, wherein the second representation of the traffic environment is an egocentric scene graph corresponding to a specific viewing point, viewing direction, and field of view.
  7. The method (20) according any one of claims 1 to 6, wherein the detected road user is a vehicle or a pedestrian.
  8. The method (20) according to any one of claims 1 to 7, wherein processing all the second representations of the traffic environment (23) comprises sending them to a model trained on data from a perspective of the road user in order to estimate said collision hazard level.
  9. The method (20) according to any one of claims 1 to 8, wherein the traffic environment is an intersection with vehicles and pedestrians crossing, and/or a highway interchange with multiple lanes of traffic, and/or a parking lot entrance.
  10. A computer program set including instructions for executing the steps of the method (20) of any of claims 1 to 9 when said program set is executed by at least one computer.
  11. A recording medium readable by at least one computer and having recorded thereon at least one computer program including instructions for executing the steps of the method (20) of any one of claims 1 to 9.
  12. An apparatus (10) for estimating a collision hazard level in a traffic environment captured from a first point of view, comprising:
    - a constructing module (11) configured to construct a first representation of the traffic environment comprising at least one predetermined feature of each detected road user in the traffic environment to generate a second representation of the traffic environment from a second point of view of each detected road user; and
    - a processing module (12) configured to process all the second representations of the traffic environment to estimate a collision hazard level between each detected road user and every other detected road user.
  13. The apparatus (10) according to claim 12, wherein the first point of view is that of an infrastructure camera.
EP24185927.1A 2024-07-02 2024-07-02 Method for estimating a collision hazard level in a traffic environment Pending EP4675590A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
EP24185927.1A EP4675590A1 (en) 2024-07-02 2024-07-02 Method for estimating a collision hazard level in a traffic environment

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
EP24185927.1A EP4675590A1 (en) 2024-07-02 2024-07-02 Method for estimating a collision hazard level in a traffic environment

Publications (1)

Publication Number Publication Date
EP4675590A1 true EP4675590A1 (en) 2026-01-07

Family

ID=91782431

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24185927.1A Pending EP4675590A1 (en) 2024-07-02 2024-07-02 Method for estimating a collision hazard level in a traffic environment

Country Status (1)

Country Link
EP (1) EP4675590A1 (en)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20190287395A1 (en) * 2018-03-19 2019-09-19 Derq Inc. Early warning and collision avoidance
US20230082079A1 (en) * 2021-09-16 2023-03-16 Waymo Llc Training agent trajectory prediction neural networks using distillation

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20190287395A1 (en) * 2018-03-19 2019-09-19 Derq Inc. Early warning and collision avoidance
US20230082079A1 (en) * 2021-09-16 2023-03-16 Waymo Llc Training agent trajectory prediction neural networks using distillation

Non-Patent Citations (4)

* Cited by examiner, † Cited by third party
Title
GUPTA, P.BISWAS, A.ADMONI, H.HELD, D.: "Object Importance Estimation using Counterfactual Reasoning for Intelligent Driving", IEEE ROBOTICS AND AUTOMATION LETTERS, 2024
LI, C.CHAN, S. H.CHEN, Y. T.: "Droid: Driver-centric risk object identification", IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2023
LI, C.CHAN, S. H.CHEN, Y. T.: "IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS", October 2020, IEEE, article "Who make drivers stop? towards driver-centric risk assessment: Risk object identification via causal inference", pages: 10711 - 10718
XIAOSONG JIA ET AL: "HDGT: Heterogeneous Driving Graph Transformer for Multi-Agent Trajectory Prediction via Scene Encoding", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 30 April 2022 (2022-04-30), XP091211996 *

Similar Documents

Publication Publication Date Title
US12351184B2 (en) Multiple exposure event determination
US11741696B2 (en) Advanced path prediction
JP7140922B2 (en) Multi-sensor data fusion method and apparatus
CN112700470B (en) A method of target detection and trajectory extraction based on traffic video streams
US10657391B2 (en) Systems and methods for image-based free space detection
EP3789920A1 (en) Performance testing for robotic systems
Zhao et al. On-road vehicle trajectory collection and scene-based lane change analysis: Part i
Dueholm et al. Trajectories and maneuvers of surrounding vehicles with panoramic camera arrays
CN111771207A (en) Enhanced vehicle tracking
US20180089515A1 (en) Identification and classification of traffic conflicts using live video images
Revilloud et al. An improved approach for robust road marking detection and tracking applied to multi-lane estimation
CN117795566A (en) Perception of three-dimensional objects in sensor data
CN114842660B (en) Unmanned lane track prediction method and device and electronic equipment
Li et al. Exploring the causality of end-to-end autonomous driving
CN113158779B (en) Walking method, walking device and computer storage medium
Ballardini et al. Ego-lane estimation by modeling lanes and sensor failures
EP4675590A1 (en) Method for estimating a collision hazard level in a traffic environment
US20240119830A1 (en) System and method for predicting a future position of a road user
Mao et al. Bird’s Eye View Map for End-to-end Autonomous Driving Using Reinforcement Learning
DE102023114042A1 (en) Image-based pedestrian speed estimation
Alvarez et al. Autonomous vehicles: Vulnerable road user response to visual information using an analysis framework for shared spaces
US20260105757A1 (en) Visual language model instruction tuning for enhanced spatial reasoning
Ahmad et al. Review of Self-Driving Vehicles
Venkateswarlu et al. Navigating the future: Deep reinforcement learning for smart autonomous vehicles
EP4645249A1 (en) Three-dimensional annotation of traffic management objects

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20240702

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR