EP4631027A1 - Image synthesis apparatus and method - Google Patents

Image synthesis apparatus and method

Info

Publication number
EP4631027A1
EP4631027A1 EP23820825.0A EP23820825A EP4631027A1 EP 4631027 A1 EP4631027 A1 EP 4631027A1 EP 23820825 A EP23820825 A EP 23820825A EP 4631027 A1 EP4631027 A1 EP 4631027A1
Authority
EP
European Patent Office
Prior art keywords
image
scene
objects
segmentation map
network
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP23820825.0A
Other languages
German (de)
French (fr)
Inventor
Mark Leslie BLAXALL
Hans Wolff
Lev Markhasin
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sony Europe BV
Sony Semiconductor Solutions Corp
Original Assignee
Sony Europe BV
Sony Semiconductor Solutions Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sony Europe BV, Sony Semiconductor Solutions Corp filed Critical Sony Europe BV
Publication of EP4631027A1 publication Critical patent/EP4631027A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00Scenes; Scene-specific elements
    • G06V20/50Context or environment of the image
    • G06V20/56Context or environment of the image exterior to a vehicle by using sensors mounted on the vehicle
    • G06V20/58Recognition of moving objects or obstacles, e.g. vehicles or pedestrians; Recognition of traffic objects, e.g. traffic signs, traffic lights or roads
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00Arrangements for image or video recognition or understanding
    • G06V10/70Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/77Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation

Definitions

  • the present disclosure relates to augmented reality, particularly for a windshield of a vehicle or smartglasses.
  • Augmented reality systems offer a real-time interactive experience that displays computergenerated content with a matching alignment to a real-world view.
  • Virtual content can be constructive, wherein content is added to the view, or destructive, wherein content is reduced or removed from the view.
  • the virtual content may often be seamlessly interwoven with the real -world view, such that its experience gives a more natural and intuitive feel for the viewer. This may play an important role in safety for certain applications, such as driving.
  • a human driver is prone to making mistakes.
  • Reasons could be a distraction, a blind spot or reacting too late to a danger.
  • some accidents happen because drivers are distracted by advertisements or by accidents on another side of a highway.
  • Augmented reality systems particularly heads-up displays, are well established to provide the driver with useful information related to a planned route or a surrounding environment. While offering useful information, it is also the case that alarms, messages, and additional monitor bounding boxes may cause additional distraction for the driver and might actually create more danger, as opposed to reducing it.
  • augmented reality systems may help visually-impared people to read smaller text or read text from a farther distance. It may also help them follow the movement of small objects. The use of such a system could be so intuitive that the user forgets that there is an interface, encouraging further use for maintaining safety or overcoming challenges related to vision.
  • the present disclosure relates to an image synthesis apparatus.
  • the image synthesis apparatus comprises a first interface configured to receive an input image of a scene comprising objects of different object types.
  • the image synthesis apparatus further comprises a semantic segmentation network configured to map the input image to a segmentation map of the scene, wherein each pixel of the segmentation map is assigned to one of the different object types.
  • the image synthesis apparatus further comprises an object detection network configured to map the input image of the scene to one or more object locations of detected objects, wherein each object location is assigned one object type.
  • the image synthesis apparatus further comprises an image completion network configured to merge the one or more detected objects into the segmentation map of the scene based on the respective object locations and object types to obtain a modified segmentation map and to synthesize an output image based on the modified segmentation map.
  • the present disclosure relates to an image synthesis method.
  • the image synthesis method includes receiving an input image of a scene comprising objects of different object types.
  • the image synthesis method further includes mapping, using a semantic segmentation network, the input image to a segmentation map of the scene, wherein each pixel of the segmentation map is assigned to one of the different object types.
  • the image synthesis method further includes detecting, using an object detection network, one or more objects in the scene, wherein each detected object has associated therewith a respective object location and object type.
  • the image synthesis method further includes modifying the segmentation map of the scene based on the one or more detected objects and synthesizing an output image based on the modified segmentation map.
  • the present disclosure also relates to a computer program having computer-readable instructions for carrying out the above method, when the computer program is executed on a programmable hardware device.
  • Fig. 1 shows a scene as viewed through a windshield of a vehicle including buildings, a traffic light, multiple pedestrians, multiple other vehicles, an advertisement, a pothole, and a skyline;
  • Fig. 2 shows a block diagram of an apparatus for image synthesis according to a first embodiment
  • Fig. 3 shows a semantic segmentation map of the scene in Fig. 1 generated by a semantic segmentation network, wherein the objects listed above among others are assigned an object type to generate the semantic segmentation map;
  • Fig. 4 shows a labeled input image of the scene in Fig. 1 with labels of objects at corresponding object locations that have been detected by an object detection network;
  • Fig. 5 shows a modified version of the segmentation map of Fig. 3 that has been modified by an image completion network based on information provided by the object detection network;
  • Fig. 6 shows the scene of Fig. 1 in an augmented reality form as viewed on the windshield of the vehicle, wherein certain objects have been enhanced or removed depending on their features;
  • Fig. 7A shows a highway scene as viewed through a windshield of a vehicle;
  • Fig. 7B shows the highway scene in an augmented reality form as viewed on the windshield of the vehicle, wherein an input image based on the highway scene has been modified to a synthesized output image with an enlarged road sign and with a view of a blindspot;
  • Fig. 8 A shows a sport event scene as viewed by means of a television signal
  • Fig. 8B shows the sport event scene of the television signal in an augmented reality form with an enlarged ball to demonstrate how the image synthesis apparatus can assist people in other daily activities, particularly through the use of smartglasses;
  • Fig. 9 shows an image synthesis method according to a first embodiment
  • Fig. 10 shows a functional block diagram for an image synthesis apparatus with a first input from a first interface and a second input from a second interface for a vehicle setting;
  • an AR system can seamlessly integrate detected objects from an input image of a scene captured e.g. in front of the vehicle into a synthesized output image viewed by the driver. If the objects in the output image can be enhanced or removed by an image synthesis apparatus and then presented as an intuitive AR experience, driving safety may be dramatically improved.
  • Fig- 1 shows a scene 100a from the view of a driver through a windshield of a vehicle while driving on a road with traffic.
  • the scene 100a includes buildings 1-1, 1-2, and 1-3, a traffic light 2-1, multiple pedestrians 3-1, 3-2, and 3-3, multiple other vehicles 4-1 and 4-2, an advertisement 5-1, a pothole 6-1, a skyline 8, a sky 9, grass 10, and a road 11.
  • Some objects of the scene 100a are unlikely to cause a distraction or present a danger.
  • the sky 9, the grass 10, and the road 11 are all immobile objects that usually do not require special attention from the driver. Other objects may be more important for the driver to keep in mind.
  • the traffic light 2-1 and multiple pedestrians 3-1, 3-2, and 3-3 crossing the street at an intersection may be such important objects.
  • Other objects may indeed cause a distraction or present a danger to the driver.
  • there is a pothole 6-1 in the road presenting a danger. This danger may be challenging to avoid, especially since a pothole is often too small to be seen from far away and its view may be obstructed by other vehicles.
  • the advertisement 5-1 may cause a distraction for the driver. If a distraction is present, other dangerous objects, such as the pothole 6-1, may become even more dangerous since the driver may not react in time with less attention on the road.
  • the scene 100a also has a skyline 8 in the background. This may also be considered distracting since it may divert attention from the driver.
  • An AR experience while driving can help avoid dangerous situations, which may arise quickly and unexpectedly in any typical vehicle setting, such as the scene 100a depicted in Fig. 1.
  • Fig- 2 shows a block diagram of an image synthesis apparatus 200, which can provide such an AR experience in accordance with the present disclosure.
  • the image synthesis apparatus 200 comprises an interface 210 configured to receive an input image 212 of a scene 100 comprising objects of different object types.
  • the scene 100 may be the scene 100a or another scene.
  • the image synthesis apparatus 200 may optionally comprise a second interface 214, which will be discussed in Fig. 7A, 7B, and 10.
  • the scene 100 may be captured by one or more environmental sensors, such as a camera, radar, LiDAR, or combinations thereof, leading to the generation of the input image 212 that may be sent to the interface 210.
  • the input image 212 may comprise at least one of a camera image, a radar image, and a LiDAR image. Under certain conditions, a radar or LiDAR sensor may provide visual information of the scene 100 that a camera may fail to obtain, such as in foggy or dark conditions.
  • the input image 212 may depict a scene similar to the scene 100a in Fig. 1 and may thus comprise objects relevant to the driver that may be organized by pre-defined object types relevant to driving.
  • an object type may be any particular group of objects that share similar characteristics.
  • pre-defined object types may be useful to organize visual information within an image.
  • pre-defined object types may include “building”, “traffic light”, “pedestrian”, “vehicle”, “advertisement”, “pothole”, “road sign”, “skyline”, “sky”, “grass”, and “road”, etc.
  • separate objects of an object type may appear in an image, each with a respective location.
  • the image synthesis apparatus 200 comprises a semantic segmentation network 220 that is configured to receive the input image 212 from the interface 210 and to map the input image 212 to a segmentation map 222 of the scene 100. Each pixel of the segmentation map 222 is assigned to one of the different object types.
  • the image synthesis apparatus 200 further comprises an object detection network 230 that is configured to receive the input image 212 from the interface 210 and to map the input image 212 of the scene 100 to one or more object locations of detected objects. Each object location is assigned one object type.
  • the object detection network 230 may be configured to generate a labeled input image 232 of the scene 100 with labels of objects at corresponding object locations.
  • the object detection network 230 may comprise an artificial neural network, for example.
  • the image synthesis apparatus 200 further comprises an image completion network 240 configured to merge the one or more detected objects of the labeled input image 232 into the segmentation map 222 of the scene 100 based on the respective object locations and object types to obtain a modified segmentation map 244 and to synthesize an output image 248 based on the modified segmentation map 244.
  • the image completion network 240 may comprise an artificial neural network, for example.
  • the image completion network 240 may comprise a map modifier unit 242 and an image output unit 246.
  • the map modifier unit 242 may be configured to receive the segmentation map 222 and the labeled input image 232, to generate the modified segmentation map 244 based on the segmentation map 222 and the labeled input image 232, and to send the modified segmentation map 244 to the image output unit 246.
  • the image output unit 246 may be configured to receive the modified segmentation map 244 and to synthesize the output image 248 based on the modified segmentation map 244.
  • the image synthesis apparatus 200 may comprise a display 250 configured to display the synthesized output image 248 to a user.
  • the image output unit 246 of the image completion network 240 may be configured to send the output image 248 to the display 250.
  • the display 250 may be configured to display the synthesized output image 248 on a transparent member, for example.
  • the transparent member may be the windshield of a vehicle, a pair of glasses, or another means to display the synthesized output image 248.
  • the display 250 may also comprise a structure that is not transparent, such as a projection screen or another flat object that enables a user to view the synthesized output image 248.
  • the scene 100a may be captured by one or more environmental sensors to generate a first input image 212a.
  • the input image 212a may be received by the interface 210 and sent to the semantic segmentation network 220 and the object detection network 230 to generate a corresponding segmentation map 222a and a corresponding labeled input image 232a. These may be received by the image completion network 240 to generate a corresponding modified segmentation map 244a and a corresponding synthesized output image 248a, which may be displayed on the display 250.
  • a second scene 100b may lead to the synthesis of a second output image 248b, with all corresponding maps and images generated. This will be discussed in Fig. 7A and 7B.
  • a third scene 100c may lead to the synthesis of a third output image 248c, with all corresponding maps and images generated. This will be discussed in Fig. 8A and Fig. 8B.
  • the image synthesis apparatus 200 may repeat the procedure of receiving input images 212 and synthesizing corresponding output images 248. In this way, a stream of synthesized output images 248 may be generated based on a stream of input images 212.
  • the stream of synthesized output images 248 may be used to generate an AR expereince for a user.
  • the AR experience provided by the image synthesis apparatus 200 cannot be easily ignored by the driver. It enables an intuitive interaction because it augments reality, enabling the user’s natural reflexes to be activated.
  • the features of the image synthesis apparatus 200 that enables the transformation of the input image 212 to the synthesized output image 248 to be displayed on the display 250 are described in further detail below.
  • the transformation of the input image 212 to the synthesized output image 248 includes semantic segmentation of the input image 212.
  • the process of semantic segmentation is performed by the semantic segmentation network 220.
  • the process of semantic segmentation includes a categorization of each pixel of an image to an object type according to multiple relevant object types.
  • the object types may be pre-defined according to a setting of the image, such as a vehicle setting as previously described.
  • Each pixel of the image may be assigned to an object type based on the context of the surrounding pixels and the entire image. With each pixel of the image categorized according to the pre-defined object types, the image can be manipulated more easily towards a useful purpose.
  • Semantic segmentation performed within the means of conventional engineering requires an enormous domain knowledge database and enormous computational time.
  • an artificial neural network as the semantic segmentation network 220 to perform semantic segmentation can overcome such limitations by modeling the domain knowledge from a dataset of labeled pixels.
  • a neural network may be a convolutional neural network (CNN), which is applied to analyze visual imagery and uses relatively little pre-processing compared to other image classification algorithms.
  • CNN includes multiple layers, including convolutional layers, pooling layers, and fully-connected (FC) layers. Progressing through each layer, features of the input image 212 are identified, such as colors and edges, and eventually larger elements or shapes of the object, until it finally identifies the intended object.
  • Each convolutional layer comprises a feature detector, also known as a filter (or kernel), which can perform a process known as a convolution by moving across the receptive fields of the image 212, checking if s specific feature is present.
  • a convolutional neural network learns to optimize the filter through automated learning, whereas in traditional algorithms the filters are hand-engineered. This independence from prior knowledge and human intervention in feature extraction is a major advantage for a convolutional neural network. Examples of convolutional neural network architectures that may be used are U- Net, AlexNet, VGGNet, GoogLeNet, ResNet, and ZFNet.
  • the semantic segmentation network 220 may be a neural network, particularly a convolutional neural network.
  • the semantic segmentation network 220 can be trained with known input images to categorize objects to pre-defined object types. Semantic segmentation networks are often trained to differentiate objects in a foreground from a background. Images with a specific context, such as a vehicle setting, often have repeating object types in the foreground and background and the semantic segmentation network 220 can be trained to recognize and categorize such objects in the input image 212. For example, traffic lights are usually associated with a foreground and a sky is usually associated with a background, and the semantic segmentation network 220 can be trained to recognize traffic lights and a sky and to recognize the boundary between the two. With the recognition of the boundary, semantic segmentation can also lead to the determination of the shape of an object in the input image 212.
  • Fig- 3 shows an example segmentation map 222a of the input image 212a based on the scene 100a.
  • Each pixel of the scene 100 may be categorized by the semantic segmentation network 220 to an object type that was chosen for a specified setting, such as the vehicle setting in the scene 100a.
  • Example object types for the scene 100a include buildings as object type 1, traffic lights as object type 2, pedestrians as object type 3, vehicles as object type 4, advertisements as object type 5, potholes as object type 6, road signs as object type 7, skyline as object type 8, sky as object type 9, grass as object type 10, and roads as object type 11.
  • Each number depicted represents a pixel that has been categorized to an object type of that number.
  • the segmentation map 222a shows how certain areas have been labeled according to the object type that is depicted in the corresponding area according to the input image 212a of the scene 100a.
  • the semantic segmentation network 220 did not label any pixel as object type 7 because there was no road sign in the scene 100a.
  • Not every area of the segmentation map 222a in Fig. 3 is labeled with an object type number for clarity purposes.
  • each pixel of the input image 212 can be categorized to produce a segmentation map 222, wherein Fig. 3 does not illustrate every area with a labeled number. Rather, Fig.
  • the transformation of input image 212 to synthesized output image 248 also includes a process of detection of objects in the input image 212.
  • the process of detection of objects is performed by the object detection network 230.
  • Fig- 4 shows an example of a labeled input image 232a based on the input image 212a of the scene 100a with labels of objects at their corresponding object locations that have been detected by the object detection network 230.
  • the object detection network 230 may be configured to detect objects that are relevant in a vehicle setting, wherein each has associated therewith an object location.
  • pedestrians 3-1, 3-2, and 3-3 may be separate detected objects, each of object type 3 and each with a respective object location.
  • the object location may be expressed with 2-dimensional coordinates.
  • Vehicles 4-1 and 4-2 may be detected as separate objects of object type 4, while buildings 1- 1, 1-2, and 1-3 may be detected as separate objects of object type 1, each with a respective object location.
  • Objects that are the only instance within their object type in the input image 212a may also be detected objects with a respective object location.
  • the skyline 8, the sky 9, the grass 10, and the road 11 do not have associated therewith a detected object.
  • the object detection network 230 may also be configured as a neural network, including a convolutional neural network, as described above.
  • the semantic segmentation network 220 and the object detection network 230 may be trained in tandem, such that any object in any input image 212 will be categorized into the same object type by both networks.
  • the semantic segmentation network 220 may provide information related to how the input image 212 has been categorized according to the object types while generating the segmentation map 222.
  • the image synthesis apparatus 200 may comprise one or more neural networks configured to perform semantic segmentation and/or object detection.
  • Images generated directly from a segmentation map may be color-coded, with a unique color corresponding to each object type.
  • the detected objects of the object type might not be distinguished between each other.
  • this information would be lost in the output image 248 based on the segmentation map 222.
  • the vehicles 4-1 and 4-2 of object type 4 would not only lose many details, including brake lights and windows, but they would be represented by the same color and would have no visible feature to distinguish them other than their separate locations.
  • the buildings 1-1 to 1-3 and the pedestrians 3-1 to 3-3 are the same.
  • the image completion network 240 may be configured to receive information related to the input image 212 from both the semantic segmentation network 220 and the object detection network 230 and to synthesize an output image 248, wherein one or more detected objects are merged into the segmentation map 222 according to their respective object locations.
  • the image completion network 240 may be configured to merge the one or more detected objects into the segmentation map 222 of the scene 100 on the condition that the object type of the respective detected object meets a predetermined criterion of importance.
  • the pre-determined criterion of importance of an object may include the object presenting a danger to the driver and an object presenting relevant information for driving safety. For example, in the scene 100a, the pothole 6-1 may present a danger.
  • a danger may also be a possible future danger to the driver.
  • This may include pedestrians 3-1 to 3-3, for example.
  • the pedestrians 3-1 to 3-3 may not be dangerous crossing the road while the traffic light 2-1 is red but may become dangerous if still crossing when the traffic light 2-2 turns green.
  • Relevant information for driving safety may be anything that guides the driver to maintain a safe interaction with other vehicles and the entire surrounding. This may include the traffic light 2-1 and road signs. Relevant information may also include physical obstacles to be avoided, whether on or off the road. This may include buildings 1-1 to 1-3 off the road and vehicles 4-1 and 4-2 on the road. If only detected objects meeting the pre-determined criterion of importance are included in the output image 248, the driver only need focus on these obj ects and distracting objects that have no value when seen may be removed out of view.
  • the image completion network 240 may also be configured as an artificial neural network, including a convolutional neural network, as described above.
  • the image completion network 240 may be trained in tandem with the semantic segmentation network 220 and/or the object detection network 230. While the training for the semantic segmentation network 220 and the object detection network 230 may focus more on categorization and feature extraction, the image completion network 240 may focus more on how to enhance or remove objects based on a category or feature of an object. For example, the image completion network 240 may be trained by manually constructed output images that have had objects manually enhanced or removed based on a category or feature of the object, which may have been determined based on segmentation maps 222 generated by the semantic segmentation network 220 and labeled input images 232 generated by the object detection network 230.
  • the manually constructed output images may provide multiple different contexts or situations to train the image completion network 240 to enhance or remove objects depending on the respective context or situation. For example, in a vehicle setting, an object may be enhanced differently or may or may not be removed depending on factors related to the vehicle, such as whether the vehicle is in motion, or depending on factors related to the surrounding, such as weather conditions or darkness.
  • the image synthesis apparatus 200 may comprise one or more neural networks configured to perform semantic segmentation, object detection and/or image completion.
  • the image completion network 240 may include the map modifier unit 242 and the image output unit 246.
  • the map modifier unit 242 may be a neural network of the image completion network 240 that is configured to modify the segmentation map 222 as described above.
  • the map modifier unit 242 may be trained to complete an image based on input from the semantic segmentation network 220 and the object detection network 230 and manually constructed output images, as described above.
  • the image output unit 246 may be a neural network configured to synthesize an output image 248 based on the modified segmentation map 244.
  • the image output unit 246 may be configured to synthesize contents so that the output image 248 looks visually realistic. This may include traditional patch-based methods or deep learning-based methods. It may be a convolutional neural network that includes a spatially adaptive normalization (SPADE) layer, known to synthesize photorealistic images from a segmentation map by minimizing the loss of information by normalization layers.
  • SPADE spatially adaptive normalization
  • a method of compositing the output image 248 by linearly blending it with the input image 212 may include a method of compositing the output image 248 by linearly blending it with the input image 212.
  • Other methods of image synthesis from the segmentation map 222 may include a generative adversarial network (GAN) or a generative transformer model.
  • GAN generative adversarial network
  • a bidirectional transformer model for image synthesis offers the possibility to generate new tokens from previously generated tokens in all directions, as opposed to a sequence form, leading to significantly faster processing.
  • Fig- 5 shows an example of a modified segmentation map 244a based on the segmentation map 222a of Fig. 3, each based on scene 100a.
  • Multiple detected objects from the labeled input image 232 may be merged into the modified segmentation map 244 by the image completion network 240.
  • the image completion network 240 may be configured to merge the detected objects into the modified segmentation map 244 based on the object locations in the labeled input image 232 from the object detection network 230 and based on the segmentation map 222 from the semantic segmentation network 220.
  • the image completion network 240 may also be configured to merge the detected objects with all corresponding details related to their physical structure as captured in the input image 212 into the modified segmentation map 244.
  • the merged detected objects from the labeled input image 232a are the traffic light 2-1, pedestrians 3-1, 3-2, and 3-3, vehicles 4-1 and 4-2, and the pothole 6-1.
  • Some of the merged detected objects may have one or more characteristics changed, including the size, location, and contrast.
  • the merged detected objects may be labeled accordingly.
  • Vehicles 4-1 and 4-2 and pedestrian 3-3 may be labeled without modification, since they have maintained the same size, location, and contrast.
  • Pedestrians 3-1 and 3-2 may be labeled with an “s” as 3-1 s and 3- 2s for having a changed size.
  • Traffic light 2-1 may be labeled with an “s” and an “1” as 2-1 si for having a changed size and a changed location.
  • Pothole 6-1 may be labeled with a “c” for having a changed contrast as 6-lc.
  • 6-lcslb associated with pothole 6-lc is a separate enlarged image box, 6-lcslb, labeled with a “c”, “s”, “1”, and “b” for having a changed contrast, size, and location, and having a separate box.
  • the detected objects that are merged into the modified segmentation map 244a in Fig. 5 are labeled with parentheses, such as (4-1) for the merged detected object 4-1, to differentiate these labels from numbers labeled as examples of pixels corresponding to an object type, such as “4”.
  • Other detected objects from the scene 100a may not be included in the modified segmentation map 244a, such as buildings 1-1, 1-2, and 1-3 and advertisement 5-1. Pixels corresponding to the skyline 8 have had their object type changed, such that no pixels in the modified segmentation map 244a correspond to the skyline 8 and no skyline will be found within the output image 248a. Other pixels have been changed from one object type to another based on changes in size of a merged detected object.
  • the image completion network 240 may be configured to merge detected objects detected by the object detection network 230 into the segmentation map 222 generated by the semantic segmentation network 220 with modifications to the size, location, and/or contrast of the detected objects, depending on characteristics of the detected object and the context of the scene 100.
  • Each example of how a detected object is merged into or removed from the segmentation map 222 to obtain a modified segmentation map 244 serves to illustrate how the image completion network 240 of the image synthesis apparatus 200 may be configured.
  • Vehicles 4-1 and 4-2 may be objects of interest to the driver because the driver must keep track of them to avoid a collision. Their corresponding pixels have been categorized into object type 4 in the segmentation map 222a. While the vehicles could be clearly seen and distinguished as a single-color output by the segmentation map 222a, seeing the actual form of the detected objects as captured in the input image 212a can better maintain an intuitive interaction with the road while driving. More importantly, the two vehicles 4-1 and 4-2 may become difficult to distinguish from each other in subsequent input images 212a of the scene 100a, which would lead to a dangerous situation. For example, if one vehicle merged into the same lane as the other vehicle, it may be partially blocked by the other and it may become difficult to distinguish the two. As such, any vehicle driving on the road would meet the predetermined criterion of importance for the image completion network 240 to merge the detected object into the segmentation map 222 so that the output image 248 includes the structural details of each vehicle as captured by the input image 212.
  • the image completion network 240 may be configured to enhance the one or more detected objects and to merge the one or more enhanced objects into the segmentation map 222 of the scene 100 based on the respective object locations and object types.
  • the image completion network 240 may be configured to change a size of the one or more detected objects and to merge the one or more objects with the changed size into the segmentation map 222 of the scene 100 based on the respective object locations and object types.
  • the image completion network 240 may be configured to enlarge the one or more detected objects. This may present the detected object in the synthesized output image 248 in a form that is easier for the driver to perceive.
  • the image completion network 240 may be configured to merge one or more detected objects into the segmentation map 222 in an enlarged form, wherein the number of pixels corresponding to the respective object to be enlarged is increased from its original number of pixels in the segmentation map 222 to a larger number of pixels in the modified segmentation map 244.
  • the modified segmentation map 244 shows an example of how the image completion network 240 may be configured to selectively enhance one or more objects of an object type depending on the behavior and location of the respective object.
  • the pedestrians 3- Is and 3 -2s have been enlarged from their original size because they are located in the middle of the road, while pedestrian 3-3 has maintained his original size because he is on the side of the road and standing.
  • This principle may apply to other objects as well. For example, a vehicle that is determined to have associated therewith dangerous driving behavior may be merged and presented differently than other vehicles.
  • the image completion network 240 may selectively enhance the pedestrians on the side of the road that are children, particularly playing children, to enable the driver to perceive the children more easily until passing them. The same may also be applied to animals.
  • the object detection network 230 may be a neural network that is trained to detect information related to the behavior and motion patterns of a mobile object, given multiple input images 212.
  • the image completion network 240 may be configured to merge such an object in an enhanced form based on the location, behavior, or motion patterns of the object. This may help the driver see such a mobile object more quickly, thereby reducing the risk of injury/damage for both the driver and the object.
  • a detected object only presents a distraction and provides no relevant information to the driver, it can be removed from the segmentation map 222 by the image completion network 240.
  • the image completion network 240 can reassign its corresponding pixels to another object type, such as those of a neighboring background.
  • the pixels assigned to object type 5 by the semantic segmentation network 220 have been re-assigned to the object type of its neighboring background, the sky 9 and the grass 10, such that it is not seen in the corresponding synthesized output image 248a.
  • the image completion network 240 may be configured to change the object type assigned to the one or more detected objects and to merge the one or more objects with the changed object type into the segmentation map 222 based on the respective object locations. More specifically, the pixels associated with an object to be removed may be re-assigned to one or more other object types, such as those of a neighboring background. Merging the detected object with the changed object type into the segmentation map 222 based on the respective object location may then result in the detected object to remain unseen in the synthesized output image 248. In the case of the modified segmentation map 244a, the driver will be unaware that the advertisement 5-1 is present in the scene 100a and any distraction associated with it will be prevented.
  • Pixels that have been categorized by the semantic segmentation network 220 to an object type and do not have associated therewith any detected object may also be changed or removed by the image completion network 240.
  • the skyline 8 in scene 100a may also be considered a distraction for a driver, particularly for a driver that is unfamiliar with the geographical area.
  • all pixels corresponding to object type 8 for the skyline have been removed and replaced by pixels corresponding to object type 9 for the sky, or other object types associated with an enlargement and/or relocation of another detected object.
  • the image completion network 240 may be configured to change the object type of a pixel to any other object type.
  • the changed object type may also be associated with a merged detected object.
  • the traffic light 2-1 in the segmentation map 222a was also enlarged by the image completion network 240 to be more easily perceived by the driver, as shown in modified segmentation map 244a, wherein the number of pixels corresponding to the traffic light 2-1 is increased to a larger number of pixels in the modified segmentation map 244a. If a certain portion of the traffic light 2-1 is determined to be irrelevant for the driver by the image completion network 240, such as the metal post supporting it, the image completion network 240 may remove this portion, as shown in the modified segmentation map 244a.
  • the traffic light 2-1 can be merged to be in a different location in the modified segmentation map 244a where it may be more easily perceived by the driver. This may be at a more central location, as depicted in the modified segmentation map 244a.
  • the image completion network 240 may be configured to exchange pixels corresponding to a merged detected object with other pixels that are not associated with any other merged detected objects and are determined to not convey information relevant to the driver.
  • Certain objects on the road may be important for the driver to perceive with great detail while driving, such as potholes. While some potholes are large and must be avoided, other potholes are smaller and need not be avoided. This may be important in certain situations, such as needing to avoid another vehicle by driving near or over the pothole 6-1 in scene 100a. It may be useful for the driver to see the pothole 6-1 with a greater contrast compared to other objects to decide whether it must be avoided or not.
  • the image completion network 240 may be configured to synthesize the output image 248 with an increased contrast of the one or more detected objects compared to remaining regions of the output image.
  • the pothole 6-1 in scene 100a is depicted as pothole 6-lc in the modified segmentation map 244a for being shown with greater contrast compared to the other detected objects (as can be seen in Fig. 6).
  • the object detection network 230 may also be configured to obtain distance information of one or more detected objects to be used by the image completion network 240. Such distance information may be obtained by comparing the size of the detected object to other objects in the input image 212 with a well-known size range, such as the width and height of other vehicles or the width of the road lanes. Additionally or alternatively, distance information may be obtained by using distance or ranging sensors, such as radar, ultrasonic, and/or lidar sensors. The image completion network 240 may be configured to change the size of the one or more detected objects based on distance information associated with the respective object.
  • the image completion network 240 may be configured to merge an enlarged image of the detected object to appear as a separate image box in an appropriate location.
  • the pothole 6-1 in scene 100a is depicted in the modified segmentation map 244a as 6-lc and also as an enlarged image box 6-lcslb at a separate location (also seen in Fig. 6).
  • the separate location of such an enlarged box may preferably be a corner of the synthesized output image 248.
  • the separately located enlarged image box 6-lcslb combined with the contrasted depiction of the pothole 6- 1c at the original location may be useful for the driver to see exactly where the pothole is located and to simultaneously perceive it better.
  • the image completion network 240 may also be configured to not include certain detected objects detected by the object detection network 230 based on whether or not the semantic segmentation network 220 has provided enough information related to the detected object. For example, buildings 1-1, 1-2, and 1-3 in the scene 100a are detected objects in the labeled input image 232a and they present physical obstacles to the driver to be avoided. But since the semantic segmentation network 220 has already classified all pixels of objects 1-1, 1-2 and 1-3 to the object type 1, it is already possible to distinguish where the buildings are located in relation to the road. Also, the buildings are very stable structures and will not change their absolute location. The driver can thus easily perceive these locations to avoid a collision without having them merged into the corresponding segmentation map 222a.
  • the modified segmentation map 244a in Fig. 5 illustrates this, wherein the pixels corresponding to the object locations of detected objects 1-1, 1-2, and 1-3 are labeled with the object type 1, but not with detected objects.
  • the object detection network 230 may be a neural network configured to learn to detect smaller objects corresponding to specified details within larger detected objects, such as buildings 1-1, 1-2, and 1-3 in scene 100a.
  • the image completion network 240 may be configured to include such smaller objects within larger objects detected by the object detection network 230 into the modified segmentation map 244 based on a pre-determined criterion of importance. If a detected image is determined to be distracting but provides relevant information, the image completion network 240 may be configured to merge generic images of a simpler form including the relevant information into the modified segmentation map 244. This would convey the relevant information in a less distracting way to the driver. Configurations related to the inclusion or exclusion of detected objects may be customized according to the preferences of a user or to evolving safety standards.
  • Fig- 6 shows an example of the synthesized output image 248a of the scene 100a that may be viewed by the driver on the windshield of the vehicle or on glasses worn by the driver, for example.
  • the synthesized output image 248a is based on the modified segmentation map 244a depicted by Fig. 5 that was modified from the segmentation map 222a depicted by Fig. 3.
  • the detected objects 2-1, 3-1, 3-2, 3-3, 4-1, 4-2, and 6-1 were all merged into the modified segmentation map 244a by the image completion network 240 and are thus depicted in the output image 248a.
  • the image completion network 240 also merged a separate enlarged image box 6-lcslb located in the bottom right corner of the modified segmentation map 244a. Detected objects 1-1, 1-2, and 1-3 (buildings) were not merged into the modified segmentation map 244a.
  • the synthesized output image 248a comprises pixels categorized to object type 1 for buildings, object type 9 for the sky, object type 10 for grass, and object type 11 for the road, wherein pixels corresponding to object type 8 for the skyline have been removed.
  • the image synthesis apparatus 200 may optionally comprise a second interface 214 which is configured to receive additional information 216 beyond the input image 212 of the (first) interface 210.
  • the additional information 216 may yield one or more additional object locations of additional objects which are not visible in the input image 212. Each additional object location is assigned one object type.
  • the image completion network 240 may be configured to merge the one or more additional objects into the segmentation map 222 of the scene 100 based on the respective additional object locations and object types in order to obtain the modified segmentation map 244.
  • the first interface 210 may be coupled to a first environmental sensor, such as a camera, to obtain the input image 212.
  • the second interface 214 may be coupled to a different second environmental sensor and/or a communication network to obtain the additional information 216 yielding one or more additional object locations of additional objects.
  • the second interface 214 may be coupled to a second camera having a different field of view than the camera coupled to the first interface 210.
  • the second interface 214 may be coupled to one or more other environmental sensors, such as radar and/or lidar sensors covering the same or a different field of view (e.g., a blind spot) as the camera coupled to the first interface 210.
  • the additional information 216 received by the second interface 214 may comprise at least one of a camera image, a radar image, and a LiDAR image. Additionally or alternatively, the second interface 214 may be coupled to a wireless communication network, such as a WiFi- or mobile communications network, to receive the additional information 216 yielding one or more additional object locations of additional objects. In such an example, the additional information 216 may include real-time traffic information, weather information, information on points of interest, or the like. That is, the second interface 214 may be configured to receive the additional information 216 from at least one of a camera signal, a radar signal, a lidar signal, and a radio signal.
  • a wireless communication network such as a WiFi- or mobile communications network
  • the image completion network 240 may be configured to merge the one or more additional objects into the segmentation map 222 of the scene 100 on a condition that the object type(s) of the respective additional object(s) meets a pre-determined criterion of importance.
  • the one or more additional objects can correspond to objects (e.g., persons, other vehicles, etc.) in a blind spot of the driver.
  • the pre-determined criterion of importance for the second interface 214 may be defined the same or differently from the pre-determined criterion of importance for the first interface 210.
  • the pre-determined criterion of importance of an object may include the object presenting a danger to the driver and an object presenting relevant information for driving safety.
  • the pre-determined criterion of importance of the second interface 214 for merging the pedestrian into the modified segmentation map 244 can be based on a different requirement of distance from the vehicle, which is unlikely to be driven backwards. Other possible requirements may relate to the behavior and motion patterns of the pedestrian.
  • the pre-determined criterion of importance for the second interface 214 may depend on its orientation in comparison to the first interface 210 and the type of input images 212 or the additional information 216 it may receive, among other factors.
  • Fig. 7A shows the second scene 100b, which is also from the perspective of a driver in a vehicle.
  • An input image based on the scene 100b may be received by the first interface 210 and may then be sent to the semantic segmentation network 220 and the object detection network 230.
  • the image synthesis apparatus 200 may generate for the scene 100b a segmentation map analogous to the segmentation map 222 and a labeled input image analogous to the labeled input image 232. These may be received by the image completion network 240, which may generate a modified segmentation map analogous to 244 and a synthesized output image 248b, shown in Fig. 7B.
  • the synthesized output image 248b may then be displayed by the display 250.
  • Fig. 7B shows the synthesized output image 248b based on the scene 100b.
  • a road sign is depicted in an enlarged form as 7-ls, which was not visible in the scene 100b.
  • the enlarged road sign 7-ls labeled as corresponding to object type 7 for road signs, shows a red frame around a white circle reading the number 30 in digits of black color.
  • the number 30 may be a signal to the driver that he must slow down to a speed of 30 mph or 30 km/h in a specific range of a highway. While no such road sign is visible in the scene 100b, it may be that there was such a road sign specifying a speed limit that the vehicle has already driven past and the image synthesis apparatus 200 has maintained the road sign in view after passing.
  • the image completion network 240 may be configured to maintain merged objects in future output images 248 in situations useful to the driver.
  • the image synthesis apparatus 200 may comprise a memory to temporarily store objects detected from the object detection network 230 and may be configured to merge such stored objects, as necessary.
  • the enlarged road sign 7- Is may also be an example of the additional information 216 that may be received by the second interface 214 by a communication network. While such road signs have a limited physical range of communication to the driver, the image synthesis apparatus 200 may convey information to the driver in a much broader range.
  • the enlarged road sign 7- Is may be a generic version of a road sign chosen by the image synthesis apparatus 200 after receiving additional information 216 related to the speed limit.
  • the additional information 216 may also be received by the second interface 214 by means of a blindspot camera.
  • a sub-image box 4-3b depicts a separate blindspot scene comprising another vehicle.
  • the sub-image box 4-3b may be labeled with a “b” because it is shown in the form of a separate box.
  • the blindspot scene may be captured by a camera or environmental sensor and provided to the optional second interface 214.
  • the sub-image box 4-3b comprising the other vehicle may then be depicted in the output image 248b after having been merged into the corresponding segmentation map.
  • the merging of a vehicle in such a blindspot scene into the corresponding segmentation map may be dependent on the pre-determined criterion of importance that may be defined for the second interface 214, as previously described.
  • the image completion network 240 may be configured to merge an entire separate image received by the second interface 214, such as 4-3b, into a corresponding segmentation map in the form of a sub-image box.
  • the image completion network 240 may also be configured to merge one or more objects detected within the separate image received by the second interface 214 alone into an appropriate location of the corresponding segmentation map.
  • Objects detected from the second interface 214 and merged into the modified segmentation map 244, either as a sub-image box or as objects alone, may be labeled with text, including a brief description of context, as necessary.
  • the second interface 214 may be configured to obtain the additional information 216 by means of the communication network related to an object that is not yet among the surroundings of the vehicle.
  • the vehicle may approach an object, such as a dangerous object blocking the road, and the object may be minutes away from being perceived by means of the first or second environmental sensor.
  • the second interface 214 may be configured to obtain the additional information 216 in the form of GPS coordinates.
  • the GPS coordinates may be the object’s GPS coordinates and the vehicle’s GPS coordinates.
  • the relevant information 216 may also include information related to traffic and an average speed of the vehicle.
  • the image synthesis apparatus 200 may be configured to calculate an approximate time in which the object may be perceived by the first or second environmental sensor.
  • the image completion network 240 may be configured to portray generic versions of the object in the synthesized output image 248 based on the relevant information 216 provided to the second interface 214. For example, if the object has not yet been photographed, a synthetic generic image of the dangerous object may be inserted into the segmentation map 222 and viewed in the lower left comer, either in a box as portrayed in Fig. 8B or as a labeled object.
  • the image completion network 240 may be configured to label the object with text, including a brief description of context, as necessary.
  • the image completion network 240 may be configured to remove detected objects from view to prevent distraction to the driver. It may also be configured to display a detected object smaller if the object is determined to be a distraction.
  • a benefit to displaying a detected object determined to be a distraction in a smaller form is that the driver does not have as much detail to focus on, but still gains an understanding of context for the vehicle surrounding. For example, it may explain why there is a certain amount of traffic or why other vehicles are driving at a certain speed, if applicable.
  • an accident may happen because a driver is focused on a distraction, such as an accident pileup on the side of the highway. While the accident pile-up may be removed from view, it may be reduced in size in the field of view for the driver to gain an understanding of context.
  • the image synthesis apparatus 200 may be configured to reduce the size of the portrayal of a physical phenomenon, such as a road reflection or a mirage, or remove it entirely. In reducing the size of the portrayal, the driver may still be aware of the conditions around the vehicle that led to the phenomenon, such as being aware that it is a hot and dry day, while being less distracted by it. Configurations of the image synthesis apparatus 200 that control how objects are enhanced, removed, or reduced in size may be customized according to driver preferences or evolving safety standards.
  • Embodiments of the present disclosure may not only be relevant to vehicle settings, but also to other applications, an example of which is illustrated in Fig. 8A.
  • Fig. 8A shows a scene 100c, which is a view of a television signal of a football match on a corresponding display.
  • the scene 100c comprises a ball 12-1, which is a small object that may be difficult to follow during the match for a viewer, particularly if the viewer is visually impaired.
  • the first interface 210 may be configured to receive a television signal or any other visual signal as the input image 212. An input image based on the scene 100c may be received by the first interface 210 and may then be sent to the semantic segmentation network 220 and the object detection network 230.
  • the image synthesis apparatus 200 may generate for the scene 100c a segmentation map analogous to the segmentation map 222 and a labeled input image analogous to the labeled input image 232. These may be received by the image completion network 240, which may generate a modified segmentation map analogous to 244 and a synthesized output image 248c, shown in Fig. 8B. The synthesized output image 248c may then be displayed by the display 250.
  • Fig- 8 shows the synthesized output image 248c of the scene 100c.
  • the ball 12-1 is presented as a size-enhanced ball 12- 1 s.
  • This demonstrates an example of how augmented reality can also assist a user of the image synthesis apparatus 200 in daily life beyond driving, particularly with a display on a pair of transparent glasses.
  • Such an embodiment can be used in a wide variety of other settings to enhance objects. While displaying the output image on a windshield may be optimal for driving situations, displaying the output image on glasses allows a flexibility of use by the user in nearly any situation.
  • the image synthesis apparatus 200 may comprise a semantic segmentation neural network 220, object detection neural network 230, and/or an image completion neural network 240 that are trained for many different specific tasks.
  • the image synthesis apparatus 200 with a camera and a glasses display may be configured to capture a stream of input images 212 depicting a sport event scene such as 100c, detect the ball of the sport event within each input image 212, synthesize output images 248 with an enlarged version of the ball, and continuously display the output images 248 with the enlarged ball on the glasses, particularly as a video stream. This may enable the user to follow the ball and the entire sport event more easily. This may be an even greater help in other sports associated with smaller objects, such as ice hockey, baseball, or tennis.
  • the image synthesis apparatus 200 may be trained to enhance other objects of the sport event scene 100c, such as the players, referee, or boundary lines of the field of play.
  • the user may also benefit from the image synthesis apparatus 200 when viewing a football match next to its field of play.
  • the user may be located far away from the field, where it may be difficult to follow the ball.
  • the input image 212 may be in the form of a camera image capturing the sport event scene 100c as seen from the perspective within the stadium.
  • the same benefits that the user experiences with the television signal as the input image 212 can also be experienced with a camera image as the input image 212.
  • the semantic segmentation network 220, object detection network 230, and image completion network 240 may be trained to provide an AR experience corresponding to any setting, which may be a sport event, as well as a city for a tourist or a natural landscape for a hiker, among others.
  • the user may specify a trained setting from among the total number of settings that the networks 220; 230; 240 have been trained to use the image synthesis apparatus 200 to synthesize and display output images 248. Different groups of pre-defined object types and/or typical objects may be used to train the networks 220; 230; 240 under different trained settings.
  • the image synthesis apparatus 200 may comprise various modes, each mode corresponding to a respective trained setting, and the image synthesis apparatus 200 may switch to a different mode based on an input by a user.
  • a synthetic image may also be received as the input image 212 by the image synthesis apparatus 200, which may then synthesize the output image 248 with certain enhanced or removed objects within the received synthetic image in the same manner as when the input image 212 depicts a real-life scene.
  • the image synthesis apparatus may be configured to receive any collection of input images 212 that depicts visual information, whether captured directly by a camera in proximity to the user or captured elsewhere, or whether real or synthetic, and then to synthesize corresponding output images 248 that enhances or removes certain objects based on a training of the semantic segmentation network 220, the object detection network 230, and the image completion network 240, and to display the corresponding output images 248 on a display 250, particularly on a transparent member, such as glasses.
  • the examples and embodiments of the image synthesis apparatus 200 demonstrate a flexibility in configuration, comprising a semantic segmentation network 220, an object detection network 230, and an image completion network 240 that are able to work together to receive input images 212 and generate corresponding output images 248.
  • the image synthesis apparatus 200 may comprise one interface 210 configured to receive input images 212, a first and second interface 210; 214 configured to receive input images 212 of the same scene 100 or different scenes, or it may comprise any number of interfaces offering information of any number of scenes, wherein the first interface 210 is configured to receive the input image 212 and the further interfaces are configured to receive additional information 216, enabling the image synthesis apparatus 200 to synthesize an output image 248 with certain enhanced or removed objects and display it for the user based on a specific training, setting, or context. As such, the image synthesis apparatus 200 may be trained for any specific task to provide an AR experience that supports a user in safety or participation in activities of daily life.
  • Method 1000 includes a step SI of receiving an input image 212 of a scene 100 comprising objects of different object types.
  • the scene 100 may correspond to a vehicle setting, such as scene 100a or 100b, or it may correspond to a football match, such as scene 100c, among others.
  • the scene 100 may comprise various objects and object types related to the context of the scene 100.
  • the method also includes a step S2 of mapping, using a semantic segmentation network 220, the input image 212 to a segmentation map 222 of the scene 100, wherein each pixel of the segmentation map 222 is assigned to one of the different object types.
  • the method 1000 also includes a step S3 of detecting, using an object detection network 230, one or more objects in the scene 100, wherein each detected object has associated therewith a respective object location and object type.
  • Such an object detection network 230 may generate a labeled input image 232.
  • the method also includes a step S4 of modifying the segmentation map 222 of the scene 100 based on the one or more detected objects to generate a modified segmentation map 244.
  • the method also includes a step S5 to synthesize an output image 248 based on the modified segmentation map 244.
  • Fig. 10 outlines a procedure 1000 wherein the image synthesis apparatus 200 includes the second interface 214.
  • SI to S4 and S7 describe processing acts of the image synthesis apparatus 200 including the (first) interface 210 that have previously been outlined in Figures 1 to 6.
  • S5 and S6 describe processing acts of the image synthesis apparatus 200 additionally including the second interface 214.
  • the first interface 210 is connected to a camera, or alternatively a radar or LiDAR sensor, and may receive one or more input images 212 depicting the scene 100,
  • the scene 100 may comprise objects of different object types.
  • the one or more input images 212 may be in the form of video frames that may collectively be processed to generate one or more output images 248 or augmented frames.
  • the augmented frames may form a video stream to be continuously viewed as an AR experience.
  • the input images 212 may be provided as input to the semantic segmentation network 220 and the object detection network 230.
  • the semantic segmentation network 220 and object detection network 230 may detect certain objects and object types to be enhanced or removed in S4 by the image completion network 240.
  • the semantic segmentation network 220, the object detection network 230, and/or the image completion network 240 may be a neural network configured to detect the objects and object types and to remove them or merge them in enhanced form into the segmentation map 222, as previously described. If such objects or object types are found, the image completion network 240 merges or removes the objects or object types into the segmentation map 222 to obtain a modified segmentation map 244. For example, in Fig. 6, S4 results in the enlarged pedestrians 3-ls and 3-2s in the synthesized output image 248a. If such objects or object types are not found, the procedure skips S4 and proceeds to S5.
  • the image synthesis apparatus 200 may further comprise the second interface 214 configured to receive additional information 216 beyond the input image 212 of the scene 100, the additional information 216 yielding one or more additional object locations of additional objects, wherein each additional object location is assigned one object type.
  • the image completion network 240 may be configured to merge the one or more additional objects into the segmentation map 222 of the scene 100 based on the respective additional object locations and object types in order to obtain the modified segmentation map 244.
  • a second environmental sensor and/or a communication network may capture a blindspot scene or any scene comprising additional information 216.
  • the second interface 214 of the image synthesis apparatus 200 may receive the additional information 216 from the second environmental sensor and/or communication network.
  • the additional information 216 may be in the form of one or more input images, which may be sent to the object detection network 230 and/or the image completion network 240.
  • one or more objects included with the additional information 216 may be merged into the segmentation map 222 of the scene 100 to be included in the modified segmentation map 244 and the synthesized output image 248.
  • the image completion network 240 may be configured to merge the one or more additional objects into the segmentation map 222 on a condition that the object type of the respective additional object meets a pre-determined criterion of importance for the second interface 214 and to merge the one or more objects included with the additional information 216 in an enhanced form, as previously described.
  • the synthesized output image 248 comprising information obtained from both the first interface 210 and the second interface 214 is displayed on a display 250.
  • Example 1 is an image synthesis apparatus comprising a first interface configured to receive an input image of a scene comprising objects of different object types, a semantic segmentation network configured to map the input image to a segmentation map of the scene, wherein each pixel of the segmentation map is assigned to one of the different object types, an object detection network configured to map the input image of the scene to one or more object locations of detected objects, wherein each object location is assigned one object type, and an image completion network configured to merge the one or more detected objects into the segmentation map of the scene based on the respective object locations and object types to obtain a modified segmentation map and to synthesize an output image based on the modified segmentation map.
  • Example 2 the image completion network of Example 1 is configured to merge the one or more detected objects into the segmentation map of the scene on the condition that the object type of the respective detected object meets a pre-determined criterion of importance.
  • Example 3 the image completion network of any one of the previous Examples is configured to enhance the one or more detected objects and to merge the one or more enhanced objects into the segmentation map of the scene based on the respective object locations and object types.
  • Example 4 the image completion network of any one of the previous Examples is configured to change a size of the one or more detected objects and to merge the one or more objects with the changed size into the segmentation map of the scene based on the respective object locations and object types.
  • Example 5 the image completion network of Example 4 is configured to change the size of the one or more detected objects based on distance information associated with the respective object.
  • Example 6 the image completion network of Examples 4 or 5 is configured to enlarge the one or more detected objects.
  • Example 7 the image completion network of any one of the previous Examples is configured to synthesize the output image with an increased contrast of the one or more detected objects compared to remaining regions of the output image.
  • Example 8 the image completion network of any one of the previous Examples is configured to change the object type assigned to the one or more detected objects and to merge the one or more objects with the changed object type into the segmentation map of the scene based on the respective object locations.
  • Example 9 the image synthesis apparatus of any one of the previous Examples further comprises a second interface configured to receive additional information beyond the input image of the scene, wherein the additional information yields one or more additional object locations of additional objects, wherein each additional object location is assigned one object type, wherein the image completion network is configured to merge the one or more additional objects into the segmentation map of the scene based on the respective additional object locations and object types in order to obtain the modified segmentation map.
  • Example 10 the first interface of the image synthesis apparatus of Example 9 is coupled to a first environmental sensor and the second interface of Example 9 is coupled to a different second environmental sensor and/or a communication network.
  • Example 11 the image completion network of Examples 9 or 10 is configured to merge the one or more additional objects into the segmentation map of the scene on a condition that the object type of the respective additional object meets a pre-determined criterion of importance.
  • Example 12 the second interface of Examples 9 or 10 is configured to receive the additional information from at least one of a camera signal, a radar signal, a lidar signal, and a radio signal.
  • Example 13 the input image of any one of the previous Examples comprises at least one of a camera image, a radar image, and a LiDAR image.
  • the image synthesis apparatus of one of the previous Examples further comprises a display configured to display the synthesized output image to a user.
  • Example 15 the display of Example 14 is configured to display the synthesized output image on a transparent member.
  • Example 16 is an image synthesis method, the method comprising receiving an input image of a scene comprising objects of different object types, mapping, using a semantic segmentation network, the input image to a segmentation map of the scene, wherein each pixel of the segmentation map is assigned to one of the different object types, detecting, using an object detection network, one or more objects in the scene, wherein each detected object has associated therewith a respective object location and object type, modifying the segmentation map of the scene based on the one or more detected objects, and synthesizing an output image based on the modified segmentation map.
  • Example 17 is a computer program having computer readable instructions for carrying out the method according to Example 16, when the computer program is executed on a programmable hardware device.
  • Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a processor, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, some one or more of the most important method steps may be executed by such an apparatus.
  • embodiments of the invention can be implemented in hardware or in software.
  • the implementation can be performed using a non- transitory storage medium such as a digital storage medium, for example a floppy disc, a DVD, a Blu-Ray, a CD, a ROM, a PROM, and EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
  • Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
  • embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer.
  • the program code may, for example, be stored on a machine readable carrier.
  • inventions comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
  • an embodiment of the present invention is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
  • a further embodiment of the present invention is, therefore, a storage medium (or a data carrier, or a computer-readable medium) comprising, stored thereon, the computer program for performing one of the methods described herein when it is performed by a processor.
  • the data carrier, the digital storage medium or the recorded medium are typically tangible and/or non-transitionary.
  • a further embodiment of the present invention is an apparatus as described herein comprising a processor and the storage medium.
  • a further embodiment of the invention is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein.
  • the data stream or the sequence of signals may, for example, be configured to be transferred via a data communication connection, for example, via the internet.
  • a further embodiment comprises a processing means, for example, a computer or a programmable logic device, configured to, or adapted to, perform one of the methods described herein.
  • a further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
  • a further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver.
  • the receiver may, for example, be a computer, a mobile device, a memory device or the like.
  • the apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
  • a programmable logic device for example, a field programmable gate array
  • a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein.
  • the methods are preferably performed by any hardware apparatus.
  • Embodiments may be based on using a machine-learning model or machine-learning algorithm.
  • Machine learning may refer to algorithms and statistical models that computer systems may use to perform a specific task without using explicit instructions, instead relying on models and inference.
  • a transformation of data may be used, that is inferred from an analysis of historical and/or training data.
  • the content of images may be analyzed using a machine-learning model or using a machine-learning algorithm.
  • the machine-learning model may be trained using training images as input and training content information as output.
  • the machine-learning model "learns” to recognize the content of the images, so the content of images that are not included in the training data can be recognized using the machine-learning model.
  • the same principle may be used for other kinds of sensor data as well: By training a machine-learning model using training sensor data and a desired output, the machine-learning model "learns” a transformation between the sensor data and the output, which can be used to provide an output based on non-training sensor data provided to the machine-learning model.
  • the provided data e.g. sensor data, meta data and/or image data
  • Machine-learning models may be trained using training input data.
  • the examples specified above use a training method called "supervised learning".
  • supervised learning the machine-learning model is trained using a plurality of training samples, wherein each sample may comprise a plurality of input data values, and a plurality of desired output values, i.e. each training sample is associated with a desired output value.
  • the machine-learning model "learns" which output value to provide based on an input sample that is similar to the samples provided during the training.
  • semi-supervised learning may be used. In semi-supervised learning, some of the training samples lack a corresponding desired output value.
  • Supervised learning may be based on a supervised learning algorithm (e.g.
  • Classification algorithms may be used when the outputs are restricted to a limited set of values (categorical variables), i.e. the input is classified to one of the limited set of values.
  • Regression algorithms may be used when the outputs may have any numerical value (within a range).
  • Similarity learning algorithms may be similar to both classification and regression algorithms but are based on learning from examples using a similarity function that measures how similar or related two objects are.
  • unsupervised learning may be used to train the machine-learning model. In unsupervised learning, (only) input data might be supplied and an unsupervised learning algorithm may be used to find structure in the input data (e.g.
  • Clustering is the assignment of input data comprising a plurality of input values into subsets (clusters) so that input values within the same cluster are similar according to one or more (pre-defined) similarity criteria, while being dissimilar to input values that are included in other clusters.
  • Reinforcement learning is a third group of machine-learning algorithms.
  • reinforcement learning may be used to train the machine-learning model.
  • one or more software actors (called “software agents") are trained to take actions in an environment. Based on the taken actions, a reward is calculated.
  • Reinforcement learning is based on training the one or more software agents to choose the actions such, that the cumulative reward is increased, leading to software agents that become better at the task they are given (as evidenced by increasing rewards).
  • Feature learning may be used.
  • the machine-learning model may at least partially be trained using feature learning, and/or the machine-learning algorithm may comprise a feature learning component.
  • Feature learning algorithms which may be called representation learning algorithms, may preserve the information in their input but also transform it in a way that makes it useful, often as a pre-processing step before performing classification or predictions.
  • Feature learning may be based on principal components analysis or cluster analysis, for example.
  • anomaly detection i.e. outlier detection
  • the machine-learning model may at least partially be trained using anomaly detection, and/or the machine-learning algorithm may comprise an anomaly detection component.
  • the machine-learning algorithm may use a decision tree as a predictive model.
  • the machine-learning model may be based on a decision tree.
  • observations about an item e.g. a set of input values
  • an output value corresponding to the item may be represented by the leaves of the decision tree.
  • Decision trees may support both discrete values and continuous values as output values. If discrete values are used, the decision tree may be denoted a classification tree, if continuous values are used, the decision tree may be denoted a regression tree.
  • Association rules are a further technique that may be used in machine-learning algorithms.
  • the machine-learning model may be based on one or more association rules.
  • Association rules are created by identifying relationships between variables in large amounts of data.
  • the machine-learning algorithm may identify and/or utilize one or more relational rules that represent the knowledge that is derived from the data. The rules may e.g. be used to store, manipulate or apply the knowledge.
  • Machine-learning algorithms are usually based on a machine-learning model.
  • the term "machine-learning algorithm” may denote a set of instructions that may be used to create, train or use a machine-learning model.
  • the term "machine-learning model” may denote a data structure and/or set of rules that represents the learned knowledge (e.g.
  • the usage of a machine-learning algorithm may imply the usage of an underlying machine-learning model (or of a plurality of underlying machine-learning models).
  • the usage of a machine-learning model may imply that the machine-learning model and/or the data structure/set of rules that is the machine-learning model is trained by a machine-learning algorithm.
  • the machine-learning model may be an artificial neural network (ANN).
  • ANNs are systems that are inspired by biological neural networks, such as can be found in a retina or a brain.
  • ANNs comprise a plurality of interconnected nodes and a plurality of connections, so-called edges, between the nodes.
  • Each node may represent an artificial neuron.
  • Each edge may transmit information, from one node to another.
  • the output of a node may be defined as a (non-linear) function of its inputs (e.g. of the sum of its inputs).
  • the inputs of a node may be used in the function based on a "weight" of the edge or of the node that provides the input.
  • the weight of nodes and/or of edges may be adjusted in the learning process.
  • the training of an artificial neural network may comprise adjusting the weights of the nodes and/or edges of the artificial neural network, i.e. to achieve a desired output for a given input.
  • the machine-learning model may be a support vector machine, a random forest model or a gradient boosting model.
  • Support vector machines i.e. support vector networks
  • Support vector machines are supervised learning models with associated learning algorithms that may be used to analyze data (e.g. in classification or regression analysis).
  • Support vector machines may be trained by providing an input with a plurality of training input values that belong to one of two categories. The support vector machine may be trained to assign a new input value to one of the two categories.
  • the machine-learning model may be a Bayesian network, which is a probabilistic directed acyclic graphical model.
  • a Bayesian network may represent a set of random variables and their conditional dependencies using a directed acyclic graph.
  • the machine-learning model may be based on a genetic algorithm, which is a search algorithm and heuristic technique that mimics the process of natural selection. It is further understood that the disclosure of several steps, processes, operations or functions disclosed in the description or claims shall not be construed to imply that these operations are necessarily dependent on the order described, unless explicitly stated in the individual case or necessary for technical reasons. Therefore, the previous description does not limit the execution of several steps or functions to a certain order. Furthermore, in further examples, a single step, function, process or operation may include and/or be broken up into several substeps, -functions, -processes or -operations.
  • aspects described in relation to a device or system should also be understood as a description of the corresponding method.
  • a block, device or functional aspect of the device or system may correspond to a feature, such as a method step, of the corresponding method.
  • aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Multimedia (AREA)
  • General Physics & Mathematics (AREA)
  • Physics & Mathematics (AREA)
  • Evolutionary Computation (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Medical Informatics (AREA)
  • Software Systems (AREA)
  • Databases & Information Systems (AREA)
  • Computing Systems (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Artificial Intelligence (AREA)
  • Traffic Control Systems (AREA)

Abstract

The present disclosure relates to an image synthesis apparatus, the image synthesis apparatus comprising a first interface configured to receive an input image of a scene comprising objects of different object types. Further, the present disclosure relates to an image synthesis apparatus comprising a semantic segmentation network configured to map the input image to a segmentation map of the scene, wherein each pixel of the segmentation map is assigned to one of the different object types, and an object detection network configured to map the input image of the scene to one or more object locations of detected objects, wherein each object location is assigned one object type. The image synthesis further comprises an image completion network configured to merge the one or more detected objects into the segmentation map of the scene based on the respective object locations and object types to obtain a modified segmentation map and to synthesize an output image based on the modified segmentation map.

Description

IMAGE SYNTHESIS APPARATUS AND METHOD
Field
The present disclosure relates to augmented reality, particularly for a windshield of a vehicle or smartglasses.
Background
Augmented reality systems offer a real-time interactive experience that displays computergenerated content with a matching alignment to a real-world view. Virtual content can be constructive, wherein content is added to the view, or destructive, wherein content is reduced or removed from the view. The virtual content may often be seamlessly interwoven with the real -world view, such that its experience gives a more natural and intuitive feel for the viewer. This may play an important role in safety for certain applications, such as driving.
A human driver is prone to making mistakes. Reasons could be a distraction, a blind spot or reacting too late to a danger. For example, some accidents happen because drivers are distracted by advertisements or by accidents on another side of a highway. Augmented reality systems, particularly heads-up displays, are well established to provide the driver with useful information related to a planned route or a surrounding environment. While offering useful information, it is also the case that alarms, messages, and additional monitor bounding boxes may cause additional distraction for the driver and might actually create more danger, as opposed to reducing it.
People with impaired vision are often unable to drive because they cannot perceive important objects as well as they need to, including road signs or other cars, particularly in challenging weather conditions. With the assistance of an augmented reality system offering an intuitive driving experience, people with ranging abilities in vision and driving may be assisted to drive more safely. Other aspects of daily life may also benefit from augmented reality systems. A more intuitive augmented reality system may help visually-impared people to read smaller text or read text from a farther distance. It may also help them follow the movement of small objects. The use of such a system could be so intuitive that the user forgets that there is an interface, encouraging further use for maintaining safety or overcoming challenges related to vision.
Thus, there is a demand for an augmented reality system for improving safety when driving a vehicle and in assisting visually-impaired people in their daily lives. Non-visually impaired people may also find new, useful applications from such an augmented reality system.
Summary
According to a first aspect, the present disclosure relates to an image synthesis apparatus. The image synthesis apparatus comprises a first interface configured to receive an input image of a scene comprising objects of different object types. The image synthesis apparatus further comprises a semantic segmentation network configured to map the input image to a segmentation map of the scene, wherein each pixel of the segmentation map is assigned to one of the different object types. The image synthesis apparatus further comprises an object detection network configured to map the input image of the scene to one or more object locations of detected objects, wherein each object location is assigned one object type. The image synthesis apparatus further comprises an image completion network configured to merge the one or more detected objects into the segmentation map of the scene based on the respective object locations and object types to obtain a modified segmentation map and to synthesize an output image based on the modified segmentation map.
According to a further aspect, the present disclosure relates to an image synthesis method. The image synthesis method includes receiving an input image of a scene comprising objects of different object types. The image synthesis method further includes mapping, using a semantic segmentation network, the input image to a segmentation map of the scene, wherein each pixel of the segmentation map is assigned to one of the different object types. The image synthesis method further includes detecting, using an object detection network, one or more objects in the scene, wherein each detected object has associated therewith a respective object location and object type. The image synthesis method further includes modifying the segmentation map of the scene based on the one or more detected objects and synthesizing an output image based on the modified segmentation map.
According to yet a further aspect, the present disclosure also relates to a computer program having computer-readable instructions for carrying out the above method, when the computer program is executed on a programmable hardware device.
Brief description of the Figures
Some examples of apparatuses and/or methods will be described in the following by way of example only, and with reference to the accompanying figures, in which
Fig. 1 shows a scene as viewed through a windshield of a vehicle including buildings, a traffic light, multiple pedestrians, multiple other vehicles, an advertisement, a pothole, and a skyline;
Fig. 2 shows a block diagram of an apparatus for image synthesis according to a first embodiment;
Fig. 3 shows a semantic segmentation map of the scene in Fig. 1 generated by a semantic segmentation network, wherein the objects listed above among others are assigned an object type to generate the semantic segmentation map;
Fig. 4 shows a labeled input image of the scene in Fig. 1 with labels of objects at corresponding object locations that have been detected by an object detection network;
Fig. 5 shows a modified version of the segmentation map of Fig. 3 that has been modified by an image completion network based on information provided by the object detection network;
Fig. 6 shows the scene of Fig. 1 in an augmented reality form as viewed on the windshield of the vehicle, wherein certain objects have been enhanced or removed depending on their features; Fig. 7A shows a highway scene as viewed through a windshield of a vehicle;
Fig. 7B shows the highway scene in an augmented reality form as viewed on the windshield of the vehicle, wherein an input image based on the highway scene has been modified to a synthesized output image with an enlarged road sign and with a view of a blindspot;
Fig. 8 A shows a sport event scene as viewed by means of a television signal;
Fig. 8B shows the sport event scene of the television signal in an augmented reality form with an enlarged ball to demonstrate how the image synthesis apparatus can assist people in other daily activities, particularly through the use of smartglasses;
Fig. 9 shows an image synthesis method according to a first embodiment; and
Fig. 10 shows a functional block diagram for an image synthesis apparatus with a first input from a first interface and a second input from a second interface for a vehicle setting;
Detailed Description
Some examples are now described in more detail with reference to the enclosed figures. However, other possible examples are not limited to the features of these embodiments described in detail. Other examples may include modifications of the features as well as equivalents and alternatives to the features. Furthermore, the terminology used herein to describe certain examples should not be restrictive of further possible examples.
Throughout the description of the figures same or similar reference numerals refer to same or similar elements and/or features, which may be identical or implemented in a modified form while providing the same or a similar function. The thickness of lines, layers and/or areas in the figures may also be exaggerated for clarification. When two elements A and B are combined using an “or”, this is to be understood as disclosing all possible combinations, i.e. only A, only B as well as A and B, unless expressly defined otherwise in the individual case. As an alternative wording for the same combinations, "at least one of A and B" or "A and/or B" may be used. This applies equivalently to combinations of more than two elements.
If a singular form, such as “a”, “an” and “the” is used and the use of only a single element is not defined as mandatory either explicitly or implicitly, further examples may also use several elements to implement the same function. If a function is described below as implemented using multiple elements, further examples may implement the same function using a single element or a single processing entity. It is further understood that the terms "include", "including", "comprise" and/or "comprising", when used, describe the presence of the specified features, integers, steps, operations, processes, elements, components and/or a group thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, processes, elements, components and/or a group thereof.
Improving driver safety has been a public concern since the adoption of the vehicle. Reducing distractions and dangers, while also maintaining an intuitive driver experience is possible through an augmented reality (AR) system. As opposed to a heads-up display that displays information in a way that could be distracting to a driver, an AR system can seamlessly integrate detected objects from an input image of a scene captured e.g. in front of the vehicle into a synthesized output image viewed by the driver. If the objects in the output image can be enhanced or removed by an image synthesis apparatus and then presented as an intuitive AR experience, driving safety may be dramatically improved.
Fig- 1 shows a scene 100a from the view of a driver through a windshield of a vehicle while driving on a road with traffic. The scene 100a includes buildings 1-1, 1-2, and 1-3, a traffic light 2-1, multiple pedestrians 3-1, 3-2, and 3-3, multiple other vehicles 4-1 and 4-2, an advertisement 5-1, a pothole 6-1, a skyline 8, a sky 9, grass 10, and a road 11. Some objects of the scene 100a are unlikely to cause a distraction or present a danger. For example, the sky 9, the grass 10, and the road 11 are all immobile objects that usually do not require special attention from the driver. Other objects may be more important for the driver to keep in mind. For example, the traffic light 2-1 and multiple pedestrians 3-1, 3-2, and 3-3 crossing the street at an intersection may be such important objects. Other objects may indeed cause a distraction or present a danger to the driver. For example, there is a pothole 6-1 in the road, presenting a danger. This danger may be challenging to avoid, especially since a pothole is often too small to be seen from far away and its view may be obstructed by other vehicles. The advertisement 5-1 may cause a distraction for the driver. If a distraction is present, other dangerous objects, such as the pothole 6-1, may become even more dangerous since the driver may not react in time with less attention on the road. The scene 100a also has a skyline 8 in the background. This may also be considered distracting since it may divert attention from the driver. An AR experience while driving can help avoid dangerous situations, which may arise quickly and unexpectedly in any typical vehicle setting, such as the scene 100a depicted in Fig. 1.
Fig- 2 shows a block diagram of an image synthesis apparatus 200, which can provide such an AR experience in accordance with the present disclosure. The image synthesis apparatus 200 comprises an interface 210 configured to receive an input image 212 of a scene 100 comprising objects of different object types. The scene 100 may be the scene 100a or another scene. The image synthesis apparatus 200 may optionally comprise a second interface 214, which will be discussed in Fig. 7A, 7B, and 10. The scene 100 may be captured by one or more environmental sensors, such as a camera, radar, LiDAR, or combinations thereof, leading to the generation of the input image 212 that may be sent to the interface 210. The input image 212 may comprise at least one of a camera image, a radar image, and a LiDAR image. Under certain conditions, a radar or LiDAR sensor may provide visual information of the scene 100 that a camera may fail to obtain, such as in foggy or dark conditions.
The input image 212 may depict a scene similar to the scene 100a in Fig. 1 and may thus comprise objects relevant to the driver that may be organized by pre-defined object types relevant to driving. In general, an object type may be any particular group of objects that share similar characteristics. Such pre-defined object types may be useful to organize visual information within an image. In embodiments related to a vehicle setting, such as the scene 100a, pre-defined object types may include “building”, “traffic light”, “pedestrian”, “vehicle”, “advertisement”, “pothole”, “road sign”, “skyline”, “sky”, “grass”, and “road”, etc. Also, separate objects of an object type may appear in an image, each with a respective location. In the scene 100a, pedestrians 3-1, 3-2, and 3-3 can each be labeled as a separate object of the same object type, each corresponding to a different location. The image synthesis apparatus 200 comprises a semantic segmentation network 220 that is configured to receive the input image 212 from the interface 210 and to map the input image 212 to a segmentation map 222 of the scene 100. Each pixel of the segmentation map 222 is assigned to one of the different object types.
The image synthesis apparatus 200 further comprises an object detection network 230 that is configured to receive the input image 212 from the interface 210 and to map the input image 212 of the scene 100 to one or more object locations of detected objects. Each object location is assigned one object type. The object detection network 230 may be configured to generate a labeled input image 232 of the scene 100 with labels of objects at corresponding object locations. The object detection network 230 may comprise an artificial neural network, for example.
The image synthesis apparatus 200 further comprises an image completion network 240 configured to merge the one or more detected objects of the labeled input image 232 into the segmentation map 222 of the scene 100 based on the respective object locations and object types to obtain a modified segmentation map 244 and to synthesize an output image 248 based on the modified segmentation map 244. The image completion network 240 may comprise an artificial neural network, for example. The image completion network 240 may comprise a map modifier unit 242 and an image output unit 246. The map modifier unit 242 may be configured to receive the segmentation map 222 and the labeled input image 232, to generate the modified segmentation map 244 based on the segmentation map 222 and the labeled input image 232, and to send the modified segmentation map 244 to the image output unit 246. The image output unit 246 may be configured to receive the modified segmentation map 244 and to synthesize the output image 248 based on the modified segmentation map 244.
The image synthesis apparatus 200 may comprise a display 250 configured to display the synthesized output image 248 to a user. The image output unit 246 of the image completion network 240 may be configured to send the output image 248 to the display 250. The display 250 may be configured to display the synthesized output image 248 on a transparent member, for example. The transparent member may be the windshield of a vehicle, a pair of glasses, or another means to display the synthesized output image 248. The display 250 may also comprise a structure that is not transparent, such as a projection screen or another flat object that enables a user to view the synthesized output image 248. In a first example, the scene 100a may be captured by one or more environmental sensors to generate a first input image 212a. The input image 212a may be received by the interface 210 and sent to the semantic segmentation network 220 and the object detection network 230 to generate a corresponding segmentation map 222a and a corresponding labeled input image 232a. These may be received by the image completion network 240 to generate a corresponding modified segmentation map 244a and a corresponding synthesized output image 248a, which may be displayed on the display 250. In a second example, a second scene 100b may lead to the synthesis of a second output image 248b, with all corresponding maps and images generated. This will be discussed in Fig. 7A and 7B. In a third example, a third scene 100c may lead to the synthesis of a third output image 248c, with all corresponding maps and images generated. This will be discussed in Fig. 8A and Fig. 8B.
The image synthesis apparatus 200 may repeat the procedure of receiving input images 212 and synthesizing corresponding output images 248. In this way, a stream of synthesized output images 248 may be generated based on a stream of input images 212. The stream of synthesized output images 248 may be used to generate an AR expereince for a user. Unlike traditional solutions in driving, such as heads-up displays, the AR experience provided by the image synthesis apparatus 200 cannot be easily ignored by the driver. It enables an intuitive interaction because it augments reality, enabling the user’s natural reflexes to be activated. The features of the image synthesis apparatus 200 that enables the transformation of the input image 212 to the synthesized output image 248 to be displayed on the display 250 are described in further detail below.
The transformation of the input image 212 to the synthesized output image 248 includes semantic segmentation of the input image 212. The process of semantic segmentation is performed by the semantic segmentation network 220. The process of semantic segmentation includes a categorization of each pixel of an image to an object type according to multiple relevant object types. The object types may be pre-defined according to a setting of the image, such as a vehicle setting as previously described. Each pixel of the image may be assigned to an object type based on the context of the surrounding pixels and the entire image. With each pixel of the image categorized according to the pre-defined object types, the image can be manipulated more easily towards a useful purpose. Semantic segmentation performed within the means of conventional engineering requires an enormous domain knowledge database and enormous computational time. Using an artificial neural network as the semantic segmentation network 220 to perform semantic segmentation can overcome such limitations by modeling the domain knowledge from a dataset of labeled pixels. Such a neural network may be a convolutional neural network (CNN), which is applied to analyze visual imagery and uses relatively little pre-processing compared to other image classification algorithms. A CNN includes multiple layers, including convolutional layers, pooling layers, and fully-connected (FC) layers. Progressing through each layer, features of the input image 212 are identified, such as colors and edges, and eventually larger elements or shapes of the object, until it finally identifies the intended object.
Each convolutional layer comprises a feature detector, also known as a filter (or kernel), which can perform a process known as a convolution by moving across the receptive fields of the image 212, checking if s specific feature is present. More specifically, a convolutional neural network learns to optimize the filter through automated learning, whereas in traditional algorithms the filters are hand-engineered. This independence from prior knowledge and human intervention in feature extraction is a major advantage for a convolutional neural network. Examples of convolutional neural network architectures that may be used are U- Net, AlexNet, VGGNet, GoogLeNet, ResNet, and ZFNet. To gain such advantages, the semantic segmentation network 220 may be a neural network, particularly a convolutional neural network.
The semantic segmentation network 220 can be trained with known input images to categorize objects to pre-defined object types. Semantic segmentation networks are often trained to differentiate objects in a foreground from a background. Images with a specific context, such as a vehicle setting, often have repeating object types in the foreground and background and the semantic segmentation network 220 can be trained to recognize and categorize such objects in the input image 212. For example, traffic lights are usually associated with a foreground and a sky is usually associated with a background, and the semantic segmentation network 220 can be trained to recognize traffic lights and a sky and to recognize the boundary between the two. With the recognition of the boundary, semantic segmentation can also lead to the determination of the shape of an object in the input image 212. While a traffic light may be irrelevant or categorized broadly in other settings, it may be important in a driving context and can be given its own object type. Fig- 3 shows an example segmentation map 222a of the input image 212a based on the scene 100a. Each pixel of the scene 100 may be categorized by the semantic segmentation network 220 to an object type that was chosen for a specified setting, such as the vehicle setting in the scene 100a. Example object types for the scene 100a, include buildings as object type 1, traffic lights as object type 2, pedestrians as object type 3, vehicles as object type 4, advertisements as object type 5, potholes as object type 6, road signs as object type 7, skyline as object type 8, sky as object type 9, grass as object type 10, and roads as object type 11. Each number depicted represents a pixel that has been categorized to an object type of that number. The segmentation map 222a shows how certain areas have been labeled according to the object type that is depicted in the corresponding area according to the input image 212a of the scene 100a. The semantic segmentation network 220 did not label any pixel as object type 7 because there was no road sign in the scene 100a. Not every area of the segmentation map 222a in Fig. 3 is labeled with an object type number for clarity purposes. The skilled person benefitting from the present disclosure will appreciate that each pixel of the input image 212 can be categorized to produce a segmentation map 222, wherein Fig. 3 does not illustrate every area with a labeled number. Rather, Fig. 3 is a depiction to guide in understanding the process of semantic segmentation. Beyond semantic segmentation of the input image 212, the transformation of input image 212 to synthesized output image 248 also includes a process of detection of objects in the input image 212. The process of detection of objects (object detection) is performed by the object detection network 230.
Fig- 4 shows an example of a labeled input image 232a based on the input image 212a of the scene 100a with labels of objects at their corresponding object locations that have been detected by the object detection network 230. The object detection network 230 may be configured to detect objects that are relevant in a vehicle setting, wherein each has associated therewith an object location. For example, in the input image 212a, pedestrians 3-1, 3-2, and 3-3 may be separate detected objects, each of object type 3 and each with a respective object location. For example, the object location may be expressed with 2-dimensional coordinates. Vehicles 4-1 and 4-2 may be detected as separate objects of object type 4, while buildings 1- 1, 1-2, and 1-3 may be detected as separate objects of object type 1, each with a respective object location. Objects that are the only instance within their object type in the input image 212a, such as the traffic light 2-1, advertisement 5-1, and pothole 6-1, may also be detected objects with a respective object location. The skyline 8, the sky 9, the grass 10, and the road 11 do not have associated therewith a detected object.
The object detection network 230 may also be configured as a neural network, including a convolutional neural network, as described above. The semantic segmentation network 220 and the object detection network 230 may be trained in tandem, such that any object in any input image 212 will be categorized into the same object type by both networks. Alternatively, the semantic segmentation network 220 may provide information related to how the input image 212 has been categorized according to the object types while generating the segmentation map 222. The image synthesis apparatus 200 may comprise one or more neural networks configured to perform semantic segmentation and/or object detection.
Images generated directly from a segmentation map may be color-coded, with a unique color corresponding to each object type. In such a color-coded image that comprises multiple detected objects of an object type, the detected objects of the object type might not be distinguished between each other. Also, if an object has multiple colors and shades portraying physical details of its structure in the input image 212, this information would be lost in the output image 248 based on the segmentation map 222. For example, for the input image 212a of the scene 100a, the vehicles 4-1 and 4-2 of object type 4 would not only lose many details, including brake lights and windows, but they would be represented by the same color and would have no visible feature to distinguish them other than their separate locations. The same is true for the buildings 1-1 to 1-3 and the pedestrians 3-1 to 3-3. In order to present an AR experience that enables the driver to intuitively react to what is seen, it may be important that some detected objects maintain the detail of the input image 212 as received by the interface 210. For example, a driver may wish to see pedestrians or other vehicles as normally seen, whereas seeing a single color for all pedestrians or vehicles would lead to a less intuitive interaction.
For this, the image completion network 240 may be configured to receive information related to the input image 212 from both the semantic segmentation network 220 and the object detection network 230 and to synthesize an output image 248, wherein one or more detected objects are merged into the segmentation map 222 according to their respective object locations. The image completion network 240 may be configured to merge the one or more detected objects into the segmentation map 222 of the scene 100 on the condition that the object type of the respective detected object meets a predetermined criterion of importance. The pre-determined criterion of importance of an object may include the object presenting a danger to the driver and an object presenting relevant information for driving safety. For example, in the scene 100a, the pothole 6-1 may present a danger. A danger may also be a possible future danger to the driver. This may include pedestrians 3-1 to 3-3, for example. The pedestrians 3-1 to 3-3 may not be dangerous crossing the road while the traffic light 2-1 is red but may become dangerous if still crossing when the traffic light 2-2 turns green. Relevant information for driving safety may be anything that guides the driver to maintain a safe interaction with other vehicles and the entire surrounding. This may include the traffic light 2-1 and road signs. Relevant information may also include physical obstacles to be avoided, whether on or off the road. This may include buildings 1-1 to 1-3 off the road and vehicles 4-1 and 4-2 on the road. If only detected objects meeting the pre-determined criterion of importance are included in the output image 248, the driver only need focus on these obj ects and distracting objects that have no value when seen may be removed out of view.
The image completion network 240 may also be configured as an artificial neural network, including a convolutional neural network, as described above. The image completion network 240 may be trained in tandem with the semantic segmentation network 220 and/or the object detection network 230. While the training for the semantic segmentation network 220 and the object detection network 230 may focus more on categorization and feature extraction, the image completion network 240 may focus more on how to enhance or remove objects based on a category or feature of an object. For example, the image completion network 240 may be trained by manually constructed output images that have had objects manually enhanced or removed based on a category or feature of the object, which may have been determined based on segmentation maps 222 generated by the semantic segmentation network 220 and labeled input images 232 generated by the object detection network 230. The manually constructed output images may provide multiple different contexts or situations to train the image completion network 240 to enhance or remove objects depending on the respective context or situation. For example, in a vehicle setting, an object may be enhanced differently or may or may not be removed depending on factors related to the vehicle, such as whether the vehicle is in motion, or depending on factors related to the surrounding, such as weather conditions or darkness. The image synthesis apparatus 200 may comprise one or more neural networks configured to perform semantic segmentation, object detection and/or image completion. The image completion network 240 may include the map modifier unit 242 and the image output unit 246. The map modifier unit 242 may be a neural network of the image completion network 240 that is configured to modify the segmentation map 222 as described above. The map modifier unit 242 may be trained to complete an image based on input from the semantic segmentation network 220 and the object detection network 230 and manually constructed output images, as described above. The image output unit 246 may be a neural network configured to synthesize an output image 248 based on the modified segmentation map 244. The image output unit 246 may be configured to synthesize contents so that the output image 248 looks visually realistic. This may include traditional patch-based methods or deep learning-based methods. It may be a convolutional neural network that includes a spatially adaptive normalization (SPADE) layer, known to synthesize photorealistic images from a segmentation map by minimizing the loss of information by normalization layers. It may include a method of compositing the output image 248 by linearly blending it with the input image 212. Other methods of image synthesis from the segmentation map 222 may include a generative adversarial network (GAN) or a generative transformer model. In particular, a bidirectional transformer model for image synthesis offers the possibility to generate new tokens from previously generated tokens in all directions, as opposed to a sequence form, leading to significantly faster processing.
Fig- 5 shows an example of a modified segmentation map 244a based on the segmentation map 222a of Fig. 3, each based on scene 100a. Multiple detected objects from the labeled input image 232 may be merged into the modified segmentation map 244 by the image completion network 240. The image completion network 240 may be configured to merge the detected objects into the modified segmentation map 244 based on the object locations in the labeled input image 232 from the object detection network 230 and based on the segmentation map 222 from the semantic segmentation network 220. The image completion network 240 may also be configured to merge the detected objects with all corresponding details related to their physical structure as captured in the input image 212 into the modified segmentation map 244. In the illustrated example of the modified segmentation map 244a in Fig. 5, the merged detected objects from the labeled input image 232a are the traffic light 2-1, pedestrians 3-1, 3-2, and 3-3, vehicles 4-1 and 4-2, and the pothole 6-1. Some of the merged detected objects may have one or more characteristics changed, including the size, location, and contrast. The merged detected objects may be labeled accordingly. Vehicles 4-1 and 4-2 and pedestrian 3-3 may be labeled without modification, since they have maintained the same size, location, and contrast. Pedestrians 3-1 and 3-2 may be labeled with an “s” as 3-1 s and 3- 2s for having a changed size. Traffic light 2-1 may be labeled with an “s” and an “1” as 2-1 si for having a changed size and a changed location. Pothole 6-1 may be labeled with a “c” for having a changed contrast as 6-lc. Also, associated with pothole 6-lc is a separate enlarged image box, 6-lcslb, labeled with a “c”, “s”, “1”, and “b” for having a changed contrast, size, and location, and having a separate box. The detected objects that are merged into the modified segmentation map 244a in Fig. 5 are labeled with parentheses, such as (4-1) for the merged detected object 4-1, to differentiate these labels from numbers labeled as examples of pixels corresponding to an object type, such as “4”.
Other detected objects from the scene 100a may not be included in the modified segmentation map 244a, such as buildings 1-1, 1-2, and 1-3 and advertisement 5-1. Pixels corresponding to the skyline 8 have had their object type changed, such that no pixels in the modified segmentation map 244a correspond to the skyline 8 and no skyline will be found within the output image 248a. Other pixels have been changed from one object type to another based on changes in size of a merged detected object.
The following paragraphs outline how the image completion network 240 may be configured to merge detected objects detected by the object detection network 230 into the segmentation map 222 generated by the semantic segmentation network 220 with modifications to the size, location, and/or contrast of the detected objects, depending on characteristics of the detected object and the context of the scene 100. Each example of how a detected object is merged into or removed from the segmentation map 222 to obtain a modified segmentation map 244 serves to illustrate how the image completion network 240 of the image synthesis apparatus 200 may be configured.
In the example of scene 100a, Vehicles 4-1 and 4-2 may be objects of interest to the driver because the driver must keep track of them to avoid a collision. Their corresponding pixels have been categorized into object type 4 in the segmentation map 222a. While the vehicles could be clearly seen and distinguished as a single-color output by the segmentation map 222a, seeing the actual form of the detected objects as captured in the input image 212a can better maintain an intuitive interaction with the road while driving. More importantly, the two vehicles 4-1 and 4-2 may become difficult to distinguish from each other in subsequent input images 212a of the scene 100a, which would lead to a dangerous situation. For example, if one vehicle merged into the same lane as the other vehicle, it may be partially blocked by the other and it may become difficult to distinguish the two. As such, any vehicle driving on the road would meet the predetermined criterion of importance for the image completion network 240 to merge the detected object into the segmentation map 222 so that the output image 248 includes the structural details of each vehicle as captured by the input image 212.
In addition to being merged into the segmentation map 222, such detected objects can also be merged in an enhanced form, depending on its object type and object location. The image completion network 240 may be configured to enhance the one or more detected objects and to merge the one or more enhanced objects into the segmentation map 222 of the scene 100 based on the respective object locations and object types. For example, the image completion network 240 may be configured to change a size of the one or more detected objects and to merge the one or more objects with the changed size into the segmentation map 222 of the scene 100 based on the respective object locations and object types. More specifically, the image completion network 240 may be configured to enlarge the one or more detected objects. This may present the detected object in the synthesized output image 248 in a form that is easier for the driver to perceive.
For the scene 100a, it may be advantageous to change the size, or particularly, to enlarge the form of the pedestrians 3-1, 3-2, and/or 3-3 in the synthesized output image 248a. If the pedestrians are presented in an enlarged form in the synthesized output image 248a, less attention of the driver will be required to perceive the pedestrians crossing. For this, the image completion network 240 may be configured to merge one or more detected objects into the segmentation map 222 in an enlarged form, wherein the number of pixels corresponding to the respective object to be enlarged is increased from its original number of pixels in the segmentation map 222 to a larger number of pixels in the modified segmentation map 244. The modified segmentation map 244 shows an example of how the image completion network 240 may be configured to selectively enhance one or more objects of an object type depending on the behavior and location of the respective object. The pedestrians 3- Is and 3 -2s have been enlarged from their original size because they are located in the middle of the road, while pedestrian 3-3 has maintained his original size because he is on the side of the road and standing. This principle may apply to other objects as well. For example, a vehicle that is determined to have associated therewith dangerous driving behavior may be merged and presented differently than other vehicles. The image completion network 240 may selectively enhance the pedestrians on the side of the road that are children, particularly playing children, to enable the driver to perceive the children more easily until passing them. The same may also be applied to animals. For example, a dog without a leash or a deer attempting to cross the road may exhibit unpredictable behavior and it may be important for the driver to perceive the animal more quickly to prevent an accident. The object detection network 230 may be a neural network that is trained to detect information related to the behavior and motion patterns of a mobile object, given multiple input images 212. The image completion network 240 may be configured to merge such an object in an enhanced form based on the location, behavior, or motion patterns of the object. This may help the driver see such a mobile object more quickly, thereby reducing the risk of injury/damage for both the driver and the object.
If a detected object only presents a distraction and provides no relevant information to the driver, it can be removed from the segmentation map 222 by the image completion network 240. If the advertisement 5-1 in scene 100a is determined to be such an object, for example, the image completion network 240 can reassign its corresponding pixels to another object type, such as those of a neighboring background. In the case of advertisement 5-1 in scene 100a, the pixels assigned to object type 5 by the semantic segmentation network 220 have been re-assigned to the object type of its neighboring background, the sky 9 and the grass 10, such that it is not seen in the corresponding synthesized output image 248a. To achieve this, the image completion network 240 may be configured to change the object type assigned to the one or more detected objects and to merge the one or more objects with the changed object type into the segmentation map 222 based on the respective object locations. More specifically, the pixels associated with an object to be removed may be re-assigned to one or more other object types, such as those of a neighboring background. Merging the detected object with the changed object type into the segmentation map 222 based on the respective object location may then result in the detected object to remain unseen in the synthesized output image 248. In the case of the modified segmentation map 244a, the driver will be unaware that the advertisement 5-1 is present in the scene 100a and any distraction associated with it will be prevented. Pixels that have been categorized by the semantic segmentation network 220 to an object type and do not have associated therewith any detected object may also be changed or removed by the image completion network 240. For example, the skyline 8 in scene 100a may also be considered a distraction for a driver, particularly for a driver that is unfamiliar with the geographical area. In the illustrated example of modified segmentation map 244a, all pixels corresponding to object type 8 for the skyline have been removed and replaced by pixels corresponding to object type 9 for the sky, or other object types associated with an enlargement and/or relocation of another detected object. For this, the image completion network 240 may be configured to change the object type of a pixel to any other object type. The changed object type may also be associated with a merged detected object.
In the illustrated example, the traffic light 2-1 in the segmentation map 222a was also enlarged by the image completion network 240 to be more easily perceived by the driver, as shown in modified segmentation map 244a, wherein the number of pixels corresponding to the traffic light 2-1 is increased to a larger number of pixels in the modified segmentation map 244a. If a certain portion of the traffic light 2-1 is determined to be irrelevant for the driver by the image completion network 240, such as the metal post supporting it, the image completion network 240 may remove this portion, as shown in the modified segmentation map 244a. Also, if the traffic light 2-1 is located at a distance or angle that cannot be easily perceived by the driver, the traffic light 2-1 can be merged to be in a different location in the modified segmentation map 244a where it may be more easily perceived by the driver. This may be at a more central location, as depicted in the modified segmentation map 244a. For this, the image completion network 240 may be configured to exchange pixels corresponding to a merged detected object with other pixels that are not associated with any other merged detected objects and are determined to not convey information relevant to the driver.
Certain objects on the road may be important for the driver to perceive with great detail while driving, such as potholes. While some potholes are large and must be avoided, other potholes are smaller and need not be avoided. This may be important in certain situations, such as needing to avoid another vehicle by driving near or over the pothole 6-1 in scene 100a. It may be useful for the driver to see the pothole 6-1 with a greater contrast compared to other objects to decide whether it must be avoided or not. For this, the image completion network 240 may be configured to synthesize the output image 248 with an increased contrast of the one or more detected objects compared to remaining regions of the output image. The pothole 6-1 in scene 100a is depicted as pothole 6-lc in the modified segmentation map 244a for being shown with greater contrast compared to the other detected objects (as can be seen in Fig. 6).
It may also be important that the driver perceives any pothole before driving near it and to perceive how far away it is located. Thus, while performing object detection, the object detection network 230 may also be configured to obtain distance information of one or more detected objects to be used by the image completion network 240. Such distance information may be obtained by comparing the size of the detected object to other objects in the input image 212 with a well-known size range, such as the width and height of other vehicles or the width of the road lanes. Additionally or alternatively, distance information may be obtained by using distance or ranging sensors, such as radar, ultrasonic, and/or lidar sensors. The image completion network 240 may be configured to change the size of the one or more detected objects based on distance information associated with the respective object. In order to see the exact location of the pothole as well as its structural detail, the image completion network 240 may be configured to merge an enlarged image of the detected object to appear as a separate image box in an appropriate location. The pothole 6-1 in scene 100a is depicted in the modified segmentation map 244a as 6-lc and also as an enlarged image box 6-lcslb at a separate location (also seen in Fig. 6). The separate location of such an enlarged box may preferably be a corner of the synthesized output image 248. For scene 100a, the separately located enlarged image box 6-lcslb combined with the contrasted depiction of the pothole 6- 1c at the original location may be useful for the driver to see exactly where the pothole is located and to simultaneously perceive it better.
The image completion network 240 may also be configured to not include certain detected objects detected by the object detection network 230 based on whether or not the semantic segmentation network 220 has provided enough information related to the detected object. For example, buildings 1-1, 1-2, and 1-3 in the scene 100a are detected objects in the labeled input image 232a and they present physical obstacles to the driver to be avoided. But since the semantic segmentation network 220 has already classified all pixels of objects 1-1, 1-2 and 1-3 to the object type 1, it is already possible to distinguish where the buildings are located in relation to the road. Also, the buildings are very stable structures and will not change their absolute location. The driver can thus easily perceive these locations to avoid a collision without having them merged into the corresponding segmentation map 222a. It may be beneficial to not include the buildings 1-1, 1-2, and/or 1-3 if they comprise details that may be distracting to the driver. The modified segmentation map 244a in Fig. 5 illustrates this, wherein the pixels corresponding to the object locations of detected objects 1-1, 1-2, and 1-3 are labeled with the object type 1, but not with detected objects.
If it is desirable to distinguish whether details of a detected object should be included or not in the output image 248, the object detection network 230 may be a neural network configured to learn to detect smaller objects corresponding to specified details within larger detected objects, such as buildings 1-1, 1-2, and 1-3 in scene 100a. The image completion network 240 may be configured to include such smaller objects within larger objects detected by the object detection network 230 into the modified segmentation map 244 based on a pre-determined criterion of importance. If a detected image is determined to be distracting but provides relevant information, the image completion network 240 may be configured to merge generic images of a simpler form including the relevant information into the modified segmentation map 244. This would convey the relevant information in a less distracting way to the driver. Configurations related to the inclusion or exclusion of detected objects may be customized according to the preferences of a user or to evolving safety standards.
Fig- 6 shows an example of the synthesized output image 248a of the scene 100a that may be viewed by the driver on the windshield of the vehicle or on glasses worn by the driver, for example. The synthesized output image 248a is based on the modified segmentation map 244a depicted by Fig. 5 that was modified from the segmentation map 222a depicted by Fig. 3. The detected objects 2-1, 3-1, 3-2, 3-3, 4-1, 4-2, and 6-1 were all merged into the modified segmentation map 244a by the image completion network 240 and are thus depicted in the output image 248a. While vehicles 4-1 and 4-2 and pedestrian 3-3 were merged with the same size and location, the pedestrians 3-1 and 3-2 were inserted in an enlarged form as 3- Is and 3-2s at the same location, the advertisement 5-1 was removed, the traffic light 2-1 was merged at a different location and in an enlarged form as 2-1 si, and the pothole 6-1 was merged with an increased contrast and the same size and location as 6-lc. The image completion network 240 also merged a separate enlarged image box 6-lcslb located in the bottom right corner of the modified segmentation map 244a. Detected objects 1-1, 1-2, and 1-3 (buildings) were not merged into the modified segmentation map 244a. The synthesized output image 248a comprises pixels categorized to object type 1 for buildings, object type 9 for the sky, object type 10 for grass, and object type 11 for the road, wherein pixels corresponding to object type 8 for the skyline have been removed. Referring now back to Fig. 2, the image synthesis apparatus 200 may optionally comprise a second interface 214 which is configured to receive additional information 216 beyond the input image 212 of the (first) interface 210. The additional information 216 may yield one or more additional object locations of additional objects which are not visible in the input image 212. Each additional object location is assigned one object type. The image completion network 240 may be configured to merge the one or more additional objects into the segmentation map 222 of the scene 100 based on the respective additional object locations and object types in order to obtain the modified segmentation map 244.
For example, the first interface 210 may be coupled to a first environmental sensor, such as a camera, to obtain the input image 212. The second interface 214 may be coupled to a different second environmental sensor and/or a communication network to obtain the additional information 216 yielding one or more additional object locations of additional objects. For example, the second interface 214 may be coupled to a second camera having a different field of view than the camera coupled to the first interface 210. Additionally or alternatively, the second interface 214 may be coupled to one or more other environmental sensors, such as radar and/or lidar sensors covering the same or a different field of view (e.g., a blind spot) as the camera coupled to the first interface 210. The additional information 216 received by the second interface 214 may comprise at least one of a camera image, a radar image, and a LiDAR image. Additionally or alternatively, the second interface 214 may be coupled to a wireless communication network, such as a WiFi- or mobile communications network, to receive the additional information 216 yielding one or more additional object locations of additional objects. In such an example, the additional information 216 may include real-time traffic information, weather information, information on points of interest, or the like. That is, the second interface 214 may be configured to receive the additional information 216 from at least one of a camera signal, a radar signal, a lidar signal, and a radio signal.
The image completion network 240 may be configured to merge the one or more additional objects into the segmentation map 222 of the scene 100 on a condition that the object type(s) of the respective additional object(s) meets a pre-determined criterion of importance. For example, the one or more additional objects can correspond to objects (e.g., persons, other vehicles, etc.) in a blind spot of the driver. The pre-determined criterion of importance for the second interface 214 may be defined the same or differently from the pre-determined criterion of importance for the first interface 210. As previously described for the first interface 210, the pre-determined criterion of importance of an object may include the object presenting a danger to the driver and an object presenting relevant information for driving safety. For example, if the second interface 214 can perceive a pedestrian crossing the street only if behind the vehicle, the pre-determined criterion of importance of the second interface 214 for merging the pedestrian into the modified segmentation map 244 can be based on a different requirement of distance from the vehicle, which is unlikely to be driven backwards. Other possible requirements may relate to the behavior and motion patterns of the pedestrian. The pre-determined criterion of importance for the second interface 214 may depend on its orientation in comparison to the first interface 210 and the type of input images 212 or the additional information 216 it may receive, among other factors.
Fig. 7A shows the second scene 100b, which is also from the perspective of a driver in a vehicle. An input image based on the scene 100b may be received by the first interface 210 and may then be sent to the semantic segmentation network 220 and the object detection network 230. The image synthesis apparatus 200 may generate for the scene 100b a segmentation map analogous to the segmentation map 222 and a labeled input image analogous to the labeled input image 232. These may be received by the image completion network 240, which may generate a modified segmentation map analogous to 244 and a synthesized output image 248b, shown in Fig. 7B. The synthesized output image 248b may then be displayed by the display 250.
Fig. 7B shows the synthesized output image 248b based on the scene 100b. A road sign is depicted in an enlarged form as 7-ls, which was not visible in the scene 100b. The enlarged road sign 7-ls, labeled as corresponding to object type 7 for road signs, shows a red frame around a white circle reading the number 30 in digits of black color. The number 30 may be a signal to the driver that he must slow down to a speed of 30 mph or 30 km/h in a specific range of a highway. While no such road sign is visible in the scene 100b, it may be that there was such a road sign specifying a speed limit that the vehicle has already driven past and the image synthesis apparatus 200 has maintained the road sign in view after passing. The image completion network 240 may be configured to maintain merged objects in future output images 248 in situations useful to the driver. For this, the image synthesis apparatus 200 may comprise a memory to temporarily store objects detected from the object detection network 230 and may be configured to merge such stored objects, as necessary. The enlarged road sign 7- Is may also be an example of the additional information 216 that may be received by the second interface 214 by a communication network. While such road signs have a limited physical range of communication to the driver, the image synthesis apparatus 200 may convey information to the driver in a much broader range. For example, if there is a specified speed limit in the area communicated to the driver as additional information 216, what is only communicated by the sign at one position may be continuously communicated by the image synthesis apparatus 200 for the duration of time that the driver is driving on a road or highway with the specified speed limit. This may be particularly useful in situations where the driver is unaware of driving at a speed beyond the speed limit. In such a case, the enlarged road sign 7- Is may be a generic version of a road sign chosen by the image synthesis apparatus 200 after receiving additional information 216 related to the speed limit.
The additional information 216 may also be received by the second interface 214 by means of a blindspot camera. In the synthesized output image 248b, a sub-image box 4-3b depicts a separate blindspot scene comprising another vehicle. The sub-image box 4-3b may be labeled with a “b” because it is shown in the form of a separate box. The blindspot scene may be captured by a camera or environmental sensor and provided to the optional second interface 214. The sub-image box 4-3b comprising the other vehicle may then be depicted in the output image 248b after having been merged into the corresponding segmentation map. The merging of a vehicle in such a blindspot scene into the corresponding segmentation map may be dependent on the pre-determined criterion of importance that may be defined for the second interface 214, as previously described. The image completion network 240 may be configured to merge an entire separate image received by the second interface 214, such as 4-3b, into a corresponding segmentation map in the form of a sub-image box. The image completion network 240 may also be configured to merge one or more objects detected within the separate image received by the second interface 214 alone into an appropriate location of the corresponding segmentation map. Objects detected from the second interface 214 and merged into the modified segmentation map 244, either as a sub-image box or as objects alone, may be labeled with text, including a brief description of context, as necessary.
As previously described, the second interface 214 may be configured to obtain the additional information 216 by means of the communication network related to an object that is not yet among the surroundings of the vehicle. In a further example, the vehicle may approach an object, such as a dangerous object blocking the road, and the object may be minutes away from being perceived by means of the first or second environmental sensor. The second interface 214 may be configured to obtain the additional information 216 in the form of GPS coordinates. The GPS coordinates may be the object’s GPS coordinates and the vehicle’s GPS coordinates. The relevant information 216 may also include information related to traffic and an average speed of the vehicle. The image synthesis apparatus 200 may be configured to calculate an approximate time in which the object may be perceived by the first or second environmental sensor. The image completion network 240 may be configured to portray generic versions of the object in the synthesized output image 248 based on the relevant information 216 provided to the second interface 214. For example, if the object has not yet been photographed, a synthetic generic image of the dangerous object may be inserted into the segmentation map 222 and viewed in the lower left comer, either in a box as portrayed in Fig. 8B or as a labeled object. The image completion network 240 may be configured to label the object with text, including a brief description of context, as necessary.
As previously described, the image completion network 240 may be configured to remove detected objects from view to prevent distraction to the driver. It may also be configured to display a detected object smaller if the object is determined to be a distraction. A benefit to displaying a detected object determined to be a distraction in a smaller form is that the driver does not have as much detail to focus on, but still gains an understanding of context for the vehicle surrounding. For example, it may explain why there is a certain amount of traffic or why other vehicles are driving at a certain speed, if applicable. In the example of a highway, an accident may happen because a driver is focused on a distraction, such as an accident pileup on the side of the highway. While the accident pile-up may be removed from view, it may be reduced in size in the field of view for the driver to gain an understanding of context.
Also, road reflections, or mirages, are a visual phenomenon often seen on a highway that may pose a distraction. More likely to be seen in very hot and sunny weather, this phenomenon is caused by atmospheric refraction of air. The asphalt of a highway may be much hotter than the air above it, so the image of something higher up can be refracted downward toward the hot asphalt and then quickly upward and it may appear similar to a pool of water. The image synthesis apparatus 200 may be configured to reduce the size of the portrayal of a physical phenomenon, such as a road reflection or a mirage, or remove it entirely. In reducing the size of the portrayal, the driver may still be aware of the conditions around the vehicle that led to the phenomenon, such as being aware that it is a hot and dry day, while being less distracted by it. Configurations of the image synthesis apparatus 200 that control how objects are enhanced, removed, or reduced in size may be customized according to driver preferences or evolving safety standards.
Embodiments of the present disclosure may not only be relevant to vehicle settings, but also to other applications, an example of which is illustrated in Fig. 8A.
Fig. 8A shows a scene 100c, which is a view of a television signal of a football match on a corresponding display. The scene 100c comprises a ball 12-1, which is a small object that may be difficult to follow during the match for a viewer, particularly if the viewer is visually impaired. In a further embodiment of the present disclosure, the first interface 210 may be configured to receive a television signal or any other visual signal as the input image 212. An input image based on the scene 100c may be received by the first interface 210 and may then be sent to the semantic segmentation network 220 and the object detection network 230. The image synthesis apparatus 200 may generate for the scene 100c a segmentation map analogous to the segmentation map 222 and a labeled input image analogous to the labeled input image 232. These may be received by the image completion network 240, which may generate a modified segmentation map analogous to 244 and a synthesized output image 248c, shown in Fig. 8B. The synthesized output image 248c may then be displayed by the display 250.
Fig- 8 shows the synthesized output image 248c of the scene 100c. In the output image 248c, the ball 12-1 is presented as a size-enhanced ball 12- 1 s. This demonstrates an example of how augmented reality can also assist a user of the image synthesis apparatus 200 in daily life beyond driving, particularly with a display on a pair of transparent glasses. Such an embodiment can be used in a wide variety of other settings to enhance objects. While displaying the output image on a windshield may be optimal for driving situations, displaying the output image on glasses allows a flexibility of use by the user in nearly any situation.
The image synthesis apparatus 200 may comprise a semantic segmentation neural network 220, object detection neural network 230, and/or an image completion neural network 240 that are trained for many different specific tasks. For example, the image synthesis apparatus 200 with a camera and a glasses display may be configured to capture a stream of input images 212 depicting a sport event scene such as 100c, detect the ball of the sport event within each input image 212, synthesize output images 248 with an enlarged version of the ball, and continuously display the output images 248 with the enlarged ball on the glasses, particularly as a video stream. This may enable the user to follow the ball and the entire sport event more easily. This may be an even greater help in other sports associated with smaller objects, such as ice hockey, baseball, or tennis. While watching the sport event, the user may be able to follow the event more easily, since any small object may then be easier to follow. The image synthesis apparatus 200 may be trained to enhance other objects of the sport event scene 100c, such as the players, referee, or boundary lines of the field of play.
The user may also benefit from the image synthesis apparatus 200 when viewing a football match next to its field of play. For example, in a stadium, the user may be located far away from the field, where it may be difficult to follow the ball. Instead of receiving a television signal as the input image 212, the input image 212 may be in the form of a camera image capturing the sport event scene 100c as seen from the perspective within the stadium. The same benefits that the user experiences with the television signal as the input image 212 can also be experienced with a camera image as the input image 212.
The semantic segmentation network 220, object detection network 230, and image completion network 240 may be trained to provide an AR experience corresponding to any setting, which may be a sport event, as well as a city for a tourist or a natural landscape for a hiker, among others. The user may specify a trained setting from among the total number of settings that the networks 220; 230; 240 have been trained to use the image synthesis apparatus 200 to synthesize and display output images 248. Different groups of pre-defined object types and/or typical objects may be used to train the networks 220; 230; 240 under different trained settings. The image synthesis apparatus 200 may comprise various modes, each mode corresponding to a respective trained setting, and the image synthesis apparatus 200 may switch to a different mode based on an input by a user.
A synthetic image may also be received as the input image 212 by the image synthesis apparatus 200, which may then synthesize the output image 248 with certain enhanced or removed objects within the received synthetic image in the same manner as when the input image 212 depicts a real-life scene. In other words, the image synthesis apparatus may be configured to receive any collection of input images 212 that depicts visual information, whether captured directly by a camera in proximity to the user or captured elsewhere, or whether real or synthetic, and then to synthesize corresponding output images 248 that enhances or removes certain objects based on a training of the semantic segmentation network 220, the object detection network 230, and the image completion network 240, and to display the corresponding output images 248 on a display 250, particularly on a transparent member, such as glasses.
The examples and embodiments of the image synthesis apparatus 200 demonstrate a flexibility in configuration, comprising a semantic segmentation network 220, an object detection network 230, and an image completion network 240 that are able to work together to receive input images 212 and generate corresponding output images 248. The image synthesis apparatus 200 may comprise one interface 210 configured to receive input images 212, a first and second interface 210; 214 configured to receive input images 212 of the same scene 100 or different scenes, or it may comprise any number of interfaces offering information of any number of scenes, wherein the first interface 210 is configured to receive the input image 212 and the further interfaces are configured to receive additional information 216, enabling the image synthesis apparatus 200 to synthesize an output image 248 with certain enhanced or removed objects and display it for the user based on a specific training, setting, or context. As such, the image synthesis apparatus 200 may be trained for any specific task to provide an AR experience that supports a user in safety or participation in activities of daily life.
Fig- 9 summarizes the proposed concept by illustrating a flowchart of a method 1000 for image synthesis based on the present disclosure. Method 1000 includes a step SI of receiving an input image 212 of a scene 100 comprising objects of different object types. The scene 100 may correspond to a vehicle setting, such as scene 100a or 100b, or it may correspond to a football match, such as scene 100c, among others. The scene 100 may comprise various objects and object types related to the context of the scene 100. The method also includes a step S2 of mapping, using a semantic segmentation network 220, the input image 212 to a segmentation map 222 of the scene 100, wherein each pixel of the segmentation map 222 is assigned to one of the different object types. The method 1000 also includes a step S3 of detecting, using an object detection network 230, one or more objects in the scene 100, wherein each detected object has associated therewith a respective object location and object type. Such an object detection network 230 may generate a labeled input image 232. The method also includes a step S4 of modifying the segmentation map 222 of the scene 100 based on the one or more detected objects to generate a modified segmentation map 244. The method also includes a step S5 to synthesize an output image 248 based on the modified segmentation map 244.
Fig. 10 outlines a procedure 1000 wherein the image synthesis apparatus 200 includes the second interface 214. SI to S4 and S7 describe processing acts of the image synthesis apparatus 200 including the (first) interface 210 that have previously been outlined in Figures 1 to 6. S5 and S6 describe processing acts of the image synthesis apparatus 200 additionally including the second interface 214. In SI, the first interface 210 is connected to a camera, or alternatively a radar or LiDAR sensor, and may receive one or more input images 212 depicting the scene 100, The scene 100 may comprise objects of different object types. The one or more input images 212 may be in the form of video frames that may collectively be processed to generate one or more output images 248 or augmented frames. The augmented frames may form a video stream to be continuously viewed as an AR experience. In S2, the input images 212 may be provided as input to the semantic segmentation network 220 and the object detection network 230. In S3, the semantic segmentation network 220 and object detection network 230 may detect certain objects and object types to be enhanced or removed in S4 by the image completion network 240.
In an embodiment of the present disclosure, the semantic segmentation network 220, the object detection network 230, and/or the image completion network 240 may be a neural network configured to detect the objects and object types and to remove them or merge them in enhanced form into the segmentation map 222, as previously described. If such objects or object types are found, the image completion network 240 merges or removes the objects or object types into the segmentation map 222 to obtain a modified segmentation map 244. For example, in Fig. 6, S4 results in the enlarged pedestrians 3-ls and 3-2s in the synthesized output image 248a. If such objects or object types are not found, the procedure skips S4 and proceeds to S5.
The image synthesis apparatus 200 may further comprise the second interface 214 configured to receive additional information 216 beyond the input image 212 of the scene 100, the additional information 216 yielding one or more additional object locations of additional objects, wherein each additional object location is assigned one object type. The image completion network 240 may be configured to merge the one or more additional objects into the segmentation map 222 of the scene 100 based on the respective additional object locations and object types in order to obtain the modified segmentation map 244. In S5, a second environmental sensor and/or a communication network may capture a blindspot scene or any scene comprising additional information 216. The second interface 214 of the image synthesis apparatus 200 may receive the additional information 216 from the second environmental sensor and/or communication network.
The additional information 216 may be in the form of one or more input images, which may be sent to the object detection network 230 and/or the image completion network 240. In S6, one or more objects included with the additional information 216 may be merged into the segmentation map 222 of the scene 100 to be included in the modified segmentation map 244 and the synthesized output image 248. The image completion network 240 may be configured to merge the one or more additional objects into the segmentation map 222 on a condition that the object type of the respective additional object meets a pre-determined criterion of importance for the second interface 214 and to merge the one or more objects included with the additional information 216 in an enhanced form, as previously described. In S7, the synthesized output image 248 comprising information obtained from both the first interface 210 and the second interface 214 is displayed on a display 250.
Note that the present technology can also be configured as described below.
Example 1 is an image synthesis apparatus comprising a first interface configured to receive an input image of a scene comprising objects of different object types, a semantic segmentation network configured to map the input image to a segmentation map of the scene, wherein each pixel of the segmentation map is assigned to one of the different object types, an object detection network configured to map the input image of the scene to one or more object locations of detected objects, wherein each object location is assigned one object type, and an image completion network configured to merge the one or more detected objects into the segmentation map of the scene based on the respective object locations and object types to obtain a modified segmentation map and to synthesize an output image based on the modified segmentation map.
In Example 2, the image completion network of Example 1 is configured to merge the one or more detected objects into the segmentation map of the scene on the condition that the object type of the respective detected object meets a pre-determined criterion of importance.
In Example 3, the image completion network of any one of the previous Examples is configured to enhance the one or more detected objects and to merge the one or more enhanced objects into the segmentation map of the scene based on the respective object locations and object types.
In Example 4, the image completion network of any one of the previous Examples is configured to change a size of the one or more detected objects and to merge the one or more objects with the changed size into the segmentation map of the scene based on the respective object locations and object types.
In Example 5, the image completion network of Example 4 is configured to change the size of the one or more detected objects based on distance information associated with the respective object.
In Example 6, the image completion network of Examples 4 or 5 is configured to enlarge the one or more detected objects. In Example 7, the image completion network of any one of the previous Examples is configured to synthesize the output image with an increased contrast of the one or more detected objects compared to remaining regions of the output image.
In Example 8, the image completion network of any one of the previous Examples is configured to change the object type assigned to the one or more detected objects and to merge the one or more objects with the changed object type into the segmentation map of the scene based on the respective object locations.
In Example 9, the image synthesis apparatus of any one of the previous Examples further comprises a second interface configured to receive additional information beyond the input image of the scene, wherein the additional information yields one or more additional object locations of additional objects, wherein each additional object location is assigned one object type, wherein the image completion network is configured to merge the one or more additional objects into the segmentation map of the scene based on the respective additional object locations and object types in order to obtain the modified segmentation map.
In Example 10, the first interface of the image synthesis apparatus of Example 9 is coupled to a first environmental sensor and the second interface of Example 9 is coupled to a different second environmental sensor and/or a communication network.
In Example 11, the image completion network of Examples 9 or 10 is configured to merge the one or more additional objects into the segmentation map of the scene on a condition that the object type of the respective additional object meets a pre-determined criterion of importance.
In Example 12, the second interface of Examples 9 or 10 is configured to receive the additional information from at least one of a camera signal, a radar signal, a lidar signal, and a radio signal.
In Example 13, the input image of any one of the previous Examples comprises at least one of a camera image, a radar image, and a LiDAR image. In Example 14, the image synthesis apparatus of one of the previous Examples further comprises a display configured to display the synthesized output image to a user.
In Example 15, the display of Example 14 is configured to display the synthesized output image on a transparent member.
Example 16 is an image synthesis method, the method comprising receiving an input image of a scene comprising objects of different object types, mapping, using a semantic segmentation network, the input image to a segmentation map of the scene, wherein each pixel of the segmentation map is assigned to one of the different object types, detecting, using an object detection network, one or more objects in the scene, wherein each detected object has associated therewith a respective object location and object type, modifying the segmentation map of the scene based on the one or more detected objects, and synthesizing an output image based on the modified segmentation map.
Example 17 is a computer program having computer readable instructions for carrying out the method according to Example 16, when the computer program is executed on a programmable hardware device.
The aspects and features described in relation to a particular one of the previous examples may also be combined with one or more of the further examples to replace an identical or similar feature of that further example or to additionally introduce the features into the further example.
Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a processor, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, some one or more of the most important method steps may be executed by such an apparatus.
Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software. The implementation can be performed using a non- transitory storage medium such as a digital storage medium, for example a floppy disc, a DVD, a Blu-Ray, a CD, a ROM, a PROM, and EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may, for example, be stored on a machine readable carrier.
Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
In other words, an embodiment of the present invention is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
A further embodiment of the present invention is, therefore, a storage medium (or a data carrier, or a computer-readable medium) comprising, stored thereon, the computer program for performing one of the methods described herein when it is performed by a processor. The data carrier, the digital storage medium or the recorded medium are typically tangible and/or non-transitionary. A further embodiment of the present invention is an apparatus as described herein comprising a processor and the storage medium.
A further embodiment of the invention is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may, for example, be configured to be transferred via a data communication connection, for example, via the internet.
A further embodiment comprises a processing means, for example, a computer or a programmable logic device, configured to, or adapted to, perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
In some embodiments, a programmable logic device (for example, a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.
Embodiments may be based on using a machine-learning model or machine-learning algorithm. Machine learning may refer to algorithms and statistical models that computer systems may use to perform a specific task without using explicit instructions, instead relying on models and inference. For example, in machine-learning, instead of a rule-based transformation of data, a transformation of data may be used, that is inferred from an analysis of historical and/or training data. For example, the content of images may be analyzed using a machine-learning model or using a machine-learning algorithm. In order for the machinelearning model to analyze the content of an image, the machine-learning model may be trained using training images as input and training content information as output. By training the machine-learning model with a large number of training images and/or training sequences (e.g. words or sentences) and associated training content information (e.g. labels or annotations), the machine-learning model "learns" to recognize the content of the images, so the content of images that are not included in the training data can be recognized using the machine-learning model. The same principle may be used for other kinds of sensor data as well: By training a machine-learning model using training sensor data and a desired output, the machine-learning model "learns" a transformation between the sensor data and the output, which can be used to provide an output based on non-training sensor data provided to the machine-learning model. The provided data (e.g. sensor data, meta data and/or image data) may be preprocessed to obtain a feature vector, which is used as input to the machine-learning model.
Machine-learning models may be trained using training input data. The examples specified above use a training method called "supervised learning". In supervised learning, the machine-learning model is trained using a plurality of training samples, wherein each sample may comprise a plurality of input data values, and a plurality of desired output values, i.e. each training sample is associated with a desired output value. By specifying both training samples and desired output values, the machine-learning model "learns" which output value to provide based on an input sample that is similar to the samples provided during the training. Apart from supervised learning, semi-supervised learning may be used. In semi-supervised learning, some of the training samples lack a corresponding desired output value. Supervised learning may be based on a supervised learning algorithm (e.g. a classification algorithm, a regression algorithm or a similarity learning algorithm. Classification algorithms may be used when the outputs are restricted to a limited set of values (categorical variables), i.e. the input is classified to one of the limited set of values. Regression algorithms may be used when the outputs may have any numerical value (within a range). Similarity learning algorithms may be similar to both classification and regression algorithms but are based on learning from examples using a similarity function that measures how similar or related two objects are. Apart from supervised or semi-supervised learning, unsupervised learning may be used to train the machine-learning model. In unsupervised learning, (only) input data might be supplied and an unsupervised learning algorithm may be used to find structure in the input data (e.g. by grouping or clustering the input data, finding commonalities in the data). Clustering is the assignment of input data comprising a plurality of input values into subsets (clusters) so that input values within the same cluster are similar according to one or more (pre-defined) similarity criteria, while being dissimilar to input values that are included in other clusters.
Reinforcement learning is a third group of machine-learning algorithms. In other words, reinforcement learning may be used to train the machine-learning model. In reinforcement learning, one or more software actors (called "software agents") are trained to take actions in an environment. Based on the taken actions, a reward is calculated. Reinforcement learning is based on training the one or more software agents to choose the actions such, that the cumulative reward is increased, leading to software agents that become better at the task they are given (as evidenced by increasing rewards).
Furthermore, some techniques may be applied to some of the machine-learning algorithms. For example, feature learning may be used. In other words, the machine-learning model may at least partially be trained using feature learning, and/or the machine-learning algorithm may comprise a feature learning component. Feature learning algorithms, which may be called representation learning algorithms, may preserve the information in their input but also transform it in a way that makes it useful, often as a pre-processing step before performing classification or predictions. Feature learning may be based on principal components analysis or cluster analysis, for example.
In some examples, anomaly detection (i.e. outlier detection) may be used, which is aimed at providing an identification of input values that raise suspicions by differing significantly from the majority of input or training data. In other words, the machine-learning model may at least partially be trained using anomaly detection, and/or the machine-learning algorithm may comprise an anomaly detection component.
In some examples, the machine-learning algorithm may use a decision tree as a predictive model. In other words, the machine-learning model may be based on a decision tree. In a decision tree, observations about an item (e.g. a set of input values) may be represented by the branches of the decision tree, and an output value corresponding to the item may be represented by the leaves of the decision tree. Decision trees may support both discrete values and continuous values as output values. If discrete values are used, the decision tree may be denoted a classification tree, if continuous values are used, the decision tree may be denoted a regression tree.
Association rules are a further technique that may be used in machine-learning algorithms. In other words, the machine-learning model may be based on one or more association rules. Association rules are created by identifying relationships between variables in large amounts of data. The machine-learning algorithm may identify and/or utilize one or more relational rules that represent the knowledge that is derived from the data. The rules may e.g. be used to store, manipulate or apply the knowledge. Machine-learning algorithms are usually based on a machine-learning model. In other words, the term "machine-learning algorithm" may denote a set of instructions that may be used to create, train or use a machine-learning model. The term "machine-learning model" may denote a data structure and/or set of rules that represents the learned knowledge (e.g. based on the training performed by the machine-learning algorithm). In embodiments, the usage of a machine-learning algorithm may imply the usage of an underlying machine-learning model (or of a plurality of underlying machine-learning models). The usage of a machine-learning model may imply that the machine-learning model and/or the data structure/set of rules that is the machine-learning model is trained by a machine-learning algorithm.
For example, the machine-learning model may be an artificial neural network (ANN). ANNs are systems that are inspired by biological neural networks, such as can be found in a retina or a brain. ANNs comprise a plurality of interconnected nodes and a plurality of connections, so-called edges, between the nodes. There are usually three types of nodes, input nodes that receiving input values, hidden nodes that are (only) connected to other nodes, and output nodes that provide output values. Each node may represent an artificial neuron. Each edge may transmit information, from one node to another. The output of a node may be defined as a (non-linear) function of its inputs (e.g. of the sum of its inputs). The inputs of a node may be used in the function based on a "weight" of the edge or of the node that provides the input. The weight of nodes and/or of edges may be adjusted in the learning process. In other words, the training of an artificial neural network may comprise adjusting the weights of the nodes and/or edges of the artificial neural network, i.e. to achieve a desired output for a given input.
Alternatively, the machine-learning model may be a support vector machine, a random forest model or a gradient boosting model. Support vector machines (i.e. support vector networks) are supervised learning models with associated learning algorithms that may be used to analyze data (e.g. in classification or regression analysis). Support vector machines may be trained by providing an input with a plurality of training input values that belong to one of two categories. The support vector machine may be trained to assign a new input value to one of the two categories. Alternatively, the machine-learning model may be a Bayesian network, which is a probabilistic directed acyclic graphical model. A Bayesian network may represent a set of random variables and their conditional dependencies using a directed acyclic graph. Alternatively, the machine-learning model may be based on a genetic algorithm, which is a search algorithm and heuristic technique that mimics the process of natural selection. It is further understood that the disclosure of several steps, processes, operations or functions disclosed in the description or claims shall not be construed to imply that these operations are necessarily dependent on the order described, unless explicitly stated in the individual case or necessary for technical reasons. Therefore, the previous description does not limit the execution of several steps or functions to a certain order. Furthermore, in further examples, a single step, function, process or operation may include and/or be broken up into several substeps, -functions, -processes or -operations.
If some aspects have been described in relation to a device or system, these aspects should also be understood as a description of the corresponding method. For example, a block, device or functional aspect of the device or system may correspond to a feature, such as a method step, of the corresponding method. Accordingly, aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.
The following claims are hereby incorporated in the detailed description, wherein each claim may stand on its own as a separate example. It should also be noted that although in the claims a dependent claim refers to a particular combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of any other dependent or independent claim. Such combinations are hereby explicitly proposed, unless it is stated in the individual case that a particular combination is not intended. Furthermore, features of a claim should also be included for any other independent claim, even if that claim is not directly defined as dependent on that other independent claim.

Claims

Claims An image synthesis apparatus, comprising: a first interface configured to receive an input image of a scene comprising objects of different object types; a semantic segmentation network configured to map the input image to a segmentation map of the scene, wherein each pixel of the segmentation map is assigned to one of the different object types; an object detection network configured to map the input image of the scene to one or more object locations of detected objects, wherein each object location is assigned one object type; and an image completion network configured to merge the one or more detected objects into the segmentation map of the scene based on the respective object locations and object types to obtain a modified segmentation map and to synthesize an output image based on the modified segmentation map. The image synthesis apparatus of claim 1, wherein the image completion network is configured to merge the one or more detected objects into the segmentation map of the scene on the condition that the object type of the respective detected object meets a pre-determined criterion of importance. The image synthesis apparatus of claim 1, wherein the image completion network is configured to enhance the one or more detected objects and to merge the one or more enhanced objects into the segmentation map of the scene based on the respective object locations and object types. The image synthesis apparatus of claim 1, wherein the image completion network is configured to change a size of the one or more detected objects and to merge the one or more objects with the changed size into the segmentation map of the scene based on the respective object locations and object types. 5. The image synthesis apparatus of claim 4, wherein the image completion network is configured to change the size of the one or more detected objects based on distance information associated with the respective object.
6. The image synthesis apparatus of claim 4, wherein the image completion network is configured to enlarge the one or more detected objects.
7. The image synthesis apparatus of claim 1, wherein the image completion network is configured to synthesize the output image with an increased contrast of the one or more detected objects compared to remaining regions of the output image.
8. The image synthesis apparatus of claim 1, wherein the image completion network is configured to change the object type assigned to the one or more detected objects and to merge the one or more objects with the changed object type into the segmentation map of the scene based on the respective object locations.
9. The image synthesis apparatus of claim 1, further comprising a second interface configured to receive additional information beyond the input image of the scene, the additional information yielding one or more additional object locations of additional objects, wherein each additional object location is assigned one object type, wherein the image completion network is configured to merge the one or more additional objects into the segmentation map of the scene based on the respective additional object locations and object types in order to obtain the modified segmentation map.
10. The image synthesis apparatus of claim 9, wherein the first interface is coupled to a first environmental sensor and wherein the second interface is coupled to a different second environmental sensor and/or a communication network.
11. The image synthesis apparatus of claim 9, wherein the image completion network is configured to merge the one or more additional objects into the segmentation map of the scene on a condition that the object type of the respective additional object meets a pre-determined criterion of importance.
12. The image synthesis apparatus of claim 9, wherein the second interface is configured to receive the additional information from at least one of a camera signal, a radar signal, a lidar signal, and a radio signal.
13. The image synthesis apparatus of claim 1, wherein the input image comprises at least one of a camera image, a radar image, and a LiDAR image.
14. The image synthesis apparatus of claim 1, further comprising a display configured to display the synthesized output image to a user.
15. The image synthesis apparatus of claim 14, wherein the display is configured to display the synthesized output image on a transparent member.
16. An image synthesis method, comprising: receiving an input image of a scene comprising objects of different object types; mapping, using a semantic segmentation network, the input image to a segmentation map of the scene, wherein each pixel of the segmentation map is assigned to one of the different object types; detecting, using an object detection network, one or more objects in the scene, wherein each detected object has associated therewith a respective object location and object type; modifying the segmentation map of the scene based on the one or more detected objects; and synthesizing an output image based on the modified segmentation map. A computer program having computer readable instructions for carrying out the method according to claim 16, when the computer program is executed on a programmable hardware device.
EP23820825.0A 2022-12-07 2023-12-05 Image synthesis apparatus and method Pending EP4631027A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP22211978 2022-12-07
PCT/EP2023/084334 WO2024121143A1 (en) 2022-12-07 2023-12-05 Image synthesis apparatus and method

Publications (1)

Publication Number Publication Date
EP4631027A1 true EP4631027A1 (en) 2025-10-15

Family

ID=84462977

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23820825.0A Pending EP4631027A1 (en) 2022-12-07 2023-12-05 Image synthesis apparatus and method

Country Status (2)

Country Link
EP (1) EP4631027A1 (en)
WO (1) WO2024121143A1 (en)

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115311422B (en) * 2018-06-27 2025-12-19 奈安蒂克公司 Multiple integration model for device positioning
CN112258618B (en) * 2020-11-04 2021-05-14 中国科学院空天信息创新研究院 Semantic mapping and positioning method based on fusion of prior laser point cloud and depth map

Also Published As

Publication number Publication date
WO2024121143A1 (en) 2024-06-13

Similar Documents

Publication Publication Date Title
US12236692B2 (en) Driver attention detection method
US11840261B2 (en) Ground truth based metrics for evaluation of machine learning based models for predicting attributes of traffic entities for navigating autonomous vehicles
CN113591872B (en) A data processing system, object detection method and device
Rasouli et al. Are they going to cross? a benchmark dataset and baseline for pedestrian crosswalk behavior
Fang et al. An automatic road sign recognition system based on a computational model of human recognition processing
JP2023525585A (en) Turn Recognition Machine Learning for Traffic Behavior Prediction
Dewangan et al. Towards the design of vision-based intelligent vehicle system: methodologies and challenges
Alvarez et al. Road geometry classification by adaptive shape models
Mukhopadhyay et al. A hybrid lane detection model for wild road conditions
CN120198884A (en) Traffic sign target detection method, electronic device and medium
CN116311252B (en) A collision warning method and system based on visual road environment diagram
Dewi et al. Image enhancement method utilizing YOLO models to recognize road markings at night
Ng et al. Real-time detection of objects on roads for autonomous vehicles using deep learning
Liu et al. Saliency difference based objective evaluation method for a superimposed screen of the HUD with various background
EP4631027A1 (en) Image synthesis apparatus and method
KR102705894B1 (en) System and method for controlling the output of advertisements in a vehicle
KR20240162019A (en) Method, system, and non-transitory computer-readable recording medium for providing advertisement contents
CN112818858A (en) Rainy day traffic video saliency detection method based on double-channel visual mechanism
Asghar et al. Flow-guided motion prediction with semantics and dynamic occupancy grid maps
Thayalan et al. Multifocus object detector for vehicle tracking in smart cities using spatiotemporal attention map
Korakakis et al. A short survey on modern virtual environments that utilize AI and synthetic data
Al Mudawi et al. Multimodal image fusion for enhanced vehicle identification in intelligent transport
Sujoy et al. Depth-aware object detection and region filtering for autonomous vehicles: a monocular camera-based novel integration of MiDaS and YOLO for complex road scenarios with irregular traffic
Pai et al. Forward Collision Warning and Lane-mark Recognition Systems Based on Deep Learning.
CN113392677A (en) Target object detection method and device, storage medium and terminal

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250707

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

DAV Request for validation of the european patent (deleted)
DAX Request for extension of the european patent (deleted)