WO2025201801A1 - A device, computer program and method - Google Patents
A device, computer program and methodInfo
- Publication number
- WO2025201801A1 WO2025201801A1 PCT/EP2025/055726 EP2025055726W WO2025201801A1 WO 2025201801 A1 WO2025201801 A1 WO 2025201801A1 EP 2025055726 W EP2025055726 W EP 2025055726W WO 2025201801 A1 WO2025201801 A1 WO 2025201801A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- objects
- user
- target object
- predetermined setting
- filter screen
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/44008—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics in the video stream
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/764—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/20—Scenes; Scene-specific elements in augmented reality scenes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/41—Higher-level, semantic clustering, classification or understanding of video scenes, e.g. detection, labelling or Markovian modelling of sport events or news items
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/41—Structure of client; Structure of client peripherals
- H04N21/422—Input-only peripherals, i.e. input devices connected to specially adapted client devices, e.g. global positioning system [GPS]
- H04N21/4223—Cameras
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/442—Monitoring of processes or resources, e.g. detecting the failure of a recording device, monitoring the downstream bandwidth, the number of times a movie has been viewed, the storage space available from the internal hard disk
- H04N21/44213—Monitoring of end-user related data
- H04N21/44218—Detecting physical presence or behaviour of the user, e.g. using sensors to detect if the user is leaving the room or changes his face expression during a TV programme
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/45—Management operations performed by the client for facilitating the reception of or the interaction with the content or administrating data related to the end-user or to the client device itself, e.g. learning user preferences for recommending movies, resolving scheduling conflicts
- H04N21/454—Content or additional data filtering, e.g. blocking advertisements
- H04N21/4542—Blocking scenes or portions of the received content, e.g. censoring scenes
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/81—Monomedia components thereof
- H04N21/8126—Monomedia components thereof involving additional data, e.g. news, sports, stocks, weather forecasts
- H04N21/814—Monomedia components thereof involving additional data, e.g. news, sports, stocks, weather forecasts comprising emergency warnings
Definitions
- the present technique relates to a device, computer program and method.
- the device 200 generates a segmentation mask that draws a detailed outline around each of the identified objects 121 and uses the segmentation mask to separate out the objects from the video frame.
- the segmented images corresponding to the identified objects are then fed into a classification model contained in device 200 for sorting into various categories of objects.
- the classification model may be a zero-shot classification model, or an open vocabulary image classification model that has been trained on a large dataset of images and associated labels. The zero-shot classification model learns aligned vision-language representations of various images from the training and is capable of generalizing to new and unseen categories without further training data.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Theoretical Computer Science (AREA)
- Signal Processing (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Databases & Information Systems (AREA)
- General Health & Medical Sciences (AREA)
- General Physics & Mathematics (AREA)
- Evolutionary Computation (AREA)
- Software Systems (AREA)
- Computing Systems (AREA)
- Medical Informatics (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Artificial Intelligence (AREA)
- Social Psychology (AREA)
- Business, Economics & Management (AREA)
- Emergency Management (AREA)
- Computer Networks & Wireless Communication (AREA)
- Computational Linguistics (AREA)
- Controls And Circuits For Display Device (AREA)
Abstract
A device comprising processing circuitry configured to: receive video content; detect objects in the video content by applying a segmentation model; identify classes of the detected objects using a zero-shot classification model; and determine a target object to be filtered among the detected objects according to the identified class of the detected objects and a predetermined setting.
Description
A DEVICE, COMPUTER PROGRAM AND METHOD
BACKGROUND
Field of the Disclosure
The present technique relates to a device, computer program and method.
Description of the Related Art
The “background” description provided herein is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventors, to the extent it is described in the background section, as well as aspects of the description which may not otherwise qualify as prior art at the time of filing, are neither expressly or impliedly admitted as prior art against the present technique.
The content of a real-life scene is normally unpredictable, for example, people who are afraid of dogs may meet dog on a bus or on a street; children in a cinema may by accident pass by a hall in which thriller or sexual movie is playing; and people who are on diet may lose their discipline when a snack advertisement pops up while they are watching a television programme. Generally speaking, some people tend to suffer from stress and struggles because of uncontrollable elements that may appear in their daily life.
With naked eyes, people cannot select what they want to see, and are therefore passively exposed to all existing items in the environment. For some people who have fears of specific animals, who are addicted to food, who suffer from PTSD after wars, social phobia, or who will faint at the sight of blood, they are always under stress since they do not know when they will come across the objects that may trigger their uncontrollable negative emotions or even cause syncope. For these reasons, some people choose rather to stay at home for a very long time in order to live in a known and controllable environment.
It is an aim of the disclosure to at least address this issue.
SUMMARY
According to one aspect of the disclosure, there is provided a device comprising processing circuitry configured to: receive video content; detect objects in the video content by applying a segmentation model; identify classes of the detected objects using a zero-shot classification model; and determine a target object to be filtered among the detected objects according to the identified class of the detected objects and a predetermined setting.
The foregoing paragraphs have been provided by way of general introduction, and are not intended to limit the scope of the following claims. The described embodiments, together with further advantages,
will be best understood by reference to the following detailed description taken in conjunction with the accompanying drawings.
BRIEF DESCRIPTION OF THE DRAWINGS
A more complete appreciation of the disclosure and many of the attendant advantages thereof will be readily obtained as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings, wherein:
Figure 1 shows a schematic diagram illustrating a device for providing a vision fdter according to embodiments of the present disclosure;
Figure 2 shows a device contained in the display apparatus of Figure 1 according to embodiments of the present disclosure;
Figure 3 describes a data structure for the storage of classification result according to embodiments of the present disclosure;
Figure 4 shows the application of a vision filter in a pair of Augmented Reality (AR) glasses with respect to an object in a real-life scene according to embodiments of the present disclosure;
Figure 5 shows a schematic diagram illustrating a filter screen for providing a vision filter according to embodiments of the present disclosure;
Figure 6 shows the application of a vision filter in a filter screen with respect to an object displayed in a TV programme according to embodiments of the present disclosure;
Figure 7 shows an example scenario in which privacy and security control is provided by the filter screen for a surveillance camera system according to embodiments of the present disclosure;
Figure 8 shows an example scenario of providing alarm monitory by the filter screen for a traffic monitoring camera system according to embodiments of the present disclosure; and
Figure 9 shows a flow chart describing a process of providing a vision filter by a device according to embodiments of the present disclosure.
DESCRIPTION OF THE EMBODIMENTS
Referring now to the drawings, wherein like reference numerals designate identical or corresponding parts throughout the several views.
Numerous modifications and variations of the present disclosure are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims, the disclosure may be practiced otherwise than as specifically described herein.
The present disclosure provides a vision filter device which allows filtering or blocking unwanted contents or objects in the scenes from being viewed by the user in order to help users live in a relaxed and enjoyable state of mind. The filtering of unwanted objects is performed by firstly capturing images of the real-life scene by a camera. The captured images are then transmitted to the vision filter device, which further detects various objects in the captured images by applying a segmentation model, such as the Segment Anything Model (SAM) created by Meta Al ® to segment the image into objects and to identify the location of each object within each image of the video. The purpose of using SAM is that it can be carried out in the vision filter device and there is no requirement for the device to be connected to the internet. Although the vison filter device in embodiments uses SAM, the disclosure is not so limited and the use of any appropriate image segmentation algorithm is envisaged such as Mockups by Glorify ® or ruttl or the like is envisaged.
The detected segments for different objects are subsequently classified by a classification model. In some embodiments, the classification model is a zero-shot classification model and is able to stay offline during operation. The vision filter device next determines one or more target objects to be filtered by comparing the corresponding classes identified for each of the segments with a predetermined criteria, such as a predetermined setting configured by the user entering the target objects or a class containing the target objects. For example, the predetermined criteria may contain a group of prohibited classes, and an object is determined as a target object to be filtered if its identified class falls within the group of prohibited classes. In another example, the predetermined criteria may contain a group of permitted classes and any objects falling outside the group are considered as target objects to be filtered. In embodiments, the predetermined setting may be configured based on conditions such as parental control rules, posttraumatic stress disorder triggers, privacy protection, or a combination thereof. After the target objects are determined, the vision filter device displays a mask on a location of a transparent display screen that corresponds to the position of the target object in the real-life scene, such that when the user view the real-life scene through the transparent display screen, the mask will filter or block the target object from the user’s view. In some embodiments, the mask may be a blur distortion filter that produces a frosted glass effect to blur the target object. In some other embodiments, the mask may be a partly or fully opaque area. In some further embodiments, the mask may be an icon or an information dialogue based on the identified class of the target object.
There are a wide range of scenarios in which this vision filter can be applied, for instance, for parental control to avoid kids watching sexual, violent scenes form movies or TV shows, or for assisting the patient with Posttraumatic Stress Disorder (PTSD) to block the trigger items, e.g., guns and for protecting the privacy by blurring smartphone screen and number plates in the public space. In some embodiments,
the images are processed on an image processing chip, such as Sony IMX500, while the images are being captured. Therefore no extract image processing needs to be done after the images are captured and the data load would be reduced.
Figure 1 is a schematic diagram illustrating a vision filter device for providing a vision filter according to embodiments of the present disclosure. In these embodiments, the vision filter device is a pair of Augmented Reality (AR) glasses 100 containing a device 200 which is connected to a camera 110 in the AR glasses 100. The camera 110 is positioned on the bridge of the AR glasses 100 and captures images in front of the user of the AR glasses 100. In embodiments, the camera 110 captures a video stream of everything the user of the AR glasses sees. The device 200 receives the video stream of the real-world scene from the camera 110 and apply a segmentation model to the video frames in the video streams to identify the location of each object 121 in the video frames and generate a mask to separate the identified objects 121 from their surroundings. In embodiments, the device 200 generates a segmentation mask that draws a detailed outline around each of the identified objects 121 and uses the segmentation mask to separate out the objects from the video frame. The segmented images corresponding to the identified objects are then fed into a classification model contained in device 200 for sorting into various categories of objects. In some embodiments, the classification model may be a zero-shot classification model, or an open vocabulary image classification model that has been trained on a large dataset of images and associated labels. The zero-shot classification model learns aligned vision-language representations of various images from the training and is capable of generalizing to new and unseen categories without further training data. As such, the use of zero-shot classification model allows the user to query images or categories with free-form text descriptions of a target object or a target category of objects. Since the zero-shot classification model is pre-trained, it is able to stay offline during operation. In the scenario of Figure 1, device 200 may classify object 121 as lorry and hamburger by applying the zero-shot classification model.
The result of the classification is compared with a predetermined filtering criteria such as a filter list containing descriptors of various objects desired to be filtered. The predetermined filtering criteria maybe created according to the user’s needs. In the example scenario of Figure 1, the user having obesity and suffering from related health problems such as heart disease, diabetes, high blood pressure or high cholesterol may have his AR glasses 100 configured to filter objects falling within snack and fast food categories. Device 200 determines that the object 121 relates to hamburgers which in turn belong to the fast food category in the filter list. Object 121 is therefore determined as a target object to be filtered.
Device 200 may instruct a display device on the AR glasses 100 to display masks 104, 105 on the respective lenses 102, 103 to block the object 121 from the user’s view. Device 200 may also detect the location of the object in real 3D space, for example, by an infra-red sensor and compute the coordinates and size of the mask on the lenses 102, 103 that correspond to the position and size of the real-life object
as seen by the user. The shape of the masks 104, 105 may also be determined based on the shape of the segmentation mask obtained from the segmentation model.
According to embodiments of the present disclosure, the lenses 102, 103 of the display device on the AR glasses 100 may be transparent display screens and the masks 104, 105 are displayed by turning on the corresponding pixels on the display screens.
The masks 104, 105 may be a partly or fully opaque area on the display screens. In the case of displaying the mask as a partly opaque area, a blur distortion effect may be created to prevent the user from seeing details of the object 121 and hence recognizing the nature of the object. In the meantime, the blur distortion effect allows the user to be alerted of the presence of an “anonymous object” in the location.
In some embodiments, the display device may be a projector device disposed on the temple of the AR glasses 100 and the masks 104, 105 are generated by the projector device projecting images on the lenses 102, 103 at coordinates corresponding to the position of the object 121 such that the projected image can also block the object 121 from the user’s view.
Although the foregoing describes that the device 200 activates vision filters based on the classification of target objects, the disclosure is in no way limited to this. According to embodiments of the present disclosure, the filtering of a target object may be determined based on additional conditions apart from the classification of the target object. For example, the device 200 may comprise a real-time clock and the device 200 may be configured such that the vision filter for snack objects are only activated during certain time periods of the day, for example, after 6 p.m. to prevent the user from taking evening snacks. In some embodiments, the device 200 may comprise biosensors for monitoring the metabolites of the user and the device 200 may be configured such that the vision filters are activated based on the physiological signals and health status of the user. For example, the vision filter for sweet objects is only activated when the blood glucose level of the user is high.
In some further embodiments, the device 200 may be connected to a pedometer for tracking the daily steps of the user such that the vision filter for fast food objects are deactivated when a goal of daily step has been met. In some additional embodiments, the device 200 may comprise or connect to a location device such as a GPS sensor or IP geolocation device for identifying the location of the user. The device 200 may be configured such that the vision filters are activated based on the location information of the user. For example, the vision filters for fast food objects are only activated when the user is in a commercial district with cafes, restaurants or food stalls around.
In some embodiments, the device 200 may receive information of the additional conditions for determining the filtering of target objects from a separate device, such as a smart watch, mobile phone, or other wearable devices.
Figure 2 shows a device 200 contained in the display apparatus of Figure 1 according to embodiments of the present disclosure. The device 200 comprises a processor 210 which is embodied as circuitry and may be any solid state circuitry such as circuitry controlled by software or an application specific integrated circuit. The processor 210 is connected to the camera 110 and the display device. This connection may be over a wired or wireless connection. The processor 210 is also connected to storage 220. In embodiments, the storage 220 is solid-state storage, but is not limited and may be optically readable storage or the like. Moreover, the storage 220 may be located remote to the device 200. In embodiments, the storage 220 contains computer readable instructions which, when loaded onto the processor 210, configures the device 200 to perform a method or methods according to embodiments of the present disclosure.
In some embodiments, the device 200 may be integrated with, or connected to, additional devices such as biosensors, pedometers or location devices. Accordingly, the processor 210 of the device 200 identifies the classification of the objects in a captured video stream and instructs the display device to activate vision filters for a target object based on predetermined filtering criteria and conditions determined by the information received from the additional devices.
Figure 3 illustrates a data structure for the storage of the classification result which are, in embodiments, a database. The data structure is stored in the storage 220 of the device 200. The purpose of storing the classification result is for determining whether a vision filter shall be applied to the respective object of in the video stream received from the camera and for generating a corresponding mask in the virtual space of the display device. The object coordinates, object size, classification of object are stored in correspondence with a unique object identifier. The data structure may further comprise a flag for indicating whether a vision filter shall be applied to the relevant object based on the result of comparing with the predetermined filtering criteria.
Figure 4 illustrates the application of a vision filter in a pair of Augmented Reality (AR) glasses with respect to an object in a real-life scene according to embodiments of the present disclosure.
The personal setting of the vision filter device 400 may be configured for an individual user, so that the AR glasses can be tailored to help the user filter out unwilling-to-see objects. According to embodiments of the present disclosure, the vision filter device has a similar appearance as a normal glass, except that it has a compact camera 410 in the middle, and two eyeglasses are Augmented Reality (AR) glasses. However, the disclosure is not so limited and the vision filter device can also be smart glass, AR headset, head mount display (HMD), heads-up display (HUD) or the like.
The camera 410 is connected to, or equipped with, a smart image processing chip which supports DNN- based image processing onsite. The camera 410 captures the content of the real-life scene, the segmentation model SAM in the smart image processing chip then detects various objects 421 in the
captured video content. In some embodiments, the smart image processing chip further includes a classifier for identifying the classes of the detected objects 421. According to the personal setting, the vision filter device 400 can generate a mask or a blur layer on the lenses of the AR glasses to block unwanted objects, for example, dogs, guns, blood or hamburger. This vision filter device 400 can also be used as parent control to filter sexual or violent scene or contents in movies, TV programmes, video games so as to protect children from potentially harmful contents. In some embodiments, the vison filter device 400 may be configured to filter specific food, cigarette, alcohol or drugs to help the user overcome addiction and assist addiction recovery.
According to embodiments of the present disclosure, the vision filter device 400 may be configured specifically for a user in order to filter or mask the exact content chosen by the user. By using SAM or other similar segmentation models, a single deep neural network (DNN) can mask a particular type of objects and no change needs to be applied to the DNN itself. In some embodiments, the vision filter device 400 provides a user interface for the user to enter the content to be masked with a simple prompt (e.g. “naked person”, “dog”, “cigarette”), as models like SAM are designed to be prompted by text or similar means. However, the disclosure is not so limited and the user input may be performed by following instructions displayed on the AR screen and based on voice, gesture recognition or the like. In embodiments, the configuration of the vision filter device 400 can be performed on a separate device such as a mobile phone, tablet or computer through wired or wireless connection.
Figure 5 depicts a schematic diagram illustrating a filter screen 500 for providing a vision filter in accordance with embodiments of the present disclosure. The filter screen 500 contains the device 200 which is connected to a camera 510 in the filter screen 500. The camera 510 is positioned on the frame of the filter screen 500 facing a display screen for showing the movies, TV programmes or video games as viewed by the spectators. The camera 510 captures a video stream of the media contents shown on the display screen as would be seen by the spectators. In some embodiments, the filter screen 500 may comprise video interface circuitry (such as HDMI, USB-C, Wi-Fi, Ethernet or the like) for directly receiving the video signal that is also being played on the television.
Device 200 receives the video stream of the media contents, either from the camera 510 or interface circuitry, and applies a segmentation model to the video frames in the video streams to identify the location of each object in the video frames and generate a mask to separate the identified objects from their surroundings. In embodiments, device 200 generates a segmentation mask that draws a detailed outline around each of the identified objects and uses the segmentation mask to separate out the objects from the video frame. The segmented images corresponding to the identified objects are then fed into a classification model contained in device 200 for sorting into various categories of objects. In some embodiments, the classification model may be a zero-shot classification model, or an open vocabulary image classification model that has been trained on a large dataset of images and associated labels. As such, the use of zero-shot classification model allows the user to query images or categories with free-
form text descriptions of a target object or a target category of objects. Since the zero-shot classification model is pre-trained, device 200 is able to stay offline during operation.
The result of the classification is compared with a predetermined filtering criteria such as a filter list containing descriptors of various objects desired to be filtered. The predetermined filtering criteria may be created according to the spectator’s needs.
Device 200 may instruct a display device on the filter screen 500 to display a mask 505 on the screen to block the target object from the spectator’s view. Device 200 may also detect the distance between the television display and the filter screen 500, and the distance between the filter screen 500 and the spectators, for example, by an infra-red sensor, thereby computing the coordinates and size of the mask on the screen that correspond to the position and size of the object in the television display as seen by the spectators. The shape of the mask 505 may also be determined based on the shape of the segmentation mask obtained from the segmentation model.
According to embodiments of the present disclosure, the filter screen 500 may comprise a transparent display screen and the mask 505 is displayed by turning on the corresponding pixels on the transparent display screen.
The mask 505 may be a partly or fully opaque area on the transparent display screen. In the case of displaying the mask as a partly opaque area, a blur distortion effect may be created to prevent the spectator from seeing details of the object 121 and hence recognizing the nature of the object.
In some embodiments, the mask 505 may be formed by a projector device projecting images on the transparent display screen at coordinates corresponding to the position of the object 121 such that the projected images can also block the object 121 from the spectator’s view.
Although the foregoing describes that the device 200 activates vision filters based on the classification of target objects, the disclosure is in no way limited to this. According to embodiments of the present disclosure, the filtering of a target object may be determined based on additional conditions apart from the classification of the target object. For example, the device 200 may comprise a real-time clock and the device 200 may be configured such that the vision filter are deactivated during the time periods of children’s television programme.
Figure 6 illustrates an example scenario in which parental control is provided by the filter screen of Figure 5 according to embodiments of the present disclosure. In these embodiments, the vison filter device is a filter screen 500 for parental control. The filter screen 500 may be implemented as a standing augmented reality screen that can be placed between the television 603 and children 610 while they are watching the movie alone or together with adults 620. In such a way, the adults 620 are able to watch the full contents of the movie, while the children 610 are protected from potentially harmful contents.
According to some embodiments of the present disclosure, the configuration of the filter screen 500 can be performed through an application installed in the television set. In some embodiments, the application may be stored in an external electronic device, such as a USB dongle, connected to the television set. The parents may configure the parental control setting of the filter screen 500 by running the application on the television set and predefine the television programme categories that are considered not suitable to be seen by children according to their age range, for example, through configuration menus 601, 602. Once the configuration is completed, the parental control setting is transmitted to the filter screen 500 through wired or wireless connection. Thereafter the children can watch movies or television programmes through the filter screen 500. When the potentially harmful objects appear in the movies or television programmes, the filter screen will block, blur or replace the item with other objects as described above with reference to Figure 5.
Figure 7 illustrates an example scenario in which privacy and security control is provided by the filter screen for a surveillance camera system according to embodiments of the present disclosure. The filter screen may be attached on a monitor display for showing videos captured from one or more public monitoring cameras 710. The filter screen may include a video interface circuitry for receiving video signals from the public monitoring cameras 710. A segmentation model such as SAM is applied to the video frames in the video signals to generate segmented images corresponding to the objects in the video frames. The segmented images are then input to a classification model such as a zero-shot classification model for identifying the classification of the relevant objects. In particular, if a sensitive object, text or scene has been classified, such as a car number plate or a passenger’s face, the filter screen will apply vision filter and display a mask 702 to block or blur the sensitive object in order to protect privacy.
According to embodiments of the present disclosure, the filter screen may further recognize the number on the car number plate and cross check with an approved list. For example, the surveillance camera system is installed in a private car park that allows access only to an approved list 704 of vehicles. In the event that the filter screen recognizes a vehicle having a number not on the approved list 704, it will trigger an alarm 707 in the monitor room to alert human assistants that a vehicle has entered the car park without permission. In some embodiments, the car number plate 703 of the vehicle may not be masked in this case to allow enforcement. In some embodiments, the alarm signal is a video signal and/or an audio signal identifying the detected object to the user.
In another example scenario, if it is recognized that the number on the car number plate 705 is on a detection list 706, the filter screen may again trigger an alarm 707 to alert the security personnel, and meanwhile not to apply vision filter to the number plate 705. However, if the number is not on the detection list 706, the filter screen will apply vision filter to the number plate in order to protect the privacy of other vehicles.
Figure 8 illustrates an example scenario of providing alarm monitory by the filter screen for a traffic monitoring camera system according to embodiments of the present disclosure. The filter screen may be attached on a monitor display in a monitor room for showing videos captured from one or more traffic monitoring cameras 810. The filter screen may include video interface circuitry for receiving video signals from the traffic monitoring cameras 810. A segmentation model such as SAM is applied to the video frames in the video signals to generate segmented images corresponding to the objects in real-life traffic 801 in the video frames. The segmented images are then input to a classification model such as a zero-shot classification model for identifying the classification of the relevant objects. In particular, if an object related to a traffic accident is classified, such as a crashed car 802, the filter screen will trigger an audio alarm signal 803 in the monitor room to alert the traffic control officer. In some embodiments, an alarm signal may be triggered or generated based on the identified class of the detected object and the alarm signal may be a video signal and/or audio signal 804, 806 describing the classification of the accident to the traffic control officer. In some embodiments, since the SAM segmentation model can recognize an object and combine with a classifier, the information on the user’s smartphone or even information in the scenes may also be checked. For instance, if a dangerous item is captured by the camera, such as an object looks like a bomb appearing in the scene, the monitoring system will trigger an alarm and alert security officers to confirm.
Figure 9 shows a flow chart 900 describing a process of providing a vision filter by a device according to embodiments of the present disclosure. The process starts in step 905, where video content is received, for example from a camera or video interface circuitry. At step 910, the objects in the video content is detected by applying a segmentation model. As noted above, in embodiments, the segmentation model is SAM. At step 915, the classes of the detected objects are identified by using a zero-shot classification model. The processes then moves to step 920, a target object to be filtered is determined among the detected objects according to the identified class of the detected objects and a predetermined setting. In embodiments, the predetermined setting may be configured by the user entering the target objects or a class containing the target objects. In embodiments, the predetermined setting may be configured based on parental control rules, posttraumatic stress disorder triggers, or privacy protection. Once the target object is determined, a mask layer may be applied to block the target object. In some embodiments, an alarm signal may be triggered or generated based on the identified class of the detected object and the alarm signal may be a video signal and/or audio signal identifying the detected object to the user
In so far as embodiments of the present disclosure have been described as being implemented, at least in part, by software -controlled data processing apparatus, it will be appreciated that a non-transitory machine-readable medium carrying such software, such as an optical disk, a magnetic disk, semiconductor memory or the like, is also considered to represent an embodiment of the present disclosure.
It will be appreciated that the above description for clarity has described embodiments with reference to different functional units, circuitry and/or processors. However, it will be apparent that any suitable distribution of functionality between different functional units, circuitry and/or processors may be used without detracting from the embodiments.
Described embodiments may be implemented in any suitable form including hardware, software, firmware or any combination of these. Described embodiments may optionally be implemented at least partly as computer software running on one or more data processors and/or digital signal processors. The elements and components of any embodiment may be physically, functionally and logically implemented in any suitable way. Indeed the functionality may be implemented in a single unit, in a plurality of units or as part of other functional units. As such, the disclosed embodiments may be implemented in a single unit or may be physically and functionally distributed between different units, circuitry and/or processors.
Although the present disclosure has been described in connection with some embodiments, it is not intended to be limited to the specific form set forth herein. Additionally, although a feature may appear to be described in connection with particular embodiments, one skilled in the art would recognize that various features of the described embodiments may be combined in any manner suitable to implement the technique.
Embodiments of the present technique can generally described by the following numbered clauses:
1. A device comprising processing circuitry configured to: receive video content; detect objects in the video content by applying a segmentation model; identify classes of the detected objects using a zero-shot classification model; and determine a target object to be filtered among the detected objects according to the identified class of the detected objects and a predetermined setting.
2. The device according to clause 1, wherein the predetermined setting is configured by the user entering the target objects or a class containing the target objects.
3. The device according to any preceding clause, wherein the predetermined setting is configured based on parental control rules.
4. The device according to any preceding clause, wherein the predetermined setting is configured based on posttraumatic stress disorder triggers.
5. The device according to any preceding clause, wherein the predetermined setting is configured based on privacy protection.
6. The device according to any preceding clause, wherein the processing circuitry is further configured to apply a mask layer to block the target object.
7. The device according to any preceding clause, wherein the processing circuitry is further configured to generate a video signal with the target object masked or blurred.
8. The device according to any preceding clause, wherein the processing circuitry is further configured to trigger an alarm signal based on the identified class of the detected object.
9. The device according to clause 8, wherein the alarm signal is a video signal identifying the detected object to the user.
10. The device according to any one of clauses 8 to 9, wherein the alarm signal is an audio signal identifying the detected object to the user.
11. Augmented reality glasses comprising the device according to any one of clauses 1 to 10; and an image capturing device for capturing the video content from a real-world scene, wherein the target object is a real-world object.
12. The augmented reality glasses according to clause 11, wherein the processing circuitry is further configured to: acquire location information of the target object in the real-world scene; determine a viewing angle of the augmented reality glasses with respect to the target object; and display, based on the location information and the viewing angle, a mask layer on the augmented reality glasses to block the target object from being viewed by the user.
13. A vision filter screen for filtering images displayed on a display device, the filter screen comprising the device according to any one of clauses 1 to 10; wherein the video content is the content displayed by the display device.
14. The vision filter screen according to clause 13, further comprising interface circuitry for receiving video signals carrying the video content.
15. The vision filter screen according to any one of clauses 13 to 14, further comprising an image capturing device for capturing the content displayed by the display device.
16. The vision filter screen according to any one of clauses 13 to 15, further comprising sensor circuitry for determining the location and viewing angle of the user, wherein the processing circuitry is further configured to: display, based on the location and the viewing angle of the user, a mask layer on the filter screen to block the object from being viewed by the user.
17. A method performed in a device, the method comprising: receiving video content; detecting objects in the video content by applying a segmentation model; identifying classes of the detected objects using a zero-shot classification model; and determining a target object to be filtered among the detected objects according to the identified class of the detected objects and a predetermined setting.
18. The method according to clause 17, wherein the predetermined setting is configured by the user entering the target objects or a class containing the target objects.
19. The method according to any one of clauses 17 to 18, wherein the predetermined setting is configured based on parental control rules.
20. The method according to any one of clauses 17 to 19, wherein the predetermined setting is configured based on posttraumatic stress disorder triggers.
21. The method according to any one of clauses 17 to 20, wherein the predetermined setting is configured based on privacy protection.
22. The method according to any one of clauses 17 to 21, further comprising applying a mask layer to block the target object.
23. The method according to any one of clauses 17 to 22, further comprising generating a video signal with the target object masked or blurred.
24. The method according to any one of clauses 17 to 23, further comprising triggering an alarm signal based on the identified class of the detected object.
25. The method according to clause 24, wherein the alarm signal is a video signal identifying the detected object to the user.
26. The method according to any one of clauses 24 to 25, wherein the alarm signal is an audio signal identifying the detected object to the user.
27. A computer program product comprising computer readable instructions which, when loaded onto a computer, configures the computer to perform a method according to any one of clauses 17 to 26.
Claims
1. A device comprising processing circuitry configured to: receive video content; detect objects in the video content by applying a segmentation model; identify classes of the detected objects using a zero-shot classification model; and determine a target object to be filtered among the detected objects according to the identified class of the detected objects and a predetermined setting.
2. The device according to claim 1, wherein the predetermined setting is configured by the user entering the target objects or a class containing the target objects.
3. The device according to claim 1, wherein the predetermined setting is configured based on parental control rules.
4. The device according to claim 1, wherein the predetermined setting is configured based on posttraumatic stress disorder triggers.
5. The device according to claim 1, wherein the predetermined setting is configured based on privacy protection.
6. The device according to claim 1, wherein the processing circuitry is further configured to apply a mask layer to block the target object.
7. The device according to claim 1, wherein the processing circuitry is further configured to generate a video signal with the target object masked or blurred.
8. The device according to claim 1, wherein the processing circuitry is further configured to generate an alarm signal based on the identified class of the detected object.
9. The device according to claim 8, wherein the alarm signal is a video signal identifying the detected object to the user.
10. The device according to claim 8, wherein the alarm signal is an audio signal identifying the detected object to the user.
11. Augmented reality glasses comprising the device according to claim 1 ; and
an image capturing device for capturing the video content from a real-world scene, wherein the target object is a real-world object.
12. The augmented reality glasses according to claim 11, wherein the processing circuitry is further configured to: acquire location information of the target object in the real-world scene; determine a viewing angle of the augmented reality glasses with respect to the target object; and display, based on the location information and the viewing angle, a mask layer on the augmented reality glasses to block the target object from being viewed by the user.
13. A vision filter screen for filtering images displayed on a display device, the filter screen comprising the device according to claim 1 ; wherein the video content is the content displayed by the display device.
14. The vision filter screen according to claim 13, further comprising interface circuitry for receiving video signals carrying the video content.
15. The vision filter screen according to claim 13, further comprising an image capturing device for capturing the content displayed by the display device.
16. The vision filter screen according to claim 13, further comprising sensor circuitry for determining the location and viewing angle of the user, wherein the processing circuitry is further configured to: display, based on the location and the viewing angle of the user, a mask layer on the filter screen to block the object from being viewed by the user.
17. A method performed in a device, the method comprising: receiving video content; detecting objects in the video content by applying a segmentation model; identifying classes of the detected objects using a zero-shot classification model; and determining a target object to be filtered among the detected objects according to the identified class of the detected objects and a predetermined setting.
18. The method according to claim 17, wherein the predetermined setting is configured by the user entering the target objects or a class containing the target objects.
19. The method according to claim 17, wherein the predetermined setting is configured based on parental control rules.
20. The method according to claim 17, wherein the predetermined setting is configured based on posttraumatic stress disorder triggers.
21. The method according to claim 17, wherein the predetermined setting is configured based on privacy protection.
22. The method according to claim 17, further comprising applying a mask layer to block the target object.
23. The method according to claim 17, further comprising generating a video signal with the target object masked or blurred.
24. The method according to claim 17, further comprising triggering an alarm signal based on the identified class of the detected object.
25. The method according to claim 24, wherein the alarm signal is a video signal identifying the detected object to the user.
26. The method according to any one of claim 24, wherein the alarm signal is an audio signal identifying the detected object to the user.
27. A computer program product comprising computer readable instructions which, when loaded onto a computer, configures the computer to perform a method according to claim 17.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24165978.8 | 2024-03-25 | ||
| EP24165978 | 2024-03-25 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025201801A1 true WO2025201801A1 (en) | 2025-10-02 |
Family
ID=90473263
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/EP2025/055726 Pending WO2025201801A1 (en) | 2024-03-25 | 2025-03-03 | A device, computer program and method |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025201801A1 (en) |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20160300388A1 (en) * | 2015-04-10 | 2016-10-13 | Sony Computer Entertainment Inc. | Filtering And Parental Control Methods For Restricting Visual Activity On A Head Mounted Display |
| US20170289624A1 (en) * | 2016-04-01 | 2017-10-05 | Samsung Electrônica da Amazônia Ltda. | Multimodal and real-time method for filtering sensitive media |
| US20180268240A1 (en) * | 2017-03-20 | 2018-09-20 | Conduent Business Services, Llc | Video redaction method and system |
| US20180376205A1 (en) * | 2015-12-17 | 2018-12-27 | Thomson Licensing | Method and apparatus for remote parental control of content viewing in augmented reality settings |
-
2025
- 2025-03-03 WO PCT/EP2025/055726 patent/WO2025201801A1/en active Pending
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20160300388A1 (en) * | 2015-04-10 | 2016-10-13 | Sony Computer Entertainment Inc. | Filtering And Parental Control Methods For Restricting Visual Activity On A Head Mounted Display |
| US20180376205A1 (en) * | 2015-12-17 | 2018-12-27 | Thomson Licensing | Method and apparatus for remote parental control of content viewing in augmented reality settings |
| US20170289624A1 (en) * | 2016-04-01 | 2017-10-05 | Samsung Electrônica da Amazônia Ltda. | Multimodal and real-time method for filtering sensitive media |
| US20180268240A1 (en) * | 2017-03-20 | 2018-09-20 | Conduent Business Services, Llc | Video redaction method and system |
Non-Patent Citations (2)
| Title |
|---|
| KIRILLOV ALEXANDER ET AL: "Segment Anything", 2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 5 April 2023 (2023-04-05), pages 3992 - 4003, XP093256960, ISBN: 979-8-3503-0718-4, Retrieved from the Internet <URL:https://arxiv.org/pdf/2304.02643> [retrieved on 20230405], DOI: 10.1109/ICCV51070.2023.00371 * |
| XIAODAN HU ET AL: "Smart Dimming Sunglasses for Photophobia Using Spatial Light Modulator", ARXIV.ORG, CORNELL UNIVERSITY LIBRARY, 201 OLIN LIBRARY CORNELL UNIVERSITY ITHACA, NY 14853, 9 July 2023 (2023-07-09), XP091562673, Retrieved from the Internet <URL:arXiv.org> * |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP3692461B1 (en) | Removing personally identifiable data before transmission from a device | |
| US10685496B2 (en) | Saving augmented realities | |
| Grusin | Premediation | |
| Simons et al. | Evidence for preserved representations in change blindness | |
| US9858474B2 (en) | Object tracking and best shot detection system | |
| JP6550460B2 (en) | System and method for identifying eye signals, and continuous biometric authentication | |
| US9600715B2 (en) | Emotion detection system | |
| US10614564B2 (en) | Method to determine impaired ability to operate a motor vehicle | |
| EP3026904A1 (en) | System and method of contextual adjustment of video fidelity to protect privacy | |
| JP6402219B1 (en) | Crime prevention system, crime prevention method, and robot | |
| US20160328015A1 (en) | Methods and devices for detecting and responding to changes in eye conditions during presentation of video on electronic devices | |
| JP2005315802A (en) | User support device | |
| JP2007200298A (en) | Image processing device | |
| US20200380243A1 (en) | Face Quality of Captured Images | |
| US20190356939A1 (en) | Systems and Methods for Displaying Synchronized Additional Content on Qualifying Secondary Devices | |
| FR3058534A1 (en) | INDIVIDUAL VISUAL IMMERSION DEVICE FOR MOVING PERSON WITH OBSTACLE MANAGEMENT | |
| TW202318155A (en) | Systems and methods for performing behavior detection and behavioral intervention | |
| CN114596636A (en) | Abnormal behavior identification method, device, electronic device and readable storage medium | |
| CN104050785A (en) | Safety alert method based on virtualized boundary and face recognition technology | |
| Baldry et al. | From Embodied Abuse to Mass Disruption: Generative, Inter-Reality Threats in Social, Mixed-Reality Platforms | |
| CN110262663B (en) | Schedule generation method and related products based on eye tracking technology | |
| WO2025201801A1 (en) | A device, computer program and method | |
| US10740624B2 (en) | Method for monitoring consumption of content | |
| Topinka | Terrorism, governmentality and the simulated city: the Boston Marathon bombing and the search for suspect two | |
| GB2622625A (en) | An image sensor device, method and computer program |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25708802 Country of ref document: EP Kind code of ref document: A1 |