WO2023045239A1 - 行为识别方法、装置、设备、介质、芯片、产品及程序 - Google Patents
行为识别方法、装置、设备、介质、芯片、产品及程序 Download PDFInfo
- Publication number
- WO2023045239A1 WO2023045239A1 PCT/CN2022/077461 CN2022077461W WO2023045239A1 WO 2023045239 A1 WO2023045239 A1 WO 2023045239A1 CN 2022077461 W CN2022077461 W CN 2022077461W WO 2023045239 A1 WO2023045239 A1 WO 2023045239A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- information
- face
- human body
- behavior
- event
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T7/00—Image analysis
- G06T7/20—Analysis of motion
- G06T7/246—Analysis of motion using feature-based methods, e.g. the tracking of corners or segments
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/10—Image acquisition modality
- G06T2207/10016—Video; Image sequence
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T2207/00—Indexing scheme for image analysis or image enhancement
- G06T2207/30—Subject of image; Context of image processing
- G06T2207/30196—Human being; Person
- G06T2207/30201—Face
Definitions
- Embodiments of the present disclosure relate to but are not limited to the technical field of computer vision, and particularly relate to a behavior recognition method, device, device, medium, chip, product and program.
- Embodiments of the present disclosure provide a behavior recognition method, device, equipment, medium, chip, product and program.
- a behavior recognition method comprising: acquiring a sequence of image frames; tracking at least one object detected in the sequence of image frames to obtain tracking information of each object; based on each Object tracking information, performing behavior recognition on the at least one object to obtain an identification of a behavior event in the image frame sequence; detecting object information of at least one object participating in the behavior event in the image frame sequence; The identifier of the behavior event is associated with the object information of the at least one object to obtain a behavior recognition result.
- the detecting the object information of at least one object participating in the behavior event in the sequence of image frames includes: determining an event area where the behavior event exists in each image of the sequence of image frames ; determining the object information of the at least one object from the event area in each of the images in which the behavior event exists.
- the event area where the behavior event exists in each image of the image frame sequence can be determined, at least one object can be determined in the obtained event area, and the detection area for detecting the object is reduced, so that the event can be quickly obtained from the event area. At least one object is detected in the area, and the speed of obtaining object information of the at least one object is improved.
- the object information of the at least one object includes a face identifier and a human body identifier of the at least one object; associating the identifier of the behavior event with the object information of the at least one object obtains the behavior
- the identification result includes: associating the face identification and human body identification belonging to the same object among the face identification and body identification of the at least one object to obtain an association result; performing the identification of the behavior event with the association result associated to obtain the behavior recognition result.
- the associating the face identifiers and body identifiers belonging to the same object among the face identifiers and body identifiers of the at least one object, and obtaining the association result includes: in each image frame sequence In an image, determine the positional relationship between each human face and each human body participating in the behavior event; based on the positional relationship, identify the person belonging to the same object among the human face identification and human body identification of the at least one object
- the face identification is associated with the human body identification to obtain an association result.
- the face identifiers and human body identifiers belonging to the same object are associated, thereby not only providing an association scheme between human faces and human bodies, And can accurately determine the associated face and human body.
- the associating the face identifiers and body identifiers belonging to the same object among the face identifiers and body identifiers of the at least one object, and obtaining the association result includes: from the person of the at least one object In the face identification and the human body identification, the face identification and the human body identification belonging to the same object are obtained; the face identification and the human body identification belonging to the same object correspond to the same tag value; wherein; the face identification of different objects corresponds to different tag values, The human body identifiers of different objects correspond to different tag values; the face identifiers and human body identifiers corresponding to the same tag value are associated to obtain the association result.
- the face identifiers and body identifiers belonging to the same object correspond to the same tag value, and the face identifiers and body identifiers corresponding to the same tag value are associated to obtain the association As a result, it is possible to easily associate face identifiers and human body identifiers belonging to the same object.
- the object information includes: face feature information and/or face attribute information;
- the detecting the object information of at least one object participating in the behavior event in the image frame sequence includes: detecting all The face feature information of the at least one object participating in the behavior event in the image frame sequence; based on the face feature information of the at least one object, determine the face attribute information of the at least one object;
- the method also The method includes: determining the identity information of the at least one object based on the face feature information and/or the face attribute information of the at least one object.
- the identity information of at least one object is determined, so that the identity information of the participating behavior event can be determined through the face information of the object, and then the participating behavior event can be determined.
- the object of the event is traced afterwards.
- the object information includes: human body feature information and/or human body attribute information
- the detecting the object information of at least one object participating in the behavior event in the image frame sequence includes: detecting the image The human body characteristic information of the at least one object participating in the behavior event in the frame sequence; based on the human body characteristic information of the at least one object, determining the human body attribute information of the at least one object; the method also includes: based on the The human body characteristic information and/or the human body attribute information of at least one object determine the identity information of the at least one object.
- the identity information of at least one object is determined, so that the identity information of the participating behavior event can be determined through the object's human body information, and then the object participating in the behavior event can be identified. Perform post-mortem traceability.
- the method further includes: when the behavior event is a preset event, outputting warning information; the warning information includes at least one of the following shown in the sequence of image frames: The event area of the behavior event, the spatio-temporal information of the behavior event, the face frame of the object participating in the behavior event, the body frame of the object participating in the behavior event, the face image in the face frame, the human body The image of the human body in the frame, the attribute information of the face and/or human body of the object participating in the behavior event, and the attribute information of the family members associated with the object participating in the behavior event.
- the alarm information can be output, so that the staff can determine that there is a preset event, and then can process the preset event in time.
- a behavior recognition device which includes: an acquisition part configured to acquire a sequence of image frames; a tracking part configured to track at least one object detected in the sequence of image frames, and obtain each The tracking information of the object; the identification part is configured to perform behavior recognition on the at least one object based on the tracking information of each object, and obtain the identification of the behavior event in the image frame sequence; the detection part is configured to detect the Object information of at least one object participating in the behavior event in the image frame sequence; an associating part configured to associate the identifier of the behavior event with the object information of the at least one object to obtain a behavior recognition result.
- the detection part is further configured to: determine the event area where the behavior event exists in each image of the image frame sequence; determine the event area where the behavior event exists in each image In the event area, object information of the at least one object is determined.
- the object information of the at least one object includes a face identification and a body identification of the at least one object
- the associating part is further configured to: associate the face identifier and the body identifier belonging to the same object among the face identifiers and body identifiers of the at least one object to obtain an association result; combine the identifier of the behavior event with the The above association results are associated to obtain the behavior recognition results.
- the association part is further configured to: in each image of the sequence of image frames, determine the positional relationship between each human face and each human body participating in the behavior event; based on the The positional relationship is associating the face identifiers and human body identifiers belonging to the same object among the face identifiers and body identifiers of the at least one object to obtain an association result.
- the associating part is further configured to: acquire the face ID and body ID belonging to the same object from the face ID and body ID of the at least one object; Corresponding to the same label value as the human body identification; wherein; the face identifications of different objects correspond to different label values, and the human body identifications of different objects correspond to different label values; Association to obtain the association result.
- the object information includes: face feature information and/or face attribute information
- the detection part is further configured to: detect the face feature information of the at least one object participating in the behavior event in the image frame sequence; determine the at least one face feature information based on the face feature information of the at least one object The face attribute information of the object;
- the behavior recognition device further includes a determining part; the determining part is configured to determine the identity information of the at least one object based on the facial feature information and/or the facial attribute information of the at least one object.
- the object information includes: human body feature information and/or human body attribute information,
- the detection part is further configured to: detect the human body feature information of the at least one object participating in the behavior event in the image frame sequence; determine the at least one object's body feature information based on the human body feature information of the at least one object Personal attribute information;
- the behavior recognition device further includes a determining part; the determining part is configured to determine the identity information of the at least one object based on the human body characteristic information and/or the human body attribute information of the at least one object.
- the behavior recognition device further includes an output part; the output part is configured to output warning information when the behavior event is a preset event;
- the warning information includes at least one of the following shown in the sequence of image frames:
- the event area where the behavior event occurs the spatio-temporal information of the behavior event, the face frame of the object participating in the behavior event, the body frame of the object participating in the behavior event, the face image in the face frame, the The human body image in the human body frame, the attribute information of the face and/or human body of the object participating in the behavior event, and the attribute information of the family members associated with the object participating in the behavior event.
- an electronic device including: a memory and a processor, the memory stores a computer program that can run on the processor, and the processor implements the computer program described in the first aspect above when executing the computer program steps in the method described above.
- a computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method described in the first aspect above in the steps.
- a chip including: a processor, configured to call and run a computer program from a memory, so that a device installed with the chip executes the steps in the method described in the first aspect.
- a computer program product carries a program code, and the program code includes instructions that can be configured to execute the steps in the method as described in the first aspect.
- a computer program including computer-readable codes.
- a processor in the electronic device executes the method described in the first aspect. step.
- the identification of the behavior event in the image frame sequence is obtained by performing behavior recognition on at least one object, and then the identification of the behavior event is associated with the object information of at least one object participating in the behavior event, it is possible to The obtained behavior recognition results trace the source of the objects involved in the behavior event afterwards.
- FIG. 1 is a schematic diagram of an implementation flow of a behavior recognition method provided by an embodiment of the present disclosure
- FIG. 2 is a schematic diagram of an implementation flow of another behavior recognition method provided by an embodiment of the present disclosure
- FIG. 3 is a schematic diagram of an implementation flow of another behavior recognition method provided by an embodiment of the present disclosure.
- FIG. 4 is a schematic diagram of an implementation flow of another behavior recognition method provided by an embodiment of the present disclosure.
- FIG. 5 is a schematic diagram of an implementation flow of a behavior recognition method provided by another embodiment of the present disclosure.
- FIG. 6 is a schematic diagram of an implementation flow of a behavior recognition method provided by another embodiment of the present disclosure.
- FIG. 7 is a schematic diagram of the composition and structure of a behavior recognition device provided by an embodiment of the present disclosure.
- FIG. 8 is a schematic diagram of a hardware entity of an electronic device provided by an embodiment of the present disclosure.
- FIG. 9 is a schematic structural diagram of a chip according to an embodiment of the present disclosure.
- multiple means two or more
- multiple frames means two or more frames, unless otherwise specifically defined.
- the following method is provided to determine whether there is a behavioral event in the video: first obtain the grayscale difference value of adjacent frames, then obtain a binary image according to the comparison between the grayscale difference value and a certain threshold value, and finally obtain the binary image according to the two
- the area with a large change in the value map can be used to judge whether there is a behavioral event.
- the following method is provided to determine whether there is a behavior event in the video: train the target detector in advance, obtain the pedestrian coordinate frame in each frame of image in the video, according to the distance between the pedestrian coordinate frame in each frame of image Classify the groups of people in each frame of the image, and send the classified group of people to the abnormal behavior classifier model for processing to obtain the probability of abnormal behavior.
- the above two methods either determine whether there is a behavioral event based on the information in the adjacent image or the information of the current frame. Since the amount of information relied on to determine whether there is a behavioral event is small, the false detection rate is serious. .
- the relevant personnel need to go to the scene to deal with the behavioral event after confirming that the behavioral event has occurred, and there is a time lag. .
- An embodiment of the present disclosure provides a behavior recognition method, which uses deep learning technology to analyze pedestrians in a video scene by fully considering the temporal characteristics of the video, so as to analyze whether there is a behavior event in the video.
- the basic information of the personnel involved in the behavioral event can be retained and stored for later use by the staff.
- Behavioral incidents can be dangerous incidents, fight incidents, incidents of riding electric vehicles without helmets, speeding incidents, or incidents such as driving in violation of regulations.
- an example of a behavior event is a fight event.
- the present disclosure does not limit the specific content of the behavior event, and in other embodiments, the behavior event may also be other, which is not limited in the embodiment of the present disclosure.
- Figure 1 is a schematic diagram of the implementation flow of a behavior recognition method provided by an embodiment of the present disclosure. As shown in Figure 1, the method is applied to a behavior recognition device.
- the behavior recognition device can be a processor or a chip, processing Devices or chips can be used in electronic equipment. In some other implementation manners, the behavior recognition device may be an electronic device.
- the electronic device in the embodiment of the present disclosure may include at least one of the following: server, mobile phone (Mobile Phone), tablet computer (Pad), computer with wireless transceiver function, palmtop computer , desktop computers, personal digital assistants, portable media players, smart speakers, navigation devices, smart watches, smart glasses, smart necklaces and other wearable devices, pedometers, digital TVs, virtual reality (Virtual Reality, VR) terminal equipment , Augmented Reality (AR) terminal equipment, wireless terminals in Industrial Control, wireless terminals in Self Driving, wireless terminals in Remote Medical Surgery, smart grid ( Wireless terminals in Smart Grid, wireless terminals in Transportation Safety, wireless terminals in Smart City, wireless terminals in Smart Home, and vehicles and vehicle-mounted devices in the Internet of Vehicles system Or on-board modules, etc.
- the method includes:
- one or more cameras can shoot real-time videos respectively, and send the captured real-time videos to the behavior recognition device, so that the behavior recognition device can obtain the real-time videos sent by one camera or multiple cameras.
- the behavior recognition device may be provided with a camera, and the real scene may be captured by the camera on the behavior recognition device, so as to obtain a real-time video.
- the behavior recognition device may receive pre-stored videos sent by other devices, or obtain pre-stored videos from its own storage.
- the behavior recognition device After the behavior recognition device obtains the real-time video or the pre-stored video, it can intercept the real-time video or the pre-stored video every set duration, so that the behavior recognition device can obtain the set duration video every set duration, based on A video of a set duration defines a sequence of image frames.
- the set duration may be a fixed duration, or the set duration may be a variable duration based on the current time and/or people and/or traffic in the current shooting scene.
- the value range of the set duration may be between 1 second and 10 minutes, for example, the set duration may be 1 second, 3 seconds, 1 minute, or 10 minutes.
- the number of images included in different sequences of video frames is the same.
- the behavior recognition result of an image frame sequence is obtained at intervals, and then through an image frame sequence to determine whether a behavior event occurs in the image frame sequence, fully considering the image frame sequence
- the temporal information and spatial information of the sequence improve the reliability of the obtained behavior recognition results.
- the Video of the set duration may be determined as a sequence of image frames; in other implementations, in order to reduce the amount of calculation of the behavior recognition device, the Extract images of a set ratio or a set number of images from a video of a fixed length, and determine the images of a set ratio or a set number of images as an image frame sequence.
- the behavior recognition device obtains the real-time video sent by multiple cameras
- the real-time videos sent by the multiple cameras can be processed in parallel; real-time video processing.
- S102 Track at least one object detected in the sequence of image frames to obtain tracking information of each object.
- the objects in the at least one object may be objects of the same attribute, for example, the at least one object is all persons.
- the objects in the at least one object may be objects of different attributes, for example, the at least one object may include at least one of people, animals, vehicles, and the like.
- An implementation manner of the above S102 may include: sequentially detecting objects in each image in the sequence of image frames to obtain tracking information of each object.
- the tracking information of each object may include at least one of the following: a coordinate frame of each object in the sequence of image frames, a screenshot of the coordinate frame of each object, and the like.
- the object may include a human face and/or a human body. There may be multiple coordinate frames for each object, and the number of coordinate frames for each object may be the same as the number of occurrences of each detected object in the sequence of image frames.
- the deep neural network can be obtained first, and the tracking information of each object can be input into the deep neural network, and the behavior of the at least one object can be recognized through the deep neural network, so as to obtain whether there is behavior in the sequence of image frames.
- the detection result of the event In the case that the detection result indicates that a behavior event exists in the sequence of image frames, the identity of the behavior event can be determined.
- the deep neural network may be a human body pose estimation network, through which it is determined whether there is a fight event.
- the deep neural network can analyze each image in the image frame sequence, and obtain at least one of whether there is a fighting event in each image, the number of people participating in the fighting event, the intensity of the fighting event, etc., And based on the ratio between the images with fighting events in the image frame sequence and all the images in the image frame sequence, the number of people participating in the fighting events in each image with fighting events, and the intensity of fighting events in each image with fighting events At least one of the degrees, determine whether there is a behavior event in the sequence of image frames.
- the detection result of whether there is a behavioral event in the image frame sequence can be obtained through the deep neural network, which can make the determined detection result of whether there is a behavioral event accurate.
- an image of each object may be extracted from a sequence of image frames, and then the object is determined based on the extracted image of the human body At least one of the number of times the human arm is raised, the range of the arm swing, the number of times the leg is raised, the range of the leg is raised, the age information of the face, the expression information of the face, the wound information of the face, etc., to determine whether each object participates in a behavior event.
- An identification of a behavioral event may correspond to a sequence of image frames.
- the identification of the behavior event may be determined based on at least one of the shooting time stamp of the image frame sequence, the real shooting location information of the image frame sequence, the number information of the camera, and the like. For example, different image frame sequences have different behavior event identifiers, and there is a one-to-one correspondence between the image frame sequences and the behavior event identifiers.
- the at least one object participating in the behavior event may include: at least one human face and/or at least one human body participating in the behavior event.
- an object frame of at least one object by detecting at least one object participating in the behavior event in the sequence of image frames, an object frame of at least one object (the object frame of the object includes a coordinate frame of the object frame and/or a screenshot of the coordinate frame of the object frame),
- the object frame of at least one object may include at least one of the following: a face frame of at least one human face, a screenshot of a face frame of at least one human face, a body frame of at least one human body, and a screenshot of a human body frame of at least one human body.
- the object frame in the embodiment of the present disclosure may be the coordinate information of the object frame
- the face frame may be the coordinate information of the face frame
- the human body frame may be the coordinate information of the human body frame.
- the screenshot of the object frame in the embodiment of the present disclosure may be an image corresponding to the object frame
- the screenshot of the face frame may be the image corresponding to the face frame
- the screenshot of the human body frame may be the image corresponding to the human body frame.
- the at least one object participating in the behavior event is: at least one object participating in the behavior event determined from the sequence of image frames.
- the object information of each object in the at least one object may include at least one of the following: the position information of each object and/or the object frame of each object in each image in the sequence of image frames, the feature information of each object, each The attribute information of an object, the identification information of each object, and the object tag value of each object.
- an object may include a human face and/or a human body, at least one object may include at least one human face and/or at least one human body, and each object may include each human face and/or each human body.
- behavior recognition results may also be stored.
- the behavior recognition result may be sent to the storage device, so that the storage device stores the behavior recognition result.
- the storage device may be a device independent of the behavior recognition device, for example, the storage device may be a distributed storage device.
- storing the behavior recognition result may include: storing the behavior recognition result in the behavior recognition device.
- the field information to be retrieved may include at least one of the following: period information to be retrieved, address information to be retrieved, face information to be retrieved, human body information to be retrieved, and name information to be retrieved.
- the behavior recognition result may be used to output at least one of the following when the acquired field information to be retrieved corresponds to the identifier of the behavior event: identifier of the behavior event, object information of at least one object.
- the object identifiers of the same object in different images in the sequence of image frames are the same. In other embodiments, the object identifiers of the same object in different images in the image frame sequence are different.
- the identifier of a behavior event can be associated with the object information of at least one object to obtain an association information.
- the association information can represent the association relationship between the identifier of the behavior event and the identifier of each object, so that based on the association relationship and the identifier of the behavior event, Each object participating in a behavioral event in a sequence of image frames can be easily determined.
- association information may be included in the associations (associations1) of the event.
- associations1 ⁇ le,lf i ⁇ , ⁇ le,lp j ⁇ ,... ⁇ .
- the associated information includes: ⁇ le,lf i ⁇ and ⁇ le,lp j ⁇ and so on.
- lf i represents the i-th human face participating in the behavior event in the image frame sequence
- lp j represents the j-th human body participating in the behavior event in the image frame sequence.
- the identification of the behavior event in the image frame sequence is obtained by performing behavior recognition on at least one object, and then the identification of the behavior event is associated with the object information of at least one object participating in the behavior event, it is possible to The obtained behavior recognition results trace the source of the objects involved in the behavior event afterwards.
- FIG. 2 is a schematic flowchart of another method for storing events provided by an embodiment of the present disclosure. As shown in FIG. 2 , the method is applied to a device for storing events, and the method includes:
- S202 Track at least one object detected in the sequence of image frames to obtain tracking information of each object.
- S205 Determine the object information of the at least one object from the event area where the behavior event exists in each image.
- the tracking information of each object is input into the deep neural network, and the deep neural network can also be used to obtain: an event area where a behavior event exists in each image of the sequence of image frames. Then the behavior recognition device can determine at least one object from the event area in each image where the behavior event exists.
- the behavior recognition device may determine all objects in the event area where the behavior event exists in each image as at least one object. For example, when a human face and two human bodies are detected through the event area, the one human face and the two human bodies may be determined as at least one object. In some other implementation manners, in order to make the determined object participating in the behavior event accurate, the event area in each image where the behavior event exists may be identified to obtain at least one object.
- the event area where a behavior event exists in each image of the image frame sequence can be determined, at least one object can be determined in the obtained event area, and the detection area used to detect the object is reduced, thus At least one object can be quickly detected from the event area, and the speed of obtaining object information of the at least one object is improved.
- FIG. 3 is a schematic diagram of the implementation flow of another behavior recognition method provided by an embodiment of the present disclosure. As shown in FIG. 3 , the method is applied to a behavior recognition device, and the method includes:
- S302. Track at least one object detected in the sequence of image frames to obtain tracking information of each object.
- the object information of the at least one object includes a face identification and a human body identification of the at least one object.
- S305 can be implemented in the following manner: in each image of the image frame sequence, determine the positional relationship between each human face and each human body participating in the behavior event; based on the positional relationship and associating the face identifiers and body identifiers belonging to the same object among the face identifiers and body identifiers of the at least one object, to obtain an association result.
- the behavior recognition device can determine the face frame of each face participating in the behavior event, the position information in each image, and determine the position of each human body in each human body frame and each image participating in the behavior event Information, and then based on the two position information, determine the positional relationship between each face and each human body participating in the behavior event.
- the face identifiers and human body identifiers belonging to the same object are associated, thus not only providing a face and human body identification
- the association scheme can accurately determine the associated face and human body.
- the associating the face identifiers and body identifiers belonging to the same object among the face identifiers and body identifiers of the at least one object, and obtaining the association result may include: from the at least one object In the face identification and body identification, obtain the face identification and body identification belonging to the same object; the face identification and body identification belonging to the same object correspond to the same label value; where; the face identification of different objects corresponds to different label values , the human body identifiers of different objects correspond to different tag values; the face identifiers and human body identifiers corresponding to the same tag value are associated to obtain the association result.
- the object information of the object may include a tag value of a human face and/or a tag value of a human body.
- the tag values of the faces belonging to the same person are the same as the tag values of the human body, or, the tag values of the faces belonging to the same person have a mapping relationship with the tag values of the human body.
- the identifier of the behavior event, the identifier of at least one human face and the label value of the human face may be determined as the behavior recognition result. In some other implementation manners, the identifier of the behavior event, the identifier of at least one human body, and the tag value of each specified human body may be determined as the behavior recognition result. In some other implementations, the identifier of a behavior event, the identifier of at least one human face, the tag value of at least one human face, the identifier of at least one human body, and the tag value of at least one human body may be determined as the behavior recognition result.
- the face ID and body ID belonging to the same object correspond to the same tag value, and the face ID and body ID corresponding to the same tag value are associated. , to obtain the association result, so that the face identifiers and human body identifiers belonging to the same object can be easily associated.
- the above association information can be obtained by associating the identifier of the behavior event with the association result.
- the behavior recognition device can obtain the face information corresponding to each face ID in all face IDs, and determine whether there is a face tag value in the face information of the face, or, when determining the face If the face label value in the face information is not zero, it is determined that the face has an associated human body, so that the face identifier and the human body identifier belonging to the same object can be associated to obtain an association result.
- the behavior recognition device may determine whether the face frame corresponding to each face in all faces is included in another frame, and if yes, determine that there is an associated human body for the face, And the associated human body is a human body in another frame, so that the face identifier and the human body identifier belonging to the same object can be associated to obtain an association result; if not, it is determined that there is no associated human body for the face.
- a scenario of a human body that is not associated with a human face is that the human body is blocked, so that the human body corresponding to the human face cannot be recognized.
- association information includes ⁇ le, lf 1 ⁇ , ⁇ le, lf 2 ⁇ , ⁇ le, lf 3 ⁇
- each specified face identification of the associated human body is lf 1 and lf 3
- associate the face identifier lf 1 with the human body identifier lp 1 and get ⁇ lf 1 , lp 1 ⁇
- the identity lf i of is associated with the identity lp j of the human body.
- the association result may be associations2 or ⁇ lf i , lp j ⁇ .
- the behavior recognition device can obtain the human body information corresponding to each human body identifier in all human body identifiers, and determine whether there is a human body tag value in the human body information of the human body, or determine whether the human body tag value is present in the human body information of the determined human body. If the label value is not zero, it is determined that the human body has an associated face, so that the face identifier and the human body identifier belonging to the same object can be associated to obtain an association result.
- the behavior recognition device may determine that the human body frame corresponding to each human body in all human bodies includes another frame, and in the case of yes, determine that the human body has an associated human face, and the associated human face is The human face in another frame, so that the face identification and the human body identification belonging to the same object can be associated to obtain an association result; in the case of no, it is determined that the human body does not have an associated human face.
- a scenario where there is no human face associated with the human body is that the human face is blocked, so that the human face corresponding to the human body cannot be recognized.
- association information includes ⁇ le, lp 1 ⁇ , ⁇ le, lp 2 ⁇ , ⁇ le, lp 3 ⁇
- each designated human body with an associated face is identified as lp 1 and lp 2
- associate the human body identifier lp 1 with the face identifier lf 1 and get ⁇ lf 1 , lp 1 ⁇
- the association result may be associations3 or ⁇ lp j ,lf i ⁇ .
- the association result is obtained, and then the identification of the behavior event is associated with the association result to obtain the behavior recognition result, so that it can be obtained from the behavior recognition result Faces and human bodies involved in behavioral events make information retrieval more comprehensive.
- Fig. 4 is a schematic diagram of the implementation flow of another behavior recognition method provided by the embodiment of the present disclosure.
- the object information includes: face feature information and/or face attribute information;
- the method is applied to a behavior recognition device, and the method includes:
- S402. Track at least one object detected in the sequence of image frames to obtain tracking information of each object.
- S406 Determine identity information of the at least one object based on the face feature information and/or the face attribute information of the at least one object.
- the identity information of the at least one object when the identity information of at least one object is determined, the identity information of the at least one object can be output, so that relevant personnel can quickly determine the identity of the object participating in the behavior event based on the identity information of at least one object. identity.
- At least one of face position information, face feature information, face attribute information, and at least one face identifier is included in the object information of the object.
- the face position information may be: the face position information of at least one face in each image in the sequence of image frames.
- the face position information of the at least one face in each image may be position information of a face frame of the at least one face in each image.
- the position information of the face frame of the at least one face in each image may include: position information of the upper left corner and the lower right corner of the face frame of the at least one face in each image.
- the characteristic information of each human face in the at least one human face can be matched with one or more images in the face database to determine the belongingness of each human face in the at least one human face.
- the attribute information of the family member may include at least one of the family member's name, gender, relationship with the object participating in the behavior event, contact information, and the like.
- the attribute information of the face may include one of the following information: whether to wear glasses, whether to wear a mask, age, gender, attribute information of facial features, and the like.
- the identity information of at least one object is determined based on the face feature information and/or face attribute information of at least one object, so that the identity information of the participating behavior event can be determined through the face information of the object, In turn, it is possible to trace the source of the objects participating in the behavior event afterwards.
- Fig. 5 is a schematic diagram of the implementation flow of a behavior recognition method provided by another embodiment of the present disclosure.
- the object information includes: human body feature information and/or human body attribute information
- the method Applied to a behavior recognition device the method includes:
- S502. Track at least one object detected in the sequence of image frames to obtain tracking information of each object.
- At least one of human body position information, human body feature information, human body attribute information, and at least one human body identifier is included in the object information of the object.
- the human body position information may be: human body position information of at least one human body in each image in each image of the sequence of image frames.
- the human body position information of at least one human body in each image may be position information of a human body frame of at least one human body in each image.
- the position information of the body frame of the at least one human body in each image may include: the position information of the upper left corner and the lower right corner of the body frame of the at least one human body in each image.
- the characteristic information of each human body in the at least one human body can be matched with one or more images in the human body library, and the name of the person to whom each human body belongs to at least one human body has been determined, At least one of school, class, work unit, identity mark, attribute information of family members, etc.
- the attribute information of the human body may include one of the following: information such as height, body shape, weight, and wearing information.
- the identity information of at least one object is determined based on the human body feature information and/or human body attribute information of at least one object, so that the identity information of the participating behavior event can be determined through the object's body information, and then the identity information of the participating behavior event can be determined. Objects participating in behavior events are traced afterwards.
- FIG. 6 is a schematic diagram of the implementation flow of a behavior recognition method provided by another embodiment of the present disclosure. As shown in FIG. 6, the method is applied to a behavior recognition device, and the method includes:
- S602. Track at least one object detected in the sequence of image frames to obtain tracking information of each object.
- the preset event may be a preset dangerous event, for example, the preset event may include at least one of the following: a fight event, an event of riding an electric vehicle without a helmet, an event of speeding, and an event of driving in violation of regulations.
- outputting the alarm information may include: outputting the alarm information to a specified device, wherein the specified device includes at least one of the following: a display device, a terminal device of a staff member in the area where the behavior event occurs, or a family member of the object participating in the behavior event terminal equipment.
- the alarm information includes at least one of the following displayed in the image frame sequence: the event area where the behavior event occurs, the spatio-temporal information of the behavior event, the face frame of the object participating in the behavior event, the participating The human body frame of the behavior event object, the face image in the human face frame, the human body image in the human body frame, the attribute information of the human face and/or human body of the object participating in the behavior event, and the The object of the behavior event is associated with the attribute information of the family member.
- the alarm information includes at least one of the face frame of the object participating in the behavior event, the body frame of the object participating in the behavior event, and the event area where the behavior event occurs, so that the corresponding face frame, At least one of the human body frame and the event area of the behavior event enables the staff to easily know the occurrence of the behavior event through the displayed real-time video.
- alarm information may be output, so that the staff can determine that the preset event exists, and then process the preset event in a timely manner.
- the camera collects video streams, and sends the collected video streams to the server, so that the server executes the steps of the behavior recognition method provided by the embodiments of the present disclosure, and the server determines that a behavior event occurs in a certain image frame sequence
- an alarm may be output once.
- an alarm is output, and a plurality of detection results of key frames preset at intervals of a second duration are output.
- the behavior recognition method provided by the embodiments of the present disclosure can be realized through the following modules: video decoding module, target detection and tracking module, fight detection module, face and human body detection module, human face and human body matching module, feature attribute extraction module, event output module and an alarm display module. Wherein, these modules can be set on one device or set on different devices.
- the video decoding module is used to decode the received video stream.
- the target detection and tracking module is used to detect and track pedestrians in the video stream.
- the fight detection module is used to input the image frame sequence into the human body font estimation model (that is, the above-mentioned deep neural network) to obtain the estimation of behavioral events.
- the face and human body detection module is used to detect the face and human body of pedestrians in the fighting area.
- the human face and human body matching module is used to match the human face and human body detected by the human face and human body.
- the feature attribute extraction module is used to extract feature information and attribute information of human face and human body.
- the event output module is used to output the event occurrence area, the event, and the face and human body information participating in the event to the alarm display module.
- the alarm display module is used for alarm display.
- the target detection and tracking module is used to detect pedestrians appearing in the video, using a human body detection algorithm.
- an image detection algorithm based on deep learning is used, which has a higher detection rate than traditional algorithms, and can cope with human bodies in complex environments, and These individual objects are tracked.
- the fight detection module is used to input the detected human body into the human body pose estimation model, take the human body event sequence in the video as input, and output an estimation result of a fight event at regular intervals, fully considering the time and space information of the video, Improved reliability of result output.
- the human face and human body matching module is used to associate the detected human face and human body through the spatial geometric relationship, and associate the human face and human body belonging to the same object.
- the event output module is used to organize and output information such as facial and human features and attributes involved in fighting events to a designated location for downstream storage or later retrieval.
- the alarm display module is used to display the fighting area with a global wide area network (World Wide Web, Web) page.
- a global wide area network World Wide Web, Web
- the face and human body matching module is to find the corresponding relationship between the face set and the human body set of pedestrians appearing in the fighting area.
- each item in the face set has 0
- each item in the face set has 0
- the elements in the face set such as f i
- generate a unique mark denoted as lf i
- the elements in the human body set such as p j
- generate a unique mark denoted as lp j .
- the event output module is the overall encapsulation of the event.
- the fight detection module detects the occurrence of a fight event, it will generate an event E, and generate its unique mark for the event E, which is recorded as le.
- the embodiments of the present disclosure use deep learning to detect pedestrian trajectories, use image frame sequences to analyze and model the behavior of pedestrian models, can perform accurate behavior analysis on pedestrians in videos, and can output accurate behavior prediction results.
- the embodiment of the present disclosure outputs the entire fighting event, and completes the entire event closed loop. Start tracking from the appearance of participants, start to identify behaviors when fights occur, output fight events, and output basic information of people involved in fight behavior events, and finally store these information for use by staff.
- the embodiments of the present disclosure can apply the protection provided by the embodiments of the present disclosure in various places where cameras exist, including but not limited to traffic places, shopping malls, stations, schools, or squares.
- Embodiments of the present disclosure may also provide a behavior recognition method, which is applied to a behavior recognition device. After the behavior recognition result is obtained, the method may further include: obtaining field information to be retrieved; The identification of the target behavior event corresponding to the field information to be retrieved, and/or the object information of the object participating in the target behavior event. In some embodiments, an identifier of the target behavior event and/or object information of objects participating in the target behavior event may also be output.
- the behavior recognition device may include a display screen, through which the staff inputs the field information to be retrieved, so that the behavior recognition device obtains the field information to be retrieved.
- the worker inputs the field information to be retrieved through a terminal device different from the behavior recognition device, and the terminal device sends the field information to be retrieved to the behavior recognition device, so that the behavior recognition device can obtain the field information to be retrieved .
- the field information to be retrieved includes: time period information to be retrieved and address information to be retrieved.
- obtaining the field information to be retrieved may be implemented in the following manner: in response to the input time period information to be retrieved and address information to be retrieved, the time period information to be retrieved and the address information to be retrieved are obtained.
- the time period information may be continuous certain time period information or non-continuous at least two time period information.
- the staff can directly input the time period information to be retrieved and the address information to be retrieved on the display screen of the behavior recognition device or the display screen of the terminal device, so that the behavior recognition device can obtain the time period information to be retrieved and the address information to be retrieved.
- the address information may be Wenyi Road, or the intersection of Wenyi Road and Jianshe Road, for example.
- obtaining the field information to be retrieved may be achieved in the following manner: in response to input time period information to be retrieved, outputting at least one address information that matches the time period information to be retrieved; in response to at least one address information The trigger operation of the address information to be retrieved is obtained to obtain the address information to be retrieved.
- the information about the occurrence of behavior events within the time period information to be retrieved can pop up on the display screen.
- the staff can select the address information to be retrieved from the displayed at least one piece of address information.
- the address information to be retrieved in the embodiments of the present disclosure may be one or more address information.
- at least one piece of address information may also be related to the authority of the account logged in by the staff.
- the staff can first input the time period information to be retrieved, so that the staff can determine which addresses have behavior events based on the time period information to be retrieved, so that the staff can choose to work according to at least one address information displayed
- the address information to be retrieved that the personnel care about, and then check the object information of the event-participating object that exists in the address information to be retrieved.
- obtaining the field information to be retrieved can be achieved in the following manner: responding to input address information to be retrieved; outputting at least one period information matching the address information to be retrieved; responding to at least one period information In the trigger operation of the time period information to be retrieved, the time period information to be retrieved is obtained.
- the identification of the target behavior event acquired by the behavior recognition device is not only related to the period information to be retrieved that the staff is interested in, but also related to the address information to be retrieved that the staff is interested in, so that the staff can easily From the massive amount of information, the interested period of time to be retrieved and the object information of the object participating in the behavior event in the address information to be retrieved are retrieved.
- the action recognition result set may correspond to multiple image frame sequences.
- the behavior recognition result set may include behavior recognition results of at least one image frame sequence, and each image frame sequence behavior recognition result includes: an identification of a behavior event occurring in the image frame sequence, and object information of at least one object participating in the behavior event.
- behavior event identifiers may all represent fight event identifiers.
- the identifier of at least one behavioral event may correspond to at least one sequence of image frames one-to-one.
- the staff by obtaining the identification of the target behavior event corresponding to the field information to be retrieved, and/or the object information of the object participating in the target behavior event, the staff can easily obtain the fields to be retrieved that are of interest The behavior event or object information corresponding to the information.
- the object information of the object participating in the target behavior event includes: face information corresponding to all face identifiers associated with the target behavior event identifier, and/or, all human body IDs associated with the target behavior event identifier Identify the corresponding human body information.
- the method may further include: obtaining the field information to be retrieved; obtaining the identification of the target behavior event corresponding to the field information to be retrieved from the behavior recognition result set; From the plurality of associated information stored in the recognition result set, obtain the identification of all faces and/or the identification of all human bodies associated with the identification of the target behavior event; from the object information in the behavior recognition result set, obtain the identification of all faces The face information corresponding to the identification of the target, and/or, the human body information corresponding to the identification of all human bodies; output the identification of the target behavior event, and/or, the object information of the object participating in the target behavior event.
- a plurality of association information may include the association information between the identifier of each event and the identifier of the object, for example, the plurality of association information may include the following content: ⁇ le1,lf i ⁇ , ⁇ le1,lp j ⁇ , ⁇ le2,lf i ⁇ , ⁇ le2, lp j ⁇ and so on.
- each association information includes the association information between the identifier of the behavior event and the identifier of the object, all face identifiers and/or all human body identifiers associated with the identifier of the target behavior event can be obtained based on multiple association information.
- the identification of all faces and/or the identification of all human bodies can be obtained, and the face information corresponding to the identification of all faces and the identification of all human bodies can be obtained Corresponding human body information.
- all the face identifiers and/or all human body identifiers associated with the identifier of the target behavior event are obtained through the multiple stored associated information, so that the corresponding face information and/or human body information corresponding to all human body identifiers.
- the object information of the object participating in the target behavior event includes: human body information corresponding to an identifier of a human body associated with a specified human face among all faces.
- the method may further include: obtaining the field information to be retrieved; obtaining the identification of the target behavior event corresponding to the field information to be retrieved from the behavior recognition result set; Obtain the identifiers of all faces associated with the identifiers of the target behavior event from the plurality of associated information stored in the recognition result set; obtain the face information corresponding to the identifiers of all faces from the object information in the behavior recognition result set ; From the association results stored in the behavior recognition result set, obtain the identification of the human body associated with the specified face in all faces; specify the face as a face with an associated human body; from the object information stored in the behavior recognition result set , acquiring human body information corresponding to the associated human body identifier; outputting the identifier of the target behavior event, and/or object information of objects participating in the
- the face information corresponding to all face identifiers from the object information stored in the behavior recognition result set may also be performed: from the association results stored in the behavior recognition result set, obtain Identify the face associated with the specified human body in all human bodies; specify the human body as a human body with an associated human face; obtain the face information corresponding to the associated face identification from the object information stored in the behavior recognition result set; output An identification of the target behavior event, and/or object information of objects participating in the target behavior event.
- the staff can not only know the face information participating in the behavior event, but also can Knowing the human body information matched by the face of the person involved in the behavior event enables the staff to have a more comprehensive understanding of the objects involved in the behavior event.
- the staff can not only know the human body information participating in the behavior event, but also obtain Know the face information matched by the human body participating in the behavior event, so that the staff can more comprehensively understand the objects participating in the behavior event.
- an embodiment of the present disclosure provides an apparatus for behavior recognition, which includes various modules included therein, and may be implemented by a processor in a terminal device.
- FIG. 7 is a schematic diagram of the composition and structure of a behavior recognition device provided by an embodiment of the present disclosure.
- the behavior recognition device 700 includes: an acquisition part 701 configured to acquire a sequence of image frames; At least one object detected in the image frame sequence is tracked to obtain the tracking information of each object; the identification part 703 is configured to perform behavior recognition on the at least one object based on the tracking information of each object to obtain the The identification of the behavior event in the image frame sequence; the detection part 704 is configured to detect the object information of at least one object participating in the behavior event in the image frame sequence; the association part 705 is configured to combine the identification and the behavior event The object information of the at least one object is associated to obtain a behavior recognition result.
- the detection part 704 is further configured to determine the event area where the behavior event exists in each image of the image frame sequence; from the event area where the behavior event exists in each image , determining object information of the at least one object.
- the object information of the at least one object includes the face identification and body identification of the at least one object; the associating part 705 is further configured to combine the face identification and body identification of the at least one object, The face identification and the human body identification belonging to the same object are associated to obtain an association result; the identification of the behavior event is associated with the association result to obtain the behavior recognition result.
- the associating part 705 is further configured to, in each image of the sequence of image frames, determine the positional relationship between each human face and each human body participating in the behavior event; based on the positional relationship and associating the face identifiers and body identifiers belonging to the same object among the face identifiers and body identifiers of the at least one object, to obtain an association result.
- the associating part 705 is further configured to obtain, from the face IDs and body IDs of the at least one object, the face IDs and body IDs belonging to the same object;
- the identification corresponds to the same label value; wherein; the face identifications of different objects correspond to different label values, and the human body identifications of different objects correspond to different label values; the human face identification and the human body identification corresponding to the same label value are associated, Obtain the association result.
- the object information includes: face feature information and/or face attribute information; the detection part 704 is further configured to detect the at least one object participating in the behavior event in the sequence of image frames Face feature information; based on the face feature information of the at least one object, determine the face attribute information of the at least one object; the behavior recognition device 700 also includes: a determining part 706 configured to be based on the at least one object.
- the facial feature information and/or the facial attribute information are used to determine the identity information of the at least one object.
- the object information includes: human body characteristic information and/or human body attribute information
- the detection part 704 is further configured to detect the human body characteristics of the at least one object participating in the behavior event in the sequence of image frames information; based on the human body characteristic information of the at least one object, determine the human body attribute information of the at least one object; the determining part 706 is also configured to be based on the human body characteristic information and/or the human body attribute of the at least one object information, to determine the identity information of the at least one object.
- the behavior recognition device 700 further includes: an output part 707 configured to output warning information when the behavior event is a preset event; the warning information includes the following in the image frame sequence At least one of: the event area where the behavior event occurs, the spatio-temporal information of the behavior event, the face frame of the object participating in the behavior event, the body frame of the object participating in the behavior event, the face frame of the The face image, the human body image in the human body frame, the attribute information of the face and/or human body of the object participating in the behavior event, and the attribute information of the family members associated with the object participating in the behavior event.
- the above-mentioned behavior recognition method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer storage medium.
- the essence of the technical solutions of the embodiments of the present disclosure or the part that contributes to the related technologies can be embodied in the form of software products, the computer software products are stored in a storage medium, and include several instructions to make A terminal device executes all or part of the methods in various embodiments of the present disclosure.
- FIG. 8 is a schematic diagram of a hardware entity of an electronic device provided by an embodiment of the present disclosure. As shown in FIG. A computer program running on 801, when the processor 801 executes the program, implements the steps in the method of any of the foregoing embodiments.
- the electronic device 800 may include the above-mentioned behavior recognition apparatus.
- the memory 802 stores computer programs that can run on the processor, the memory 802 is configured to store instructions and applications executable by the processor 801, and can also cache the pending or processed data of the processor 801 and each module in the electronic device 800 Data (for example, image data, audio data, voice communication data, and video communication data) can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).
- FLASH flash memory
- RAM random access memory
- the processor 801 executes the program, the steps of any one of the above-mentioned behavior recognition methods are realized.
- the processor 801 generally controls the overall operation of the electronic device 800 .
- An embodiment of the present disclosure provides a computer storage medium, where one or more programs are stored in the computer storage medium, and the one or more programs can be executed by one or more processors to implement the behavior recognition method in any of the above embodiments. step.
- FIG. 9 is a schematic structural diagram of a chip according to an embodiment of the present disclosure.
- the chip 900 shown in FIG. 9 includes a processor 910, configured to call and run a computer program from a memory, so that a device installed with the chip executes the method in any one of the above-mentioned embodiments.
- the chip 900 may further include a memory 920 .
- the processor 910 can invoke and run a computer program from the memory 920, so as to implement the methods in the embodiments of the present disclosure.
- the memory 920 may be an independent device independent of the processor 910 , or may be integrated in the processor 910 .
- the chip 900 may further include an input interface 930 .
- the processor 910 may control the input interface 930 to communicate with other devices or chips, for example, may obtain information or data sent by other devices or chips.
- the chip 900 may further include an output interface 940 .
- the processor 910 can control the output interface 940 to communicate with other devices or chips, for example, can output information or data to other devices or chips.
- the chip can be applied to the control network element or the execution network element in the embodiments of the present disclosure, and the chip can implement the corresponding processes implemented by the control network element or the execution network element in the various methods of the embodiments of the present disclosure , for the sake of brevity, it is not repeated here.
- the chips mentioned in the embodiments of the present disclosure may also be referred to as system-on-chip, system-on-chip, system-on-a-chip, or system-on-chip.
- Embodiments of the present disclosure may also provide a computer program product, the computer program product carries program code, and instructions included in the program code can be configured to execute the steps in any one of the above methods.
- An embodiment of the present disclosure may also provide a computer program, including computer readable codes.
- a processor in the electronic device executes any of the methods described above. step.
- the above-mentioned behavior recognition device, chip or processor may include the integration of any one or more of the following: application specific integrated circuit (Application Specific Integrated Circuit, ASIC), digital signal processor (Digital Signal Processor, DSP), digital signal processing device ( Digital Signal Processing Device, DSPD), Programmable Logic Device (Programmable Logic Device, PLD), Field Programmable Gate Array (Field Programmable Gate Array, FPGA), Central Processing Unit (Central Processing Unit, CPU), Graphics Processor (Graphics Processing Unit, GPU), embedded neural network processors (neural-network processing units, NPU), controller, microcontroller, microprocessor. Understandably, the electronic device that implements the above processor function may also be other, which is not specifically limited in this embodiment of the present disclosure.
- ASIC Application Specific Integrated Circuit
- DSP Digital Signal Processor
- DSPD Digital Signal Processing Device
- PLD Programmable Logic Device
- Field Programmable Gate Array Field Programmable Gate Array
- FPGA Field Programmable Gate Array
- CPU Central Processing Unit
- GPU Graphic
- the above-mentioned computer storage medium/memory can be read-only memory (Read Only Memory, ROM), programmable read-only memory (Programmable Read-Only Memory, PROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, EPROM), Electrically Erasable Programmable Read-Only Memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), Magnetic Random Access Memory (Ferromagnetic Random Access Memory, FRAM), Flash Memory (Flash Memory), Magnetic Surface Memory, CD-ROM, or CD-ROM (Compact Disc Read-Only Memory, CD-ROM) and other memories; it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc. .
- references throughout the specification to "one embodiment” or “an embodiment” or “an embodiment of the present disclosure” or “the foregoing embodiments” or “some implementations” or “some embodiments” mean the same as implementing A specific feature, structure, or characteristic related to an example is included in at least one embodiment of the present disclosure.
- appearances of "in one embodiment” or “in an embodiment” or “embodiments of the present disclosure” or “the foregoing embodiments” or “some implementations” or “some embodiments” throughout the specification do not necessarily mean Must refer to the same embodiment.
- the particular features, structures or characteristics may be combined in any suitable manner in one or more embodiments.
- sequence numbers of the above-mentioned processes do not mean the order of execution, and the execution order of the processes should be determined by their functions and internal logic, rather than by the embodiments of the present disclosure.
- the implementation process constitutes any limitation.
- the serial numbers of the above-mentioned embodiments of the present disclosure are for description only, and do not represent the advantages and disadvantages of the embodiments.
- the behavior recognition device executes any step in the embodiments of the present disclosure, and may be a processor of the behavior recognition device executes the step. Unless otherwise specified, the embodiments of the present disclosure do not limit the order in which the behavior recognition device executes the following steps. In addition, the methods for processing data in different embodiments may be the same method or different methods. It should also be noted that any step in the embodiments of the present disclosure can be independently executed by behavior recognition, that is, when the behavior recognition device executes any step in the above embodiments, it may not depend on the execution of other steps.
- the disclosed devices and methods may be implemented in other ways.
- the device embodiments described above are only illustrative.
- the division of the modules is only a logical function division.
- the mutual coupling, or direct coupling, or communication connection between the various components shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or modules may be in electrical, mechanical or other forms of.
- modules described above as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules; they may be located in one place or distributed to multiple network modules; Part or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
- each functional module in each embodiment of the present disclosure can be integrated into one processing module, or each module can be used as a single module, or two or more modules can be integrated into one module; the above-mentioned integration
- the modules can be implemented in the form of hardware, or in the form of hardware plus software function modules.
- the above-mentioned integrated modules of the present disclosure are realized in the form of software function modules and sold or used as independent products, they may also be stored in a computer storage medium.
- the computer software products are stored in a storage medium, and include several instructions to make
- a computer device which may be a personal computer, a server, or a network device, etc.
- the aforementioned storage medium includes various media capable of storing program codes such as removable storage devices, ROMs, magnetic disks or optical disks.
- the term "and" does not affect the order of the steps.
- the electronic device executes A and then executes B. It may be that the electronic device executes A first and then B, or the electronic device executes B first. Execute A again, or execute B while the electronic device executes A.
- Embodiments of the present disclosure provide a behavior recognition method, device, device, medium, chip, product, and program, wherein, by performing behavior recognition on at least one object, the behavior event identifier in the image frame sequence is obtained, and then the behavior event The identification is associated with the object information of at least one object participating in the behavior event, so the source of the object participating in the behavior event can be traced afterwards through the behavior recognition result obtained through the association.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Image Analysis (AREA)
Abstract
本公开实施例公开了一种行为识别方法、装置、设备、介质、芯片、产品及程序,其中,该行为识别方法包括:获取图像帧序列;对所述图像帧序列中检测到的至少一个对象进行跟踪,得到每个对象的跟踪信息;基于所述每个对象的跟踪信息,对所述至少一个对象进行行为识别,得到所述图像帧序列中行为事件的标识;检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息;将所述行为事件的标识和所述至少一个对象的对象信息关联,得到行为识别结果。
Description
相关申请的交叉引用
本专利申请要求2021年09月22日提交的中国专利申请号为202111109050.X,申请名称为“行为识别方法、装置、电子设备及计算机存储介质”的优先权,该公开的全文以引用的方式并入本公开中。
本公开实施例涉及但不限于计算机视觉技术领域,尤其涉及一种行为识别方法、装置、设备、介质、芯片、产品及程序。
随着人工智能(Artificial Intelligence,AI)算法在智慧城市应用里面的逐步推广,该算法的优势逐渐显示出来:较传统算法更高的准确率,对场景更好的适应性等。由于视频中存在巨大的信息量,因此如何有效地识别视频中的信息变得非常有必要。相关技术仅关注视频中对象的检测,没有关注到需要将检测到的结果中信息的进行关联,使得无法对检测到的对象进行事后溯源。
发明内容
本公开实施例提供一种行为识别方法、装置、设备、介质、芯片、产品及程序。
第一方面,提供一种行为识别方法,所述方法包括:获取图像帧序列;对所述图像帧序列中检测到的至少一个对象进行跟踪,得到每个对象的跟踪信息;基于所述每个对象的跟踪信息,对所述至少一个对象进行行为识别,得到所述图像帧序列中行为事件的标识;检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息;将所述行为事件的标识和所述至少一个对象的对象信息关联,得到行为识别结果。
在一些实施例中,所述检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息,包括:确定所述图像帧序列的每一图像中的存在所述行为事件的事件区域;从所述每一图像中的存在所述行为事件的事件区域中,确定所述至少一个对象的对象信息。
这样,由于能够确定图像帧序列的每一图像中的存在行为事件的事件区域,从而能够在得到的事件区域中确定至少一个对象,缩小了用于检测对象的检测区域,从而能够快速地从事件区域检测到至少一个对象,提高了得到至少一个对象的对象信息的速度。
在一些实施例中,所述至少一个对象的对象信息包括所述至少一个对象的人脸标识和人体标识;所述将所述行为事件的标识和所述至少一个对象的对象信息关联,得到行为识别结果,包括:将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果;将所述行为事件的标识和所述关联结果进行关联,得到所述行为识别结果。
这样,通过将属于同一对象的人脸标识和人体标识进行关联,得到关联结果,然后将行为事件的标识和关联结果进行关联得到行为识别结果,从而能够从行为识别结果中获取参与行为事件的人脸和人体,使得信息的检索更加全面。
在一些实施例中,所述将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果,包括:在所述图像帧序列的每一图像中,确定参与所述行为事件的每一人脸与每一人体之间的位置关系;基于所述位置关系,将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果。
这样,通过确定的参与行为事件的每一人脸与每一人体之间的位置关系,将属于同一对象的人脸标识和人体标识进行关联,从而不仅提供了一种人脸和人体的关联方案,且能够准确的确定关联的人脸和人体。
在一些实施例中,所述将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果,包括:从所述至少一个对象的人脸标识和人体标识中,获取属于同一对象的人脸标识和人体标识;向属于同一对象的人脸标识和人体标识对应相同的标签值;其中;不同对象的人脸标识对应不同的标签值,不同对象的人体标识对应不同的标签值;将所述相同的标签值对应的人脸标识和人体标识进行关联,得到所述关联结果。
这样,通过获取属于同一对象的人脸标识和人体标识,向属于同一对象的人脸标识和人体标识对应相同的标签值,将相同的标签值对应的人脸标识和人体标识进行关联,得到关联结果,从而能够容易地将属于同一对象的人脸标识和人体标识进行关联。
在一些实施例中,所述对象信息包括:人脸特征信息和/或人脸属性信息;所述检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息,包括:检测所述图像帧序列中参与所述行为事件的所述至少一个对象的人脸特征信息;基于所述至少一个对象的人脸特征信息,确定所述至少一个对象的人脸属性信息;所述方法还包括:基于所述至少一个对象的所述人脸特征信息和/或所述人脸属性信息,确定所述至少一个对象的身份信息。
这样,基于至少一个对象的人脸特征信息和/或人脸属性信息,确定至少一个对象的身份信息,从而能够通过对象的人脸信息对参与行为事件的身份信息进行确定,进而能够对参与行为事件的对象进行事后溯源。
在一些实施例中,所述对象信息包括:人体特征信息和/或人体属性信息,所述检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息,包括:检测所述图像帧序列中参与所述行为事件的所述至少一个对象的人体特征信息;基于所述至少一个对象的人体特征信息,确定所述至少一个对象的人体属性信息;所述方法还包括:基于所述至少一个对象的所述人体特征信息和/或所述人体属性信息,确定所述至少一个对象的身份信息。
这样,基于至少一个对象的人体特征信息和/或人体属性信息,确定至少一个对象的身份信息,从而能够通过对象的人体信息对参与行为事件的身份信息进行确定,进而能够对参与行为事件的对象进行事后溯源。
在一些实施例中,所述方法还包括:在所述行为事件为预设事件的情况下,输出告警信息;所述告警信息包括以下在所述图像帧序列中展示的至少之一:发生所述行为事件的事件区域、所述行为事件的时空信息、参与所述行为事件对象的人脸框、参与所述行为事件对象的人体框、所述人脸框中的人脸图像、所述人体框中的人体图像、参与所述行为事件的对象的人脸和/或人体的属性信息、参与所述行为事件的对象关联家属的属性信息。
这样,在行为事件为预设事件的情况下,可以输出告警信息,从而使得工作人员能够确定存在预设事件,进而能够及时地对预设事件进行处理。
第二方面,提供一种行为识别装置,所述装置包括:获取部分,配置为获取图像帧序列;跟踪部分,配置为对所述图像帧序列中检测到的至少一个对象进行跟踪,得到每个对象的跟踪信息;识别部分,配置为基于所述每个对象的跟踪信息,对所述至少一个对象进行行为识别,得到所述图像帧序列中行为事件的标识;检测部分,配置为检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息;关联部分,配置为将所述行为事件的标识和所述至少一个对象的对象信息关联,得到行为识别结果。
在一些实施例中,所述检测部分,还配置为:确定所述图像帧序列的每一图像中的存在所述行为事件的事件区域;从所述每一图像中的存在所述行为事件的事件区域中,确定所述至少一个对象的对象信息。
在一些实施例中,所述至少一个对象的对象信息包括所述至少一个对象的人脸标识和 人体标识;
所述关联部分,还配置为:将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果;将所述行为事件的标识和所述关联结果进行关联,得到所述行为识别结果。
在一些实施例中,所述关联部分,还配置为:在所述图像帧序列的每一图像中,确定参与所述行为事件的每一人脸与每一人体之间的位置关系;基于所述位置关系,将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果。
在一些实施例中,所述关联部分,还配置为:从所述至少一个对象的人脸标识和人体标识中,获取属于同一对象的人脸标识和人体标识;向属于同一对象的人脸标识和人体标识对应相同的标签值;其中;不同对象的人脸标识对应不同的标签值,不同对象的人体标识对应不同的标签值;将所述相同的标签值对应的人脸标识和人体标识进行关联,得到所述关联结果。
在一些实施例中,所述对象信息包括:人脸特征信息和/或人脸属性信息;
所述检测部分,还配置为:检测所述图像帧序列中参与所述行为事件的所述至少一个对象的人脸特征信息;基于所述至少一个对象的人脸特征信息,确定所述至少一个对象的人脸属性信息;
所述行为识别装置还包括确定部分;所述确定部分,配置为基于所述至少一个对象的所述人脸特征信息和/或所述人脸属性信息,确定所述至少一个对象的身份信息。
在一些实施例中,所述对象信息包括:人体特征信息和/或人体属性信息,
所述检测部分,还配置为:检测所述图像帧序列中参与所述行为事件的所述至少一个对象的人体特征信息;基于所述至少一个对象的人体特征信息,确定所述至少一个对象的人体属性信息;
所述行为识别装置还包括确定部分;所述确定部分,配置为基于所述至少一个对象的所述人体特征信息和/或所述人体属性信息,确定所述至少一个对象的身份信息。
在一些实施例中,所述行为识别装置还包括输出部分;所述输出部分,配置为在所述行为事件为预设事件的情况下,输出告警信息;
所述告警信息包括以下在所述图像帧序列中展示的至少之一:
发生所述行为事件的事件区域、所述行为事件的时空信息、参与所述行为事件对象的人脸框、参与所述行为事件对象的人体框、所述人脸框中的人脸图像、所述人体框中的人体图像、参与所述行为事件的对象的人脸和/或人体的属性信息、参与所述行为事件的对象关联家属的属性信息。
第三方面,提供一种电子设备,包括:存储器和处理器,所述存储器存储有可在所述处理器上运行的计算机程序,所述处理器执行所述计算机程序时实现上述第一方面所述方法中的步骤。
第四方面,提供一种计算机存储介质,所述计算机存储介质存储有一个或者多个程序,所述一个或者多个程序可被一个或者多个处理器执行,以实现上述第一方面所述方法中的步骤。
第五方面,提供一种芯片,包括:处理器,用于从存储器中调用并运行计算机程序,使得安装有所述芯片的设备执行如第一方面所述方法中的步骤。
第六方面,提供一种计算机程序产品,所述计算机程序产品承载有程序代码,所述程序代码包括的指令可配置为执行如第一方面所述方法中的步骤。
第七方面,提供一种计算机程序,包括计算机可读代码,在所述计算机可读代码在电子设备中运行的情况下,所述电子设备中的处理器执行如第一方面所述方法中的步骤。
在本公开实施例中,由于通过对至少一个对象进行行为识别,得到图像帧序列中行为 事件的标识,然后将行为事件的标识与参与行为事件的至少一个对象的对象信息关联,因此能够通过关联得到的行为识别结果,对参与行为事件的对象进行事后溯源。
为使本公开的上述目的、特征和优点能更明显易懂,下文特举较佳实施例,并配合所附附图,作详细说明如下。
为了更清楚地说明本公开实施例的技术方案,下面将对实施例中所需要使用的附图作简单地介绍,此处的附图被并入说明书中并构成本说明书中的一部分,这些附图示出了符合本公开的实施例,并与说明书一起用于说明本公开的技术方案。
图1为本公开实施例提供的一种行为识别方法的实现流程示意图;
图2为本公开实施例提供的另一种行为识别方法的实现流程示意图;
图3为本公开实施例提供的又一种行为识别方法的实现流程示意图;
图4为本公开实施例提供的再一种行为识别方法的实现流程示意图;
图5为本公开另一实施例提供的一种行为识别方法的实现流程示意图;
图6为本公开又一实施例提供的一种行为识别方法的实现流程示意图;
图7为本公开实施例提供的一种行为识别装置的组成结构示意图;
图8为本公开实施例提供的一种电子设备的硬件实体示意图;
图9为本公开实施例的芯片的示意性结构图。
下面将通过实施例并结合附图具体地对本公开的技术方案进行详细说明。下面这几个具体的实施例可以相互结合,对于相同或相似的概念或过程可能在某些实施例中不再赘述。
需要说明的是:在本公开实例中,“第一”、“第二”等是用于区别类似的对象,而不必用于描述特定的顺序或先后次序。
另外,本公开实施例所记载的技术方案之间,在不冲突的情况下,可以任意组合。在本公开的描述中,“多个”的含义是两个或两个以上,“多帧”的含义是两帧或两个以帧,除非另有明确具体的限定。
在一些实施方式中,提供了以下方式确定视频中是否存在行为事件:先获取相邻帧的灰度差值,然后根据灰度差值与某一阈值的比较来获取二值图,最后根据二值图中变化较大的区域来判断是否存在行为事件。
在另一些实施方式中,提供了以下方式确定视频中是否存在行为事件:事先训练好目标检测器,获取视频中每一帧图像中的行人坐标框,根据每一帧图像中行人坐标框之间的距离对每一帧图像中的人群进行分类,将分类好的人群送到异常行为分类器模型中进行处理,来获取异常行为概率。
然而,上述两种方式要么是根据相邻图像中的信息,要么是根据当前帧的信息,来确定是否存在行为事件,由于确定是否存在行为事件所依赖的信息量较少,导致误检率严重。
另外,相关人员在确定到发生行为事件之后,需要到现场处理行为事件,时间上存在滞后性,例如在相关人员到达现场后,很可能行为事件已经结束,导致无法对参与行为事件的对象进行溯源。
本公开实施例提供了一种行为识别方法,利用深度学习技术,充分考虑视频的时间特性来对视频场景中的行人进行分析,从而分析视频中是否存在行为事件。另外,能够保留参与行为事件的人员的基本信息并存储,来供工作人员后期使用。行为事件可以是危险事件,打架事件、骑电动车不戴头盔事件、行驶超速事件或者不按规定标示行驶等事件。在本公开实施例中,行为事件示例为打架事件。本公开并不限定行为事件的具体内容,在其它实施例中,行为事件还可以是其它,本公开实施例对此不作限定。
图1为本公开实施例提供的一种行为识别方法的实现流程示意图,如图1所示,该方 法应用于行为识别装置,在一些实施方式中,行为识别装置可以为处理器或芯片,处理器或芯片可以应用于电子设备中。在另一些实施方式中,行为识别装置可以为电子设备。本公开实施例中的电子设备、下述的其它设备、显示设备或者终端设备可以包括以下至少之一:服务器、手机(Mobile Phone)、平板电脑(Pad)、带无线收发功能的电脑、掌上电脑、台式计算机、个人数字助理、便捷式媒体播放器、智能音箱、导航装置、智能手表、智能眼镜、智能项链等可穿戴设备、计步器、数字TV、虚拟现实(Virtual Reality,VR)终端设备、增强现实(Augmented Reality,AR)终端设备、工业控制(Industrial Control)中的无线终端、无人驾驶(Self Driving)中的无线终端、远程手术(Remote Medical Surgery)中的无线终端、智能电网(Smart Grid)中的无线终端、运输安全(Transportation Safety)中的无线终端、智慧城市(Smart City)中的无线终端、智慧家庭(Smart Home)中的无线终端以及车联网系统中的车、车载设备或车载模块等。该方法包括:
S101、获取图像帧序列。
在一些实施方式中,一路或者多路摄像头可以分别拍摄实时视频,并将拍摄到的实时视频向行为识别装置发送,从而行为识别装置能够得到一路摄像头或多路摄像头发送的实时视频。在另一些实施方式中,行为识别装置上可以设置有摄像头,通过行为识别装置上的摄像头对真实场景进行拍摄,从而得到实时视频。在又一些实施方式中,行为识别装置可以接收其它设备发送预先存储的视频,或者从自身的存储器中获取预先存储的视频。
在行为识别装置得到实时视频或者预先存储的视频之后,可以每隔设定时长对实时视频或者预先存储的视频进行截取,从而行为识别装置可以每隔设定时长获取到设定时长的视频,基于设定时长的视频确定图像帧序列。其中,设定时长可以是固定时长,或者,设定时长可以是基于当前时间和/或当前拍摄场景中的人和/或流量而变化的时长。设定时长的取值范围可以在1秒钟至10分钟之间,例如,设定时长可以为1秒钟、3秒钟、1分钟或者10分钟等。在一些实施方式中,不同的视频帧序列中包括的图像的数量相同。
这样,通过每隔一段时间获取图像帧序列,从而每隔一段时间得到一个图像帧序列的行为识别结果,进而通过一个图像帧序列来确定该图像帧序列中是否发生行为事件,充分考虑了图像帧序列的时间信息与空间信息,提高了得到的行为识别结果的可靠性。
基于设定时长的视频确定图像帧序列,在一些实施方式中,可以是将设定时长的视频确定为图像帧序列;在另一些实施方式中,为了降低行为识别装置的计算量,可以从设定时长的视频中抽取设定比例的图像或者设定数量的图像,将设定比例的图像或者设定数量的图像确定为图像帧序列。
在行为识别装置得到多路摄像头发送的实时视频的情况下,在一些实施方式中,可以对多路摄像头发送的实时视频进行并行处理;在另一些实施方式中,可以对依次对每一路摄像头发送的实时视频进行处理。
S102、对所述图像帧序列中检测到的至少一个对象进行跟踪,得到每个对象的跟踪信息。
至少一个对象中的对象可以是相同属性的对象,例如,至少一个对象均是人。或者,至少一个对象中的对象可以是不同属性的对象,例如,至少一个对象可以包括人、动物、车辆中的至少一者等事物。
上述S102的一种实施方式可以包括:依次对图像帧序列中的每个图像中的对象进行检测,得到每个对象的跟踪信息。其中,每个对象的跟踪信息可以包括以下至少之一:图像帧序列中的每个对象的坐标框、每个对象的坐标框截图等。其中,对象可以是包括人脸和/或人体。每个对象的坐标框可以有多个,每个对象的坐标框的数量,可以与检测到的每个对象在图像帧序列中出现的次数相同。
S103、基于所述每个对象的跟踪信息,对所述至少一个对象进行行为识别,得到所述图像帧序列中行为事件的标识。
在一些实施方式中,可以先获取深度神经网络,将每个对象的跟踪信息输入至深度神经网络,通过深度神经网络可以对所述至少一个对象进行行为识别,从而得到图像帧序列中是否存在行为事件的检测结果。在检测结果表征图像帧序列中存在行为事件的情况下,可以确定行为事件的标识。例如,在行为事件为打架事件的情况下,深度神经网络可以是人体姿态估计网络,通过人体姿态估计网络确定是否存在打架事件。在一些实施方式中,深度神经网络可以对图像帧序列中每一图像进行分析,并得到每一图像中是否存在打架事件、参与打架事件的人数、打架事件的激烈程度等中的至少一者,并基于图像帧序列中存在打架事件的图像与图像帧序列中所有图像之间的比值、每一存在打架事件的图像中参与打架事件的人数,以及每一存在打架事件的图像中打架事件的激烈程度中的至少一者,确定图像帧序列是否存在行为事件。
通过这种方式,由于深度神经网络是通过训练得到的,从而通过深度神经网络得到:图像帧序列中是否存在行为事件的检测结果,能够使得确定的是否存在行为事件的检测结果准确。
以行为事件为打架事件为例,在另一些实施方式中,可以从图像帧序列中提取每一对象的图像(例如人体的图像或者人脸的图像),然后基于提取的人体的图像确定该对象人体胳膊抬起的次数、胳膊挥舞的幅度、腿部抬起的次数、腿部抬起的幅度、人脸的年龄信息、人脸的表情信息、人脸的伤口信息等中的至少一者,来确定每一对象是否参与行为事件。
行为事件的标识可以与图像帧序列相对应。在一些实施方式中,可以基于图像帧序列的拍摄时间戳、图像帧序列的真实拍摄位置信息、摄像头的编号信息等中的至少一者,确定行为事件的标识。例如,图像帧序列不同,行为事件的标识不同,图像帧序列和行为事件的标识之间存在一一对应的关系。
S104、检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息。
参与行为事件的至少一个对象可以包括:参与行为事件的至少一个人脸和/或至少一个人体。在一些实施方式中,通过检测图像帧序列中参与行为事件的至少一个对象,可以得到至少一个对象的对象框(对象的对象框包括对象框的坐标框和/或对象框的坐标框截图),例如,至少一个对象的对象框可以包括以下至少之一:至少一个人脸的人脸框、至少一个人脸的人脸框截图、至少一个人体的人体框、至少一个人体的人体框截图。
其中,本公开实施例中的对象框可以是对象框的坐标信息,人脸框可以是人脸框的坐标信息,人体框可以是人体框的坐标信息。本公开实施例中的对象框截图可以是对象框所对应的图像,人脸框截图可以是人脸框所对应的图像,人体框截图可以是人体框所对应的图像。
参与行为事件的至少一个对象为:从图像帧序列中确定的参与行为事件的至少一个对象。
至少一个对象中每一对象的对象信息可以包括以下至少之一:每一对象和/或每一对象的对象框在图像帧序列中每一图像中的位置信息、每一对象的特征信息、每一对象的属性信息、每一对象的标识信息、每一对象的对象标签值。
在本公开实施例中,对象可以包括人脸和/或人体,至少一个对象可以包括至少一个人脸和/或至少一个人体,每一对象可以包括每一人脸和/或每一人体。
S105、将所述行为事件的标识和所述至少一个对象的对象信息关联,得到行为识别结果。
在一些实施例中,还可以存储行为识别结果。例如,可以向存储设备发送行为识别结果,以使存储设备存储行为识别结果。存储设备可以是独立于行为识别装置之外的设备,例如,存储设备可以是分布式存储设备。在另一些实施方式中,存储行为识别结果可以包括:将行为识别结果存储在行为识别装置中。
在对行为识别结果进行检索的情况下,可以得到以下至少之一:行为事件的标识、至少一个对象的对象信息。在一些实施方式中,待检索的字段信息可以包括以下至少之一:待检索的时段信息、待检索的地址信息、待检索的人脸信息、待检索的人体信息、待检索的人名信息。
在一些实施方式中,行为识别结果可以用于在获取的待检索的字段信息与行为事件的标识对应的情况下,输出以下至少之一:行为事件的标识、至少一个对象的对象信息。
在一些实施方式中,图像帧序列中不同图像中同一对象的对象标识相同。在另一些实施方式中,图像帧序列中不同图像中同一对象的对象标识不同。
一个行为事件的标识可以和至少一个对象的对象信息关联,得到一个关联信息,关联信息可以表征行为事件的标识与每一个对象的标识之间的关联关系,从而基于关联关系和行为事件的标识,能够容易确定出图像帧序列中参与行为事件的每一对象。
例如,对于某一个图像帧序列,生成与该图像帧序列对应的行为事件的唯一标识,该行为事件的唯一标识可以为记为le,关联信息可以包括在事件的关联项(associations1)中。例如,associations1={{le,lf
i},{le,lp
j},...}。其中,关联信息包括:{le,lf
i}和{le,lp
j}等。其中,lf
i表示图像帧序列中参与行为事件的第i个人脸,lp
j表示图像帧序列中参与行为事件的第j个人体。
在本公开实施例中,由于通过对至少一个对象进行行为识别,得到图像帧序列中行为事件的标识,然后将行为事件的标识与参与行为事件的至少一个对象的对象信息关联,因此能够通过关联得到的行为识别结果,对参与行为事件的对象进行事后溯源。
图2为本公开实施例提供的另一种对事件的存储方法的实现流程示意图,如图2所示,该方法应用于对事件的存储装置,该方法包括:
S201、获取图像帧序列。
S202、对所述图像帧序列中检测到的至少一个对象进行跟踪,得到每个对象的跟踪信息。
S203、基于所述每个对象的跟踪信息,对所述至少一个对象进行行为识别,得到所述图像帧序列中行为事件的标识。
S204、确定所述图像帧序列的每一图像中的存在所述行为事件的事件区域。
S205、从所述每一图像中的存在所述行为事件的事件区域中,确定所述至少一个对象的对象信息。
在一些实施方式中,将每个对象的跟踪信息输入至深度神经网络,通过深度神经网络还可以得到:图像帧序列的每一图像中的存在行为事件的事件区域。接着行为识别装置可以从每一图像中的存在行为事件的事件区域中,确定至少一个对象。
在一些实施方式中,行为识别装置可以将每一图像中的存在行为事件的事件区域中所有的对象,确定为至少一个对象。例如,通过事件区域检测到一个人脸和两个人体的情况下,可以将该一个人脸和两个人体确定为至少一个对象。在另一些实施方式中,为了使得确定的参与行为事件的对象准确,可以对每一图像中的存在行为事件的事件区域进行识别,以得到至少一个对象。
S206、将所述行为事件的标识和所述至少一个对象的对象信息关联,得到行为识别结果。
在本公开实施例中,由于能够确定图像帧序列的每一图像中的存在行为事件的事件区域,从而能够在得到的事件区域中确定至少一个对象,缩小了用于检测对象的检测区域,从而能够快速地从事件区域检测到至少一个对象,提高了得到至少一个对象的对象信息的速度。
图3为本公开实施例提供的又一种行为识别方法的实现流程示意图,如图3所示,该方法应用于行为识别装置,该方法包括:
S301、获取图像帧序列。
S302、对所述图像帧序列中检测到的至少一个对象进行跟踪,得到每个对象的跟踪信息。
S303、基于所述每个对象的跟踪信息,对所述至少一个对象进行行为识别,得到所述图像帧序列中行为事件的标识。
S304、检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息。
其中,所述至少一个对象的对象信息包括所述至少一个对象的人脸标识和人体标识。
S305、将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果。
在一些实施例中,S305可以通过以下方式实现:在所述图像帧序列的每一图像中,确定参与所述行为事件的每一人脸与每一人体之间的位置关系;基于所述位置关系,将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果。
行为识别装置可以确定参与行为事件的每一人脸的人脸框,在每一图像中的位置信息,以及确定参与行为事件的每一人体在每一人体的人体框,在每一图像中的位置信息,然后基于这两个位置信息,确定参与行为事件的每一人脸与每一人体之间的位置关系。
通过这种方式,通过确定的参与行为事件的每一人脸与每一人体之间的位置关系,将属于同一对象的人脸标识和人体标识进行关联,从而不仅提供了一种人脸和人体的关联方案,且能够准确的确定关联的人脸和人体。
在一些实施例中,所述将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果,可以包括:从所述至少一个对象的人脸标识和人体标识中,获取属于同一对象的人脸标识和人体标识;向属于同一对象的人脸标识和人体标识对应相同的标签值;其中;不同对象的人脸标识对应不同的标签值,不同对象的人体标识对应不同的标签值;将所述相同的标签值对应的人脸标识和人体标识进行关联,得到所述关联结果。
在一些实施方式中,对象的对象信息可以包括人脸的标签值和/或人体的标签值。其中,属于同一个人的人脸的标签值与人体的标签值相同,或者,属于同一个人的人脸的标签值与人体的标签值具有映射关系。
在一些实施方式中,可以将行为事件的标识,至少一个人脸的标识和人脸的标签值确定为行为识别结果。在另一些实施方式中,可以将行为事件的标识、至少一个人体的标识和每一指定人体的标签值确定为行为识别结果。在又一些实施方式中,可以将行为事件的标识、至少一个人脸的标识、至少一个人脸的标签值、至少一个人体的标识和至少一个人体的标签值确定为行为识别结果。
通过这种方式,通过获取属于同一对象的人脸标识和人体标识,向属于同一对象的人脸标识和人体标识对应相同的标签值,将相同的标签值对应的人脸标识和人体标识进行关联,得到关联结果,从而能够容易地将属于同一对象的人脸标识和人体标识进行关联。
S306、将所述行为事件的标识和所述关联结果进行关联,得到所述行为识别结果。
在一些实施例中,将所述行为事件的标识和所述关联结果进行关联,可以得到上述的关联信息。
在一些实施方式中,行为识别装置可以获取所有人脸的标识中每一人脸标识对应的人脸信息,在该人脸的人脸信息中确定是否有人脸标签值,或者,在确定该人脸的人脸信息中人脸标签值不为零的情况下,确定该人脸存在关联的人体,从而能够将属于同一对象的人脸标识和人体标识进行关联,得到关联结果。在另一些实施方式中,行为识别装置可以确定所有人脸中的每一个人脸对应的人脸框是否包括在另一个框之内,在是的情况下,确定该人脸存在关联的人体,且关联的人体为另一个框内的人体,从而能够将属于同一对象 的人脸标识和人体标识进行关联,得到关联结果;在否的情况下,确定该人脸不存在关联的人体。人脸不存在关联的人体的一种场景是该人体被遮挡,从而识别不到该人脸对应的人体。
例如,在关联信息包括{le,lf
1},{le,lf
2},{le,lf
3},且确定具有关联的人体的每一指定人脸标识为lf
1和lf
3的情况下,确定人脸标识lf
1关联的人体标识为lp
1,确定人脸标识lf
3关联的人体标识为lp
2,将人脸标识lf
1与人体标识为lp
1关联,得到{lf
1,lp
1},将人脸标识lf
3与人体标识为lp
2关联,得到{lf
3,lp
2}。
在一些实施方式中,关联结果可以包括在人脸的关联项(associations2)中,associations2={{le,lf
i},{lf
i,lp
j}},{lf
i,lp
j}表征人脸的标识lf
i与人体的标识lp
j关联。在一些实施方式中,关联结果可以是associations2或者{lf
i,lp
j}。
在一些实施方式中,行为识别装置可以获取所有人体的标识中每一人体标识对应的人体信息,在该人体的人体信息中确定是否有人体标签值,或者,在确定该人体的人体信息中人体标签值不为零的情况下,确定该人体存在关联的人脸,从而能够将属于同一对象的人脸标识和人体标识进行关联,得到关联结果。在另一些实施方式中,行为识别装置可以确定所有人体中的每一个人体对应的人体框包括有另一个框,在是的情况下,确定该人体存在关联的人脸,且关联的人脸为另一个框内的人脸,从而能够将属于同一对象的人脸标识和人体标识进行关联,得到关联结果;在否的情况下,确定该人体不存在关联的人脸。人体不存在关联的人脸的一种场景是该人脸被遮挡,从而识别不到该人体对应的人脸。
例如,在关联信息包括{le,lp
1},{le,lp
2},{le,lp
3},且确定具有关联的人脸的每一指定人体标识为lp
1和lp
2的情况下,确定人体标识lp
1关联的人脸标识为lf
1,确定人体标识lp
2关联的人脸标识为lf
3,将人体标识lp
1与人脸标识为lf
1关联,得到{lf
1,lp
1},将人体标识lp
2与人脸标识为lf
3关联,得到{lp
2,lf
3}。
在一些实施方式中,关联结果可以包括在人体的关联项(associations3)中,associations3={{le,lp
j},{lp
j,lf
i}},{lp
j,lf
i}表征人体的标识lp
j与人脸的标识lf
i关联。在一些实施方式中,关联结果可以是associations3或者{lp
j,lf
i}。
在本公开实施例中,通过将属于同一对象的人脸标识和人体标识进行关联,得到关联结果,然后将行为事件的标识和关联结果进行关联得到行为识别结果,从而能够从行为识别结果中获取参与行为事件的人脸和人体,使得信息的检索更加全面。
图4为本公开实施例提供的再一种行为识别方法的实现流程示意图,如图4所示,在本公开实施例中,对象信息包括:人脸特征信息和/或人脸属性信息;该方法应用于行为识别装置,该方法包括:
S401、获取图像帧序列。
S402、对所述图像帧序列中检测到的至少一个对象进行跟踪,得到每个对象的跟踪信息。
S403、基于所述每个对象的跟踪信息,对所述至少一个对象进行行为识别,得到所述图像帧序列中行为事件的标识。
S404、检测所述图像帧序列中参与所述行为事件的所述至少一个对象的人脸特征信息。
S405、基于所述至少一个对象的人脸特征信息,确定所述至少一个对象的人脸属性信息。
S406、基于所述至少一个对象的所述人脸特征信息和/或所述人脸属性信息,确定所述至少一个对象的身份信息。
在一些实施例中,在确定到至少一个对象的身份信息的情况下,可以输出该至少一个对象的身份信息,以使相关人员基于至少一个对象的身份信息能够快速的确定参与行为事件的对象的身份。
S407、将所述行为事件的标识和所述至少一个对象的对象信息关联,得到行为识别结果。
在一些实施方式中,人脸位置信息、人脸特征信息、人脸属性信息、至少一个人脸的标识中的至少一者,包括在对象的对象信息中。人脸位置信息可以是:在图像帧序列的每一图像中至少一个人脸在每一图像中的人脸位置信息。
至少一个人脸在每一图像中的人脸位置信息,可以是至少一个人脸的人脸框在每一图像中的位置信息。例如,至少一个人脸的人脸框在每一图像中的位置信息可以包括:至少一个人脸的人脸框的左上角和右下角在每一图像中的位置信息。
通过确定至少一个人脸的人脸特征信息,可以将至少一个人脸中每一人脸的特征信息与人脸库中的一个或多个图像进行匹配,以确定至少一个人脸中每一人脸所属的人的名字、学校、班级、工作单位、身份标识、家属的属性信息等中的至少一者。家属的属性信息可以包括家属的姓名、性别、与参与行为事件的对象之间的关系、联系方式等中的至少之一。
人脸的属性信息可以包括以下之一:是否戴眼镜、是否戴口罩、年龄、性别、五官的属性信息等信息。
在本公开实施例中,基于至少一个对象的人脸特征信息和/或人脸属性信息,确定至少一个对象的身份信息,从而能够通过对象的人脸信息对参与行为事件的身份信息进行确定,进而能够对参与行为事件的对象进行事后溯源。
图5为本公开另一实施例提供的一种行为识别方法的实现流程示意图,如图5所示,在本公开实施例中,对象信息包括:人体特征信息和/或人体属性信息,该方法应用于行为识别装置,该方法包括:
S501、获取图像帧序列。
S502、对所述图像帧序列中检测到的至少一个对象进行跟踪,得到每个对象的跟踪信息。
S503、基于所述每个对象的跟踪信息,对所述至少一个对象进行行为识别,得到所述图像帧序列中行为事件的标识。
S504、检测所述图像帧序列中参与所述行为事件的所述至少一个对象的人体特征信息。
S505、基于所述至少一个对象的人体特征信息,确定所述至少一个对象的人体属性信息。
S506、基于所述至少一个对象的所述人体特征信息和/或所述人体属性信息,确定所述至少一个对象的身份信息。
S507、将所述行为事件的标识和所述至少一个对象的对象信息关联,得到行为识别结果。
在一些实施方式中,人体位置信息、人体特征信息、人体属性信息、至少一个人体的标识中的至少一者,包括在对象的对象信息中。人体位置信息可以是:在图像帧序列的每一图像中至少一个人体在每一图像中的人体位置信息。
至少一个人体在每一图像中的人体位置信息,可以是至少一个人体的人体框在每一图像中的位置信息。例如,至少一个人体的人体框在每一图像中的位置信息可以包括:至少一个人体的人体框的左上角和右下角在每一图像中的位置信息。
通过确定至少一个人体的人体特征信息,可以将至少一个人体中每一人体的特征信息与人体库中的一个或多个图像进行匹配,已确定至少一个人体中每一人体所属的人的名字、学校、班级、工作单位、身份标识、家属的属性信息等中的至少一者。
人体的属性信息可以包括以下之一:身高、体型、体重、穿着信息等信息。
在本公开实施例中,基于至少一个对象的人体特征信息和/或人体属性信息,确定至少一个对象的身份信息,从而能够通过对象的人体信息对参与行为事件的身份信息进行确定,进而能够对参与行为事件的对象进行事后溯源。
图6为本公开又一实施例提供的一种行为识别方法的实现流程示意图,如图6所示,该方法应用于行为识别装置,该方法包括:
S601、获取图像帧序列。
S602、对所述图像帧序列中检测到的至少一个对象进行跟踪,得到每个对象的跟踪信息。
S603、基于所述每个对象的跟踪信息,对所述至少一个对象进行行为识别,得到所述图像帧序列中行为事件的标识。
S604、检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息。
S605、将所述行为事件的标识和所述至少一个对象的对象信息关联,得到行为识别结果。
S606、在所述行为事件为预设事件的情况下,输出告警信息。
预设事件可以为预先设定的危险事件,例如,预设事件可包括以下至少之一:打架事件、骑电动车不戴头盔事件、行驶超速事件、不按规定标示行驶的事件。
在一些实施方式中,输出告警信息可以包括:向指定设备输出告警信息,其中,指定设备包括以下至少之一:显示设备、发生行为事件区域的工作人员的终端设备、参与行为事件的对象的家属的终端设备。
其中,所述告警信息包括以下在所述图像帧序列中展示的至少之一:发生所述行为事件的事件区域、所述行为事件的时空信息、参与所述行为事件对象的人脸框、参与所述行为事件对象的人体框、所述人脸框中的人脸图像、所述人体框中的人体图像、参与所述行为事件的对象的人脸和/或人体的属性信息、参与所述行为事件的对象关联家属的属性信息。
通过告警信息包括参与行为事件的对象的人脸框、参与行为事件的对象的人体框、发生行为事件的事件区域中的至少一者,从而能够在显示的实时视频上显示相应地人脸框、人体框以及行为事件的事件区域中的至少一者,使得工作人员能够通过显示的实时视频即可容易得知发生行为事件。
在本公开实施例中,在行为事件为预设事件的情况下,可以输出告警信息,从而使得工作人员能够确定存在预设事件,进而能够及时地对预设事件进行处理。
在一些实施方式中,摄像头采集视频流,并将采集到的视频流发送至服务器,以使服务器执行本公开实施例提供的行为识别方法的步骤,服务器在确定到某一图像帧序列发生行为事件的情况下,可以通过告警页面显示发生行为事件的告警。在多个连续的图像帧序列均发生行为事件的情况下,可以输出一次告警。在一些实施方式中,事件持续达到第一预设时长后,输出一次告警,并输出多个间隔预设第二时长的关键帧的检测结果。
可以通过以下几个模块实现本公开实施例提供的行为识别方法:视频解码模块、目标检测跟踪模块、打架检测模块、人脸人体检测模块、人脸人体匹配模块、特征属性提取模块、事件输出模块以及告警显示模块。其中,这些模块可以在一个设备上设置或者在不同的设备上设置。
视频解码模块用于对接入的视频流进行解码。目标检测跟踪模块用于对视频流中的行人进行检测,并跟踪。打架检测模块用于将图像帧序列输入人体字体估计建模型(即上述的深度神经网络),来获取行为事件的估计。人脸人体检测模块用于对有打架区域的行人进行人脸人体的检测。人脸人体匹配模块用于对人脸人体检测到的人脸人体进行匹配。特征属性提取模块用于提取人脸人体的特征信息和属性信息。事件输出模块用于向告警显示模块输出事件发生区域,事件,参与事件的人脸人体信息。告警显示模块用于进行告警显示。
目标检测跟踪模块用于检测出视频中出现的行人,采用人体检测算法,本公开实施例中使用基于深度学习的图像检测算法,相比传统算法检测率要高,能够应对复杂环境的人体,并对这些人分体对象进行跟踪。
打架检测模块用于将检测到的人体输入到人体姿态估计模型中,将视频中的人体事件序列作为输入,每隔一段时间输出一个打架事件的估计结果,充分考虑了视频的时间与空间信息,提高了结果输出的可靠性。
人脸人体匹配模块用于将检测到的人脸与人体通过空间几何关系,把属于同一个对象人脸与人体关联起来。
事件输出模块用于将参与打架事件的人脸人体特征,属性等信息组织输出到指定位置,供下游存储或后面检索。
告警显示模块用于以全球广域网(World Wide Web,Web)页面展示发生打架区域。
人脸人体匹配模块是对打架区域内出现的行人的人脸集合与人体集合找到其对应关系。对于人脸集合{f
1,f
2,...,f
n}以及人体集合{p
1,p
2,...,p
m},人脸集合中的每一项在人体集合中存在0或1个对应,对人脸集合中的元素,比如f
i,生成一个唯一的标记,记为lf
i;同样对于人体集合中的元素,比如p
j,生成一个唯一的标记,记为lp
j。假如对于人脸f
i存在对应的人体p
j,那么将f
i中新增一个标记mf
i,并令mf
i=lp
j,同理在p
j中新增一个标记mp
j,令mp
j=lf
i。
人脸人体属性检测模块是对检测到的人脸集合与人体集合进行属性与特征的提取。对于人脸集合{f
1,f
2,...,f
n}中的每一个元素,比如f
i提取对应的特征fe
i,以及人脸的属性fa
i等基本信息,这样形成了一个新的集合{F
1,F
2,...,F
n},其中F
i={f
i,fe
i,fa
i,lf
i,mf
i}(如果mf
i存在),同理集合{P
1,P
2,...,P
n},其中P
i={p
i,pe
i,pa
i,lp
i,mp
i}(如果mp
i存在).
事件输出模块是对事件的整体封装,当打架检测模块检测到打架事件发生时,会生成事件E,对于事件E生成其唯一的标记,记为le,输出事件时,设置关联项associations,其中包含与事件相关的所有人脸人体,即E中的associations={{le,lf
i},{le,lp
j},...},对于参与事件的人脸,假设为F
i,设置了其中的关联项associations,其中包含了事件本身以及与之可能对应的人体,即F
i中的associations={{le,lf
i},{lf
i,lp
j}},同理对于人体P
j,其P
j的associations={{le,lp
j},{lp
j,lf
i}}。如此事件与人脸人体的关联便已完成,其中人脸,人体以及事件分别输出存储。以其关联关系则可完成事件以及人脸人体的索引。
本公开实施例利用深度学习检测行人轨迹,利用图像帧序列来分析对行人模型行为建模,可以对视频中的行人进行准确的行为分析,能够输出准确的行为预测结果。
本公开实施例输出了整个打架事件,完成了整个事件闭环。从参与人员出现开始跟踪,到出现打架事件开始识别行为,输出打架事件,并输出参与打架的行为事件的人员的基本信息,最后将这些信息存储,供工作人员使用。
本公开实施例在摄像头存在的各种场所均可应用本公开实施例提供的防范,包括但不限于交通场所、商场、车站、学校或广场等。
本公开实施例还可以提供一种行为识别方法,该方法应用于行为识别装置,在得到行为识别结果之后,该方法还可以包括:获取待检索的字段信息;从行为识别结果集合中,获取与所述待检索的字段信息对应的目标行为事件的标识,和/或,参与所述目标行为事件对象的对象信息。在一些实施例中,还可以输出目标行为事件的标识,和/或,参与目标行为事件对象的对象信息。
在一些实施方式中,行为识别装置可以包括显示屏,工作人员通过显示屏输入待检索 的字段信息,从而行为识别装置得到待检索的字段信息。在另一些实施方式中,工作人员通过不同于行为识别装置的终端设备输入待检索的字段信息,终端设备将待检索的字段信息向行为识别装置发送,以使行为识别装置得到待检索的字段信息。
待检索的字段信息包括:待检索的时段信息和待检索的地址信息。
在一些实施方式中,获取待检索的字段信息可以通过以下方式实现:响应于输入的待检索的时段信息和待检索的地址信息,得到待检索的时段信息和待检索的地址信息。
时段信息可以是连续的某一时段信息或者非连续的至少两个时段信息。
在这种实施方式中,工作人员可以直接在行为识别装置的显示屏或者终端设备的显示屏上,输入待检索的时段信息和待检索的地址信息,以使行为识别装置得到待检索的时段信息和待检索的地址信息。其中,地址信息例如可以是文艺路,或者文艺路与建设路交叉口等。
在另一些实施方式中,获取待检索的字段信息可以通过以下方式实现:响应于输入的待检索的时段信息,输出与待检索的时段信息匹配的至少一个地址信息;响应于对至少一个地址信息中待检索的地址信息的触发操作,得到待检索的地址信息。
在这种实施方式中,工作人员在行为识别装置的显示屏或者终端设备的显示屏上,输入待检索的时段信息之后,可以在显示屏上弹出在该待检索的时段信息内发生行为事件的至少一个地址信息,工作人员可以在显示的至少一个地址信息选择待检索的地址信息。
本公开实施例中的待检索的地址信息可以一个或多个地址信息。在一些实施方式中,至少一个地址信息还可以与工作人员登陆的账号的权限有关。
通过这种方式,工作人员可以先输入待检索的时段信息,从而可以基于待检索的时段信息使工作人员确定到哪些地址发生了行为事件,从而工作人员可以根据显示的至少一个地址信息来选择工作人员所关心的待检索的地址信息,进而查看待检索的地址信息所存在的参与事件对象的对象信息。
在又一些实施方式中,获取待检索的字段信息可以通过以下方式实现:响应于输入的待检索的地址信息;输出与待检索的地址信息匹配的至少一个时段信息;响应于对至少一个时段信息中待检索的时段信息的触发操作,得到待检索的时段信息。
在这种方式下,行为识别装置获取的目标行为事件的标识不仅与工作人员感兴趣的待检索的时段信息有关,还与工作人员感兴趣的待检索的地址信息有关,从而使得工作人员可以容易地在海量的信息中,检索到感兴趣的待检索的时段和待检索的地址信息中参与行为事件对象的对象信息。
行为识别结果集合可以与多个图像帧序列对应。行为识别结果集合可以包括至少一个图像帧序列的行为识别结果,每一图像帧序列的行为识别结果包括:图像帧序列中发生行为事件的标识,与参与行为事件的至少一个对象的对象信息。
需要注意的是,不同的行为事件的标识所代表的可以均是打架事件的标识。
至少一个行为事件的标识可以与至少一个图像帧序列一一对应。目标行为事件的标识可以是一个或多个。例如,在输入的时段信息所对应的时长为一小时,且设定时长为一分钟的情况下,得到的目标行为事件的标识为60个。
在本公开实施例中,通过获取与待检索的字段信息对应的目标行为事件的标识,和/或,参与目标行为事件对象的对象信息,从而工作人员能够容易得到与感兴趣的待检索的字段信息对应的行为事件或者对象信息。
在一些实施例中,参与目标行为事件对象的对象信息包括:与目标行为事件的标识关联的所有人脸的标识对应的人脸信息,和/或,与目标行为事件的标识关联的所有人体的标识对应的人体信息。在一些实施例中,在得到行为识别结果之后,该方法还可以包括:获取待检索的字段信息;从行为识别结果集合中,获取与待检索的字段信息对应的目标行为事件的标识;从行为识别结果集合中存储的多个关联信息中,获取与目标行为事件的标识 关联的所有人脸的标识和/或所有人体的标识;从行为识别结果集合中的对象信息中,获取与所有人脸的标识对应的人脸信息,和/或,与所有人体的标识对应的人体信息;输出目标行为事件的标识,和/或,参与目标行为事件对象的对象信息。多个关联信息可以包括每一事件的标识与对象的标识之间的关联信息,例如,多个关联信息可以包括以下内容:{le1,lf
i},{le1,lp
j},{le2,lf
i},{le2,lp
j}等等。其中,{le1,lf
i}表示行为事件1的标识le1与人脸标识lf
i之间的关联;{le1,lp
j}表示行为事件1的标识le1与人脸标识lp
j之间的关联;{le2,lf
i}表示行为事件2的标识le2与人脸标识lf
i之间的关联;{le2,lp
j}表示行为事件2的标识le2与人脸标识lp
j之间的关联。由于每一关联信息包括了行为事件的标识与对象的标识的关联信息,从而可以基于多个关联信息,获取与目标行为事件的标识关联的所有人脸的标识和/或所有人体的标识。
通过在行为识别装置或者存储设备中存储至少一个对象的对象信息,从而可以所有人脸的标识和/或所有人体的标识,得到与所有人脸的标识对应的人脸信息和与所有人体的标识对应的人体信息。
在本公开实施例中,通过存储的多个关联信息,获取与目标行为事件的标识关联的所有人脸的标识和/或所有人体的标识,从而能够容易地获取到与所有人脸的标识对应的人脸信息和/或与所有人体的标识对应的人体信息。
在一些实施例中,参与目标行为事件对象的对象信息包括:所有人脸中指定人脸关联的人体的标识对应的人体信息。在一些实施例中,在得到行为识别结果之后,该方法还可以包括:获取待检索的字段信息;从行为识别结果集合中,获取与待检索的字段信息对应的目标行为事件的标识;从行为识别结果集合中存储的多个关联信息中,获取与目标行为事件的标识关联的所有人脸的标识;从行为识别结果集合中的对象信息中,获取与所有人脸的标识对应的人脸信息;从行为识别结果集合中存储的关联结果中,获取所有人脸中指定人脸关联的人体的标识;指定人脸为具有关联的人体的人脸;从行为识别结果集合中存储的对象信息中,获取与关联的人体的标识对应的人体信息;输出目标行为事件的标识,和/或,参与目标行为事件对象的对象信息。
在另一些实施例中,从行为识别结果集合中存储的对象信息中,获取与所有人脸的标识对应的人脸信息之后,还可以执行:从行为识别结果集合中存储的关联结果中,获取所有人体中指定人体关联的人脸的标识;指定人体为具有关联的人脸的人体;从行为识别结果集合中存储的对象信息中,获取与关联的人脸的标识对应的人脸信息;输出目标行为事件的标识,和/或,参与目标行为事件对象的对象信息。
关联结果可以包括任一人脸的标识所关联的人体的标识,例如,关联结果可以包括associations2={{le,lf
i},{lf
i,lp
j}}或者{lf
i,lp
j},从而可以基于关联结果确定指定人脸关联的人体的标识。
在本公开实施例中,不仅获取所有人脸的标识对应的人脸信息,还获取关联的人体的标识对应的人体信息,从而使得工作人员不仅可以得知参与行为事件的人脸信息,还能够得知参与行为事件的人脸所匹配的人体信息,进而使得工作人员能够更加全面的了解到参与行为事件的对象。
关联结果可以包括任一人体的标识所关联的人脸的标识,例如,关联结果可以包括associations3={{le,lp
j},{lp
j,lf
i}}或者{lp
j,lf
i},从而可以基于关联结果确定指定人体关联的人脸的标识。
在本公开实施例中,不仅获取所有人体的标识对应的人体信息,还获取关联的人脸的标识对应的人脸信息,从而使得工作人员不仅可以得知参与行为事件的人体信息,还能够得知参与行为事件的人体所匹配的人脸信息,进而使得工作人员能够更加全面的了解到参 与行为事件的对象。
基于前述的实施例,本公开实施例提供一种行为识别装置,该装置包括所包括的各模块,可以通过终端设备中的处理器来实现。
图7为本公开实施例提供的一种行为识别装置的组成结构示意图,如图7所示,行为识别装置700包括:获取部分701,配置为获取图像帧序列;跟踪部分702,配置为对所述图像帧序列中检测到的至少一个对象进行跟踪,得到每个对象的跟踪信息;识别部分703,配置为基于所述每个对象的跟踪信息,对所述至少一个对象进行行为识别,得到所述图像帧序列中行为事件的标识;检测部分704,配置为检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息;关联部分705,配置为将所述行为事件的标识和所述至少一个对象的对象信息关联,得到行为识别结果。
在一些实施例中,检测部分704,还配置为确定所述图像帧序列的每一图像中的存在所述行为事件的事件区域;从所述每一图像中的存在所述行为事件的事件区域中,确定所述至少一个对象的对象信息。
在一些实施例中,所述至少一个对象的对象信息包括所述至少一个对象的人脸标识和人体标识;关联部分705,还配置为将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果;将所述行为事件的标识和所述关联结果进行关联,得到所述行为识别结果。
在一些实施例中,关联部分705,还配置为在所述图像帧序列的每一图像中,确定参与所述行为事件的每一人脸与每一人体之间的位置关系;基于所述位置关系,将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果。
在一些实施例中,关联部分705,还配置为从所述至少一个对象的人脸标识和人体标识中,获取属于同一对象的人脸标识和人体标识;向属于同一对象的人脸标识和人体标识对应相同的标签值;其中;不同对象的人脸标识对应不同的标签值,不同对象的人体标识对应不同的标签值;将所述相同的标签值对应的人脸标识和人体标识进行关联,得到所述关联结果。
在一些实施例中,所述对象信息包括:人脸特征信息和/或人脸属性信息;检测部分704,还配置为检测所述图像帧序列中参与所述行为事件的所述至少一个对象的人脸特征信息;基于所述至少一个对象的人脸特征信息,确定所述至少一个对象的人脸属性信息;行为识别装置700还包括:确定部分706,配置为基于所述至少一个对象的所述人脸特征信息和/或所述人脸属性信息,确定所述至少一个对象的身份信息。
在一些实施例中,所述对象信息包括:人体特征信息和/或人体属性信息,检测部分704,还配置为检测所述图像帧序列中参与所述行为事件的所述至少一个对象的人体特征信息;基于所述至少一个对象的人体特征信息,确定所述至少一个对象的人体属性信息;确定部分706,还配置为基于所述至少一个对象的所述人体特征信息和/或所述人体属性信息,确定所述至少一个对象的身份信息。
在一些实施例中,行为识别装置700还包括:输出部分707,配置为在所述行为事件为预设事件的情况下,输出告警信息;所述告警信息包括以下在所述图像帧序列中展示的至少之一:发生所述行为事件的事件区域、所述行为事件的时空信息、参与所述行为事件对象的人脸框、参与所述行为事件对象的人体框、所述人脸框中的人脸图像、所述人体框中的人体图像、参与所述行为事件的对象的人脸和/或人体的属性信息、参与所述行为事件的对象关联家属的属性信息。
以上装置实施例的描述,与上述方法实施例的描述是类似的,具有同方法实施例相似的有益效果。对于本公开装置实施例中未披露的技术细节,请参照本公开方法实施例的描述而理解。
需要说明的是,本公开实施例中,如果以软件功能模块的形式实现上述的行为识别方法,并作为独立的产品销售或使用时,也可以存储在一个计算机存储介质中。基于这样的理解,本公开实施例的技术方案本质上或者说对相关技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台终端设备执行本公开各个实施例方法的全部或部分。
图8为本公开实施例提供的一种电子设备的硬件实体示意图,如图8所示,该电子设备800的硬件实体包括:处理器801和存储器802,其中,存储器802存储有可在处理器801上运行的计算机程序,处理器801执行程序时实现上述任一实施例的方法中的步骤。电子设备800可以包括上述的行为识别装置。
存储器802存储有可在处理器上运行的计算机程序,存储器802配置为存储由处理器801可执行的指令和应用,还可以缓存待处理器801以及电子设备800中各模块待处理或已经处理的数据(例如,图像数据、音频数据、语音通信数据和视频通信数据),可以通过闪存(FLASH)或随机访问存储器(Random Access Memory,RAM)实现。处理器801执行程序时实现上述任一项的行为识别方法的步骤。处理器801通常控制电子设备800的总体操作。
本公开实施例提供一种计算机存储介质,计算机存储介质存储有一个或者多个程序,该一个或者多个程序可被一个或者多个处理器执行,以实现如上任一实施例的行为识别方法的步骤。
图9为本公开实施例的芯片的示意性结构图。图9所示的芯片900包括处理器910,用于从存储器中调用并运行计算机程序,使得安装有所述芯片的设备执行上述任一实施例中的方法。
在一些实施例中,如图9所示,芯片900还可以包括存储器920。其中,处理器910可以从存储器920中调用并运行计算机程序,以实现本公开实施例中的方法。
其中,存储器920可以是独立于处理器910的一个单独的器件,也可以集成在处理器910中。
在一些实施例中,该芯片900还可以包括输入接口930。其中,处理器910可以控制该输入接口930与其他设备或芯片进行通信,示例性地,可以获取其他设备或芯片发送的信息或数据。
在一些实施例中,该芯片900还可以包括输出接口940。其中,处理器910可以控制该输出接口940与其他设备或芯片进行通信,示例性地,可以向其他设备或芯片输出信息或数据。
在一些实施例中,该芯片可应用于本公开实施例中的控制网元或执行网元,并且该芯片可以实现本公开实施例的各个方法中由控制网元或执行网元实现的相应流程,为了简洁,在此不再赘述。应理解,本公开实施例提到的芯片还可以称为系统级芯片,系统芯片,芯片系统或片上系统芯片等。
本公开实施例还可以提供一种计算机程序产品,所述计算机程序产品承载有程序代码,所述程序代码包括的指令可配置为执行如上述任一方法中的步骤。
本公开实施例还可以提供一种计算机程序,包括计算机可读代码,在所述计算机可读代码在电子设备中运行的情况下,所述电子设备中的处理器执行如上述任一方法中的步骤。
这里需要指出的是:以上行为识别装置、电子设备、计算机存储介质、芯片、计算机程序产品实施例的描述,与上述方法实施例的描述是类似的,具有同方法实施例相似的有益效果。对于本公开行为识别装置、电子设备、计算机存储介质、芯片、计算机程序产品实施例中未披露的技术细节,请参照本公开方法实施例的描述而理解。
上述行为识别装置、芯片或处理器可以包括以下任一个或多个的集成:特定用途集成电路(Application Specific Integrated Circuit,ASIC)、数字信号处理器(Digital Signal Processor,DSP)、数字信号处理装置(Digital Signal Processing Device,DSPD)、可编程逻辑装置(Programmable Logic Device,PLD)、现场可编程门阵列(Field Programmable Gate Array,FPGA)、中央处理器(Central Processing Unit,CPU)、图形处理器(Graphics Processing Unit,GPU)、嵌入式神经网络处理器(neural-network processing units,NPU)、控制器、微控制器、微处理器。可以理解地,实现上述处理器功能的电子器件还可以为其它,本公开实施例不作具体限定。
上述计算机存储介质/存储器可以是只读存储器(Read Only Memory,ROM)、可编程只读存储器(Programmable Read-Only Memory,PROM)、可擦除可编程只读存储器(Erasable Programmable Read-Only Memory,EPROM)、电可擦除可编程只读存储器(Electrically Erasable Programmable Read-Only Memory,EEPROM)、磁性随机存取存储器(Ferromagnetic Random Access Memory,FRAM)、快闪存储器(Flash Memory)、磁表面存储器、光盘、或只读光盘(Compact Disc Read-Only Memory,CD-ROM)等存储器;也可以是包括上述存储器之一或任意组合的各种终端,如移动电话、计算机、平板设备、个人数字助理等。
应理解,说明书通篇中提到的“一个实施例”或“一实施例”或“本公开实施例”或“前述实施例”或“一些实施方式”或“一些实施例”意味着与实施例有关的特定特征、结构或特性包括在本公开的至少一个实施例中。因此,在整个说明书各处出现的“在一个实施例中”或“在一实施例中”或“本公开实施例”或“前述实施例”或“一些实施方式”或“一些实施例”未必一定指相同的实施例。此外,这些特定的特征、结构或特性可以任意适合的方式结合在一个或多个实施例中。应理解,在本公开的各种实施例中,上述各过程的序号的大小并不意味着执行顺序的先后,各过程的执行顺序应以其功能和内在逻辑确定,而不应对本公开实施例的实施过程构成任何限定。上述本公开实施例序号仅仅为了描述,不代表实施例的优劣。
在未做特殊说明的情况下,行为识别装置执行本公开实施例中的任一步骤,可以是行为识别装置的处理器执行该步骤。除非特殊说明,本公开实施例并不限定行为识别装置执行下述步骤的先后顺序。另外,不同实施例中对数据进行处理所采用的方式可以是相同的方法或不同的方法。还需说明的是,本公开实施例中的任一步骤是行为识别可以独立执行的,即行为识别装置执行上述实施例中的任一步骤时,可以不依赖于其它步骤的执行。
在本公开所提供的几个实施例中,应该理解到,所揭露的设备和方法,可以通过其它的方式实现。以上所描述的设备实施例仅仅是示意性的,例如,所述模块的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,如:多个模块或组件可以结合,或可以集成到另一个系统,或一些特征可以忽略,或不执行。另外,所显示或讨论的各组成部分相互之间的耦合、或直接耦合、或通信连接可以是通过一些接口,设备或模块的间接耦合或通信连接,可以是电性的、机械的或其它形式的。
上述作为分离部件说明的模块可以是、或也可以不是物理上分开的,作为模块显示的部件可以是、或也可以不是物理模块;既可以位于一个地方,也可以分布到多个网络模块上;可以根据实际的需要选择其中的部分或全部模块来实现本实施例方案的目的。
另外,在本公开各实施例中的各功能模块可以全部集成在一个处理模块中,也可以是各模块分别单独作为一个模块,也可以两个或两个以上模块集成在一个模块中;上述集成的模块既可以采用硬件的形式实现,也可以采用硬件加软件功能模块的形式实现。
本公开所提供的几个方法实施例中所揭露的方法,在不冲突的情况下可以任意组合,得到新的方法实施例。本公开所提供的几个产品实施例中所揭露的特征,在不冲突的情况下可以任意组合,得到新的产品实施例。
本公开所提供的几个方法或设备实施例中所揭露的特征,在不冲突的情况下可以任意组合,得到新的方法实施例或设备实施例。
本领域普通技术人员可以理解:实现上述方法实施例的全部或部分步骤可以通过程序指令相关的硬件来完成,前述的程序可以存储于计算机存储介质中,该程序在执行时,执行包括上述方法实施例的步骤;而前述的存储介质包括:移动存储设备、只读存储器(Read Only Memory,ROM)、磁碟或者光盘等各种可以存储程序代码的介质。
或者,本公开上述集成的模块如果以软件功能模块的形式实现并作为独立的产品销售或使用时,也可以存储在一个计算机存储介质中。基于这样的理解,本公开实施例的技术方案本质上或者说对相关技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机、服务器、或者网络设备等)执行本公开各个实施例所述方法的全部或部分。而前述的存储介质包括:移动存储设备、ROM、磁碟或者光盘等各种可以存储程序代码的介质。
在本公开实施例中,不同实施例中相同步骤和相同内容的说明,可以互相参照。在本公开实施例中,术语“并”不对步骤的先后顺序造成影响,例如,电子设备执行A,并执行B,可以是电子设备先执行A,再执行B,或者是电子设备先执行B,再执行A,或者是电子设备执行A的同时执行B。
在本公开实施例和所附权利要求书中所使用的单数形式的“一种”、“所述”和“该”也旨在包括多数形式,除非上下文清楚地表示其他含义。
应当理解,本文中使用的术语“和/或”仅仅是一种描述关联对象的关联关系,表示可以存在三种关系,例如,A和/或B,可以表示:单独存在A,同时存在A和B,单独存在B这三种情况。另外,本文中字符“/”,一般表示前后关联对象是一种“或”的关系。
需要说明的是,本公开所涉及的各个实施例中,可以执行全部的步骤或者可以执行部分的步骤,只要能够形成一个完整的技术方案即可。
以上所述,仅为本公开的实施方式,但本公开的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本公开揭露的技术范围内,可轻易想到变化或替换,都应涵盖在本公开的保护范围之内。因此,本公开的保护范围应以所述权利要求的保护范围为准。
本公开实施例提供一种行为识别方法、装置、设备、介质、芯片、产品及程序,其中,由于通过对至少一个对象进行行为识别,得到图像帧序列中行为事件的标识,然后将行为事件的标识与参与行为事件的至少一个对象的对象信息关联,因此能够通过关联得到的行为识别结果,对参与行为事件的对象进行事后溯源。
Claims (21)
- 一种行为识别方法,所述方法包括:获取图像帧序列;对所述图像帧序列中检测到的至少一个对象进行跟踪,得到每个对象的跟踪信息;基于所述每个对象的跟踪信息,对所述至少一个对象进行行为识别,得到所述图像帧序列中行为事件的标识;检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息;将所述行为事件的标识和所述至少一个对象的对象信息关联,得到行为识别结果。
- 根据权利要求1所述的方法,其中,所述检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息,包括:确定所述图像帧序列的每一图像中的存在所述行为事件的事件区域;从所述每一图像中的存在所述行为事件的事件区域中,确定所述至少一个对象的对象信息。
- 根据权利要求1或2所述的方法,其中,所述至少一个对象的对象信息包括所述至少一个对象的人脸标识和人体标识;所述将所述行为事件的标识和所述至少一个对象的对象信息关联,得到行为识别结果,包括:将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果;将所述行为事件的标识和所述关联结果进行关联,得到所述行为识别结果。
- 根据权利要求3所述的方法,其中,所述将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果,包括:在所述图像帧序列的每一图像中,确定参与所述行为事件的每一人脸与每一人体之间的位置关系;基于所述位置关系,将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果。
- 根据权利要求3或4所述的方法,其中,所述将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果,包括:从所述至少一个对象的人脸标识和人体标识中,获取属于同一对象的人脸标识和人体标识;向属于同一对象的人脸标识和人体标识对应相同的标签值;其中;不同对象的人脸标识对应不同的标签值,不同对象的人体标识对应不同的标签值;将所述相同的标签值对应的人脸标识和人体标识进行关联,得到所述关联结果。
- 根据权利要求1至5任一项所述的方法,其中,所述对象信息包括:人脸特征信息和/或人脸属性信息;所述检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息,包括:检测所述图像帧序列中参与所述行为事件的所述至少一个对象的人脸特征信息;基于所述至少一个对象的人脸特征信息,确定所述至少一个对象的人脸属性信息;所述方法还包括:基于所述至少一个对象的所述人脸特征信息和/或所述人脸属性信息,确定所述至少一个对象的身份信息。
- 根据权利要求1至5任一项所述的方法,其中,所述对象信息包括:人体特征信息和/或人体属性信息;所述检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息,包括:检测所述图像帧序列中参与所述行为事件的所述至少一个对象的人体特征信息;基于所述至少一个对象的人体特征信息,确定所述至少一个对象的人体属性信息;所述方法还包括:基于所述至少一个对象的所述人体特征信息和/或所述人体属性信息,确定所述至少一个对象的身份信息。
- 根据权利要求1至7任一项所述的方法,其中,所述方法还包括:在所述行为事件为预设事件的情况下,输出告警信息;所述告警信息包括以下在所述图像帧序列中展示的至少之一:发生所述行为事件的事件区域、所述行为事件的时空信息、参与所述行为事件对象的人脸框、参与所述行为事件对象的人体框、所述人脸框中的人脸图像、所述人体框中的人体图像、参与所述行为事件的对象的人脸和/或人体的属性信息、参与所述行为事件的对象关联家属的属性信息。
- 一种行为识别装置,所述装置包括:获取部分,配置为获取图像帧序列;跟踪部分,配置为对所述图像帧序列中检测到的至少一个对象进行跟踪,得到每个对象的跟踪信息;识别部分,配置为基于所述每个对象的跟踪信息,对所述至少一个对象进行行为识别,得到所述图像帧序列中行为事件的标识;检测部分,配置为检测所述图像帧序列中参与所述行为事件的至少一个对象的对象信息;关联部分,配置为将所述行为事件的标识和所述至少一个对象的对象信息关联,得到行为识别结果。
- 根据权利要求9所述的行为识别装置,其中,所述检测部分,还配置为:确定所述图像帧序列的每一图像中的存在所述行为事件的事件区域;从所述每一图像中的存在所述行为事件的事件区域中,确定所述至少一个对象的对象信息。
- 根据权利要求9或10所述的行为识别装置,其中,所述至少一个对象的对象信息包括所述至少一个对象的人脸标识和人体标识;所述关联部分,还配置为:将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果;将所述行为事件的标识和所述关联结果进行关联,得到所述行为识别结果。
- 根据权利要求11所述的行为识别装置,其中,所述关联部分,还配置为:在所述图像帧序列的每一图像中,确定参与所述行为事件的每一人脸与每一人体之间的位置关系;基于所述位置关系,将所述至少一个对象的人脸标识和人体标识中,属于同一对象的人脸标识和人体标识进行关联,得到关联结果。
- 根据权利要求11或12所述的行为识别装置,其中,所述关联部分,还配置为:从所述至少一个对象的人脸标识和人体标识中,获取属于同一对象的人脸标识和人体标识;向属于同一对象的人脸标识和人体标识对应相同的标签值;其中;不同对象的人脸标识对应不同的标签值,不同对象的人体标识对应不同的标签值;将所述相同的标签值对应的人脸标识和人体标识进行关联,得到所述关联结果。
- 根据权利要求9或13所述的行为识别装置,其中,所述对象信息包括:人脸特征信息和/或人脸属性信息;所述检测部分,还配置为:检测所述图像帧序列中参与所述行为事件的所述至少一个对象的人脸特征信息;基于所述至少一个对象的人脸特征信息,确定所述至少一个对象的人脸属性信息;所述行为识别装置还包括确定部分;所述确定部分,配置为基于所述至少一个对象的所述人脸特征信息和/或所述人脸属性信息,确定所述至少一个对象的身份信息。
- 根据权利要求9或13所述的行为识别装置,其中,所述对象信息包括:人体特 征信息和/或人体属性信息,所述检测部分,还配置为:检测所述图像帧序列中参与所述行为事件的所述至少一个对象的人体特征信息;基于所述至少一个对象的人体特征信息,确定所述至少一个对象的人体属性信息;所述行为识别装置还包括确定部分;所述确定部分,配置为基于所述至少一个对象的所述人体特征信息和/或所述人体属性信息,确定所述至少一个对象的身份信息。
- 根据权利要求9或15所述的行为识别装置,其中,所述行为识别装置还包括输出部分;所述输出部分,配置为在所述行为事件为预设事件的情况下,输出告警信息;所述告警信息包括以下在所述图像帧序列中展示的至少之一:发生所述行为事件的事件区域、所述行为事件的时空信息、参与所述行为事件对象的人脸框、参与所述行为事件对象的人体框、所述人脸框中的人脸图像、所述人体框中的人体图像、参与所述行为事件的对象的人脸和/或人体的属性信息、参与所述行为事件的对象关联家属的属性信息。
- 一种电子设备,包括:存储器和处理器,所述存储器存储有可在所述处理器上运行的计算机程序,所述处理器执行所述计算机程序时实现权利要求1至8任一项所述方法中的步骤。
- 一种计算机存储介质,所述计算机存储介质存储有一个或者多个程序,所述一个或者多个程序可被一个或者多个处理器执行,以实现权利要求1至8任一项所述方法中的步骤。
- 一种芯片,包括:处理器,用于从存储器中调用并运行计算机程序,使得安装有所述芯片的设备执行如权利要求1至8任一项所述方法中的步骤。
- 一种计算机程序产品,所述计算机程序产品承载有程序代码,所述程序代码包括的指令可配置为执行如权利要求1至8任一所述方法中的步骤。
- 一种计算机程序,包括计算机可读代码,在所述计算机可读代码在电子设备中运行的情况下,所述电子设备中的处理器执行如权利要求1至8任一所述方法中的步骤。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202111109050.X | 2021-09-22 | ||
| CN202111109050.XA CN113837066A (zh) | 2021-09-22 | 2021-09-22 | 行为识别方法、装置、电子设备及计算机存储介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2023045239A1 true WO2023045239A1 (zh) | 2023-03-30 |
Family
ID=78960367
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2022/077461 Ceased WO2023045239A1 (zh) | 2021-09-22 | 2022-02-23 | 行为识别方法、装置、设备、介质、芯片、产品及程序 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN113837066A (zh) |
| WO (1) | WO2023045239A1 (zh) |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN113837066A (zh) * | 2021-09-22 | 2021-12-24 | 深圳市商汤科技有限公司 | 行为识别方法、装置、电子设备及计算机存储介质 |
| CN114913602A (zh) * | 2022-05-20 | 2022-08-16 | 深圳市商汤科技有限公司 | 行为识别方法、装置、设备及存储介质 |
| CN116049481A (zh) * | 2023-01-30 | 2023-05-02 | 蓝思系统集成有限公司 | 视频存储的方法、装置、存储介质及电子设备 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111860430A (zh) * | 2020-07-30 | 2020-10-30 | 浙江大华技术股份有限公司 | 打架行为的识别方法和装置、存储介质及电子装置 |
| CN111985428A (zh) * | 2020-08-27 | 2020-11-24 | 上海商汤智能科技有限公司 | 一种安全检测方法、装置、电子设备及存储介质 |
| WO2021002722A1 (ko) * | 2019-07-04 | 2021-01-07 | (주)넷비젼텔레콤 | 이벤트 태깅 기반 상황인지 방법 및 그 시스템 |
| CN113223046A (zh) * | 2020-07-10 | 2021-08-06 | 浙江大华技术股份有限公司 | 一种监狱人员行为识别的方法和系统 |
| CN113837066A (zh) * | 2021-09-22 | 2021-12-24 | 深圳市商汤科技有限公司 | 行为识别方法、装置、电子设备及计算机存储介质 |
Family Cites Families (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN106845385A (zh) * | 2017-01-17 | 2017-06-13 | 腾讯科技(上海)有限公司 | 视频目标跟踪的方法和装置 |
| CN113111839A (zh) * | 2021-04-25 | 2021-07-13 | 上海商汤智能科技有限公司 | 行为识别方法及装置、设备和存储介质 |
-
2021
- 2021-09-22 CN CN202111109050.XA patent/CN113837066A/zh active Pending
-
2022
- 2022-02-23 WO PCT/CN2022/077461 patent/WO2023045239A1/zh not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2021002722A1 (ko) * | 2019-07-04 | 2021-01-07 | (주)넷비젼텔레콤 | 이벤트 태깅 기반 상황인지 방법 및 그 시스템 |
| CN113223046A (zh) * | 2020-07-10 | 2021-08-06 | 浙江大华技术股份有限公司 | 一种监狱人员行为识别的方法和系统 |
| CN111860430A (zh) * | 2020-07-30 | 2020-10-30 | 浙江大华技术股份有限公司 | 打架行为的识别方法和装置、存储介质及电子装置 |
| CN111985428A (zh) * | 2020-08-27 | 2020-11-24 | 上海商汤智能科技有限公司 | 一种安全检测方法、装置、电子设备及存储介质 |
| CN113837066A (zh) * | 2021-09-22 | 2021-12-24 | 深圳市商汤科技有限公司 | 行为识别方法、装置、电子设备及计算机存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN113837066A (zh) | 2021-12-24 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Shorfuzzaman et al. | Towards the sustainable development of smart cities through mass video surveillance: A response to the COVID-19 pandemic | |
| Bertoni et al. | Perceiving humans: from monocular 3d localization to social distancing | |
| CN111383421A (zh) | 隐私保护跌倒检测方法和系统 | |
| Rafiq et al. | Wearable sensors-based human locomotion and indoor localization with smartphone | |
| CN104317918B (zh) | 基于复合大数据gis的异常行为分析及报警系统 | |
| US11288954B2 (en) | Tracking and alerting traffic management system using IoT for smart city | |
| CN119314117B (zh) | 多模态大模型的处理方法、设备、存储介质和程序产品 | |
| Huang et al. | Enhancing multi-camera people tracking with anchor-guided clustering and spatio-temporal consistency id re-assignment | |
| Javed et al. | Face mask detection and social distance monitoring system for COVID-19 pandemic | |
| CN113837066A (zh) | 行为识别方法、装置、电子设备及计算机存储介质 | |
| An et al. | VFP290k: A large-scale benchmark dataset for vision-based fallen person detection | |
| WO2023273132A1 (zh) | 行为检测方法、装置、计算机设备、存储介质和程序 | |
| Zh Satybaldina et al. | Development of an algorithm for abnormal human behavior detection in intelligent video surveillance system | |
| Dhiraj et al. | Activity recognition for indoor fall detection in 360-degree videos using deep learning techniques | |
| Fabbri et al. | Inter-homines: Distance-based risk estimation for human safety | |
| JP2023098484A (ja) | 情報処理プログラム、情報処理方法および情報処理装置 | |
| JP2023098482A (ja) | 情報処理プログラム、情報処理方法および情報処理装置 | |
| CN113836993A (zh) | 定位识别方法、装置、设备及计算机可读存储介质 | |
| HK40064036A (zh) | 行为识别方法、装置、电子设备及计算机存储介质 | |
| WO2023124451A1 (zh) | 用于生成告警事件的方法、装置、设备和存储介质 | |
| Chen | A video surveillance system designed to detect multiple falls | |
| Yadav et al. | An intelligent system to detect violent mob activities | |
| Arivazhagan | Versatile loitering detection based on non-verbal cues using dense trajectory descriptors | |
| Sukesh Babu et al. | Robust Pedestrian Detection via Enriched Dataset | |
| Kumari et al. | Deep learning and computer vision-based social distancing detection system |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 22871305 Country of ref document: EP Kind code of ref document: A1 |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 25.09.2024) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 22871305 Country of ref document: EP Kind code of ref document: A1 |