WO2025112052A1 - 多动物自由社交行为映射分类方法、装置、电子设备及存储介质 - Google Patents

多动物自由社交行为映射分类方法、装置、电子设备及存储介质 Download PDF

Info

Publication number
WO2025112052A1
WO2025112052A1 PCT/CN2023/135942 CN2023135942W WO2025112052A1 WO 2025112052 A1 WO2025112052 A1 WO 2025112052A1 CN 2023135942 W CN2023135942 W CN 2023135942W WO 2025112052 A1 WO2025112052 A1 WO 2025112052A1
Authority
WO
WIPO (PCT)
Prior art keywords
dimensional
behavior
sequence
target
low
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2023/135942
Other languages
English (en)
French (fr)
Inventor
韩亚宁
陈可
蔚鹏飞
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Shenzhen Institute of Advanced Technology of CAS
Original Assignee
Shenzhen Institute of Advanced Technology of CAS
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Shenzhen Institute of Advanced Technology of CAS filed Critical Shenzhen Institute of Advanced Technology of CAS
Priority to PCT/CN2023/135942 priority Critical patent/WO2025112052A1/zh
Publication of WO2025112052A1 publication Critical patent/WO2025112052A1/zh
Anticipated expiration legal-status Critical
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V40/00Recognition of biometric, human-related or animal-related patterns in image or video data
    • G06V40/10Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands

Definitions

  • the present application relates to the field of computer technology, and in particular to a method, device, electronic device and storage medium for mapping and classifying free social behaviors of multiple animals.
  • the behavioral composition of animals is dynamic, hierarchical, high-dimensional and parallel. Therefore, it is difficult to determine the starting and ending time of a behavior when manually identifying behavior. The starting and ending points of behaviors observed by different experimenters are often different, resulting in inconsistent behavioral identification.
  • a behavior consists of the movement of multiple body parts of an animal. It is difficult to set enough categories and parameters in advance to precisely define each behavior, resulting in an incomplete definition of behavior.
  • experimenters need to watch the behavior of animals repeatedly. It takes about ten hours to identify a one-hour behavioral video of a single animal.
  • the identification of social behavior of multiple animals will be more complicated and time-consuming.
  • the identification of manual participation in behavioral classification is inconsistent, the definition is incomplete, and the classification efficiency is low. This situation needs to be further improved.
  • the present application provides a method, device, electronic device and storage medium for mapping and classifying free social behaviors of multiple animals, which can solve the problems of inconsistent identification, incomplete definition and low classification efficiency of manual participation behavior classification in related technologies.
  • the technical solution is as follows:
  • a method for mapping and classifying free social behaviors of multiple animals includes: obtaining corresponding body posture data based on videos of free social interactions of multiple animals; obtaining a corresponding sequence set based on the body posture data, wherein the sequence set includes a motion sequence, an action sequence, and a distance sequence; obtaining a two-dimensional time series and a corresponding target time segmentation point corresponding to each sequence in the sequence set based on each sequence in the sequence set; calling a pre-set popular feature model; inputting a manifold feature model based on all time points corresponding to all two-dimensional time series to obtain a corresponding two-dimensional manifold representation; determining a low-dimensional behavior space based on the target time segmentation point and the two-dimensional manifold representation; inputting the coordinates of the low-dimensional behavior space into a Gaussian kernel function to obtain the point density of the low-dimensional behavior space; determining the stability of the point density of the low-dimensional behavior space based on the stability of the point density of the low
  • a multi-animal free social behavior mapping and classification device includes but is not limited to:
  • Posture data acquisition module which is used to obtain corresponding body posture data based on videos of multiple animals freely socializing
  • a sequence set acquisition module used to acquire a corresponding sequence set based on the body posture data, wherein the sequence set includes a motion sequence, an action sequence and a distance sequence;
  • a time acquisition module is used to acquire the two-dimensional time series and the corresponding target time segmentation point corresponding to each sequence in the sequence set;
  • a feature model retrieval module is used to retrieve a preset popular feature model
  • a two-dimensional manifold representation acquisition module is used to acquire the corresponding two-dimensional manifold representation based on the input manifold feature model of all time points corresponding to all two-dimensional time series;
  • the behavior space determination module is used to determine the low-dimensional behavior space based on the target time segmentation points and the two-dimensional manifold representation;
  • a spatial point density acquisition module inputs the coordinates of the low-dimensional behavior space into a Gaussian kernel function to obtain the point density of the low-dimensional behavior space;
  • the cluster number model determination module is used to determine the target cluster number model based on the stability of the point density in the low-dimensional behavior space;
  • the original video acquisition module is used to obtain the original videos of social behaviors of multiple animals
  • the target behavior category acquisition module inputs the original social behavior video into the target clustering model to obtain the corresponding target behavior category.
  • a two-dimensional representation acquisition module for acquiring a two-dimensional representation corresponding to each sequence based on each sequence in the sequence set
  • the discrete time segment acquisition module decomposes the two-dimensional representation corresponding to each sequence using a dynamic time alignment kernelization algorithm to obtain the discrete time segment corresponding to each sequence;
  • a time segmentation point determination module is used to determine the corresponding target time segmentation point based on discrete time segments
  • the target time segmentation point determination module combines all the time segmentation points to determine the target time segmentation point.
  • a two-dimensional time series acquisition module is used to acquire a two-dimensional time series corresponding to each sequence
  • a six-dimensional time series determination module is used to determine the corresponding six-dimensional time series based on the two-dimensional time series of all sequences, wherein the six-dimensional time series is the two-dimensional time series corresponding to the motion sequence, the two-dimensional time series corresponding to the action sequence, and the two-dimensional time series corresponding to the distance sequence;
  • a training sequence acquisition module used for acquiring a training six-dimensional time series based on the six-dimensional time series
  • the manifold feature model determination module obtains the corresponding training two-dimensional manifold representation based on the UMAP algorithm operation on the training six-dimensional time series, and is used to determine the manifold feature model between the training six-dimensional time series and the training two-dimensional manifold representation.
  • a matrix construction module for constructing a similarity matrix based on the two-dimensional manifold representation and the time segmentation points
  • the matrix representation data acquisition module uses the UMAP algorithm to obtain the matrix representation data of the similarity matrix in two-dimensional space
  • a low-dimensional behavior space establishment module is used to establish a low-dimensional behavior space based on matrix representation data.
  • the basis vector extraction module randomly extracts a part of the time segments as basis vectors from all the time segments;
  • a calling module used to call the two-dimensional manifold representation and the time segmentation point
  • the matrix construction module measures the similarity of all time series segments and basis vectors using a dynamic time alignment kernelization algorithm based on the two-dimensional manifold representation and time segmentation points, so as to construct a similarity matrix.
  • a function set acquisition module which acquires a corresponding Gaussian kernel function set based on multiple variance parameters
  • a behavior space point density set acquisition module inputs the coordinates of the low-dimensional behavior space into a Gaussian kernel function in a Gaussian kernel function set to obtain the low-dimensional behavior space point density output by each Gaussian kernel function;
  • the maximum stable value determination module is used to determine the maximum stable value and the minimum stable value corresponding to the variance parameter based on the stability of the point density of all low-dimensional behavior spaces;
  • the target cluster number model determination module defines the maximum stable value and the minimum stable value as the upper bound and the lower bound of the cluster number function, and uses them to determine the target cluster number model.
  • An input module used to input the original video of social behavior into the target clustering model
  • the segmentation module is used to segment the original social behavior video based on the upper and lower bounds of the clustering number function
  • a segment set acquisition module is used to acquire a corresponding target video segment set
  • the saving module is used to save each target video segment in the target video segment set into a folder named after the behavior category.
  • an electronic device includes at least one processor and at least one memory, wherein the memory stores computer-readable instructions; the computer-readable instructions are executed by one or more of the processors, so that the electronic device implements the multi-animal free social behavior mapping and classification method as described above.
  • a storage medium stores computer-readable instructions thereon, and the computer-readable instructions are executed by one or more processors to implement the multi-animal free social behavior mapping and classification method as described above.
  • a computer program product includes computer-readable instructions, which are stored in a storage medium.
  • One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the multi-animal free social behavior mapping and classification method as described above.
  • an unsupervised behavior mapping and classification framework is designed according to the natural structure of social behavior, so that unknown social behaviors can be fully covered.
  • the possible behavior categories are first consistently separated by an algorithm and then manually defined, which not only effectively covers the behavior categories that are difficult to define artificially, but also obtains consistent results, and can classify the social behaviors of multiple animals, thereby improving the efficiency of classifying animal social behaviors.
  • FIG1 is a schematic diagram of an implementation environment involved in the present application.
  • FIG. 2 is a diagram showing a classification method for mapping free social behaviors of multiple animals according to an exemplary embodiment. Flowchart of the method
  • FIG3 is a flowchart of S121 to S123 in another multi-animal free social behavior mapping classification method according to an exemplary embodiment
  • FIG4 is a flow chart of steps S131 to S134 in another method for mapping and classifying free social behaviors of multiple animals according to an exemplary embodiment
  • FIG5 is a flowchart of S151 to S154 in another multi-animal free social behavior mapping classification method according to an exemplary embodiment
  • FIG6 is a flowchart of S171 to S174 in another multi-animal free social behavior mapping classification method according to an exemplary embodiment
  • FIG. 7 is a flow chart of steps S181 to S184 in another method for mapping and classifying free social behaviors of multiple animals according to an exemplary embodiment
  • FIG8 is a structural block diagram of a multi-animal free social behavior mapping and classification device according to an exemplary embodiment
  • Fig. 9 is a structural block diagram of an electronic device according to an exemplary embodiment.
  • Fig. 1 is a schematic diagram of an implementation environment involved in a multi-animal free social behavior mapping and classification method.
  • the implementation environment includes a terminal, a server, and a service system configured with a member association database.
  • the terminal can be operated by a client that provides video classification, and can be an electronic device such as a desktop computer, a laptop computer, a tablet computer, a smart phone, etc., which is not limited here.
  • the client provides a video classification function, for example, a media player, a browser, etc., which can be in the form of an application or a web page.
  • the user interface of the client for playing the video can be in the form of a program window or a web page, which is not limited here.
  • the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
  • This server is an electronic device used to provide background services.
  • the server provides cloud storage services for audio and video data for the terminal.
  • the server establishes a communication connection with the service system in advance through wired or wireless means, and realizes linkage with the service system through the communication connection.
  • the service system can be a single server or a server cluster composed of multiple servers.
  • the client running on the terminal will initiate a resource use invitation to the server, requesting the server to determine the resource location and resource usage time through resource allocation, and then issue the invitation.
  • the resource allocation process is executed for the client where the inviter is located through the resource usage invitation linkage service system, and the invitation result indicating the location of the resource and the resource usage time is returned to the client where the inviter is located, so that the client where the inviter is located can further confirm whether to issue a resource usage invitation to the client where the invitee is located.
  • server and the service system may also be integrated into the same server cluster so that resource allocation is completed by the same server cluster.
  • the present application embodiment provides a method for mapping and classifying free social behaviors of multiple animals, including:
  • a social behavior shooting module is used to shoot videos of multiple animals socializing freely, and then a deep learning posture estimation module is used to estimate the body postures of the multiple animals socializing freely.
  • the multi-animal social posture estimation in the DeepLabCut tool can be used to acquire the body posture.
  • some key video frames in the video of multiple animals socializing freely are extracted to manually annotate the body postures of multiple animals.
  • the predefined number of body points is also different. For example, mice use 16-point annotation, birds use 21-point annotation, and dogs use 17-point annotation.
  • the annotated frames are used to train the deep neural network.
  • the network training converges, the network is used to estimate the body posture of all animals in the video for subsequent processing.
  • the sequence set includes motion sequence, action sequence and distance sequence.
  • the motion sequence represents the instantaneous movement speed of each body point of the social animal, and the calculation method is the absolute value of the difference between adjacent frames multiplied by the frame rate;
  • the action sequence represents the movement of the limbs of the social animal, and the calculation method is the coordinate value of each body point minus the coordinate value of the body center point, and each animal is aligned to the Cartesian coordinate origin;
  • the distance sequence represents the position change between the limbs of the social animal during the social process, and the calculation method is the Euclidean distance between the corresponding body points of each animal and other animals.
  • the UMAP algorithm can be used to decompose the two-dimensional time series of each sequence, and then obtain the corresponding two-dimensional time series; please refer to Figure 3, the specific process of obtaining the target time segmentation point is as follows:
  • three sequences are included, namely, motion sequence, action sequence and distance sequence, so for the two-dimensional representation, there are three sequences of two-dimensional representation.
  • each discrete time segment can use the dynamic time alignment kernelization algorithm to calculate the corresponding time segmentation point, and the required target time segmentation point is determined from all the time segmentation points for subsequent use.
  • the six-dimensional time series is the two-dimensional time series corresponding to the motion sequence, the two-dimensional time series corresponding to the action sequence, and the two-dimensional time series corresponding to the distance sequence.
  • the six-dimensional time series can be determined through the two-dimensional time series corresponding to the three sequences.
  • the mapping relationship between the training six-dimensional time series and the training two-dimensional manifold representation can be learned.
  • the input is the time point corresponding to the six-dimensional time series.
  • the time segment here is the discrete time segment mentioned above, and for the selection of basis vectors, a part can be randomly selected to determine.
  • a dynamic time alignment kernelization algorithm is used to measure the similarity of all time series segments and basis vectors, and a similarity matrix is constructed.
  • the UMAP algorithm can be used to obtain matrix representation data of the similarity matrix in two-dimensional space.
  • the execution steps from S171 to S173 are to determine the Gaussian kernel function, and whether the density of the low-dimensional behavior space output by the Gaussian kernel is stable is used as the judgment criterion, wherein the variance parameter of the Gaussian kernel is used as input, and the density of the low-dimensional behavior space calculated by the Gaussian kernel is output, and two variance parameters that can make the density of the low-dimensional behavior space stable are selected as boundaries, wherein the large variance parameter is used as the upper bound and the small variance parameter is used as the lower bound, and then the variances corresponding to the lower bound and the upper bound are used as the minimum stable value and the maximum stable value of the clustering number model, thereby determining the clustering number model.
  • the difference from the prior art is that it is not only for a single animal, but a deep neural network after training convergence can be used to mark multiple animals, and no manual participation in marking is required, which can improve the marking efficiency by reducing the participation of human factors.
  • the action data of the animal is calculated from sequences such as motion sequence, action sequence and distance sequence.
  • a stable two-dimensional manifold representation can be output.
  • a low-dimensional space model can be established in combination with the target time segmentation point, and a suitable Gaussian function is used to calculate the density of the low-dimensional behavior space, and the variance parameter is selected based on the density stability of the low-dimensional behavior space point.
  • the corresponding target clustering number model can be determined using the clustering number model.
  • the videos of the social behaviors of multiple animals are separated by the target clustering number model, and the classification of animal behaviors can be realized.
  • the target clustering number model distinguishes the different social behaviors of animals, the social behaviors distinguished by the target clustering number model can be named manually later.
  • the specific behaviors include:
  • the original video of social behavior input into the target clustering model can be a video of one animal or multiple animals.
  • different animals in the video can be labeled.
  • the upper bound of the clustering number model is used to distinguish the global differences in animal social behaviors, while the lower bound of the clustering number model is used to subdivide the fine differences in animal social behaviors.
  • a target video segment set distinguished by the clustering number model can be obtained, and each behavior corresponds to a target video segment set.
  • each social behavior of the animal can be named manually, and after the original video of the social behavior is input, the system can save each target video clip in a corresponding file according to the classification of the target clustering model.
  • the following is an embodiment of the device of the present application, which can be used to execute the multi-animal free social behavior mapping classification method involved in the present application.
  • the method embodiment of the multi-animal free social behavior mapping classification method involved in the present application please refer to the method embodiment of the multi-animal free social behavior mapping classification method involved in the present application.
  • a multi-animal free social behavior mapping and classification device including but not limited to:
  • a posture data acquisition module 200 which is used to acquire corresponding body posture data based on a video of multiple animals freely socializing;
  • a sequence set acquisition module 210 for acquiring a corresponding sequence set based on the body posture data, wherein the sequence set includes a motion sequence, an action sequence, and a distance sequence;
  • a time acquisition module 220 based on each sequence in the sequence set, is used to acquire a two-dimensional time series corresponding to each sequence and a corresponding target time segmentation point;
  • the feature model retrieving module 230 is used to retrieve a preset popular feature model
  • a two-dimensional manifold representation acquisition module 240 is used to acquire a corresponding two-dimensional manifold representation based on the input manifold feature model of all time points corresponding to all two-dimensional time series;
  • a behavior space determination module 250 is used to determine a low-dimensional behavior space based on the target time segmentation points and the two-dimensional manifold representation;
  • a space point density acquisition module 260 inputs the coordinates of the low-dimensional behavior space into a Gaussian kernel function to acquire the point density of the low-dimensional behavior space;
  • a cluster number model determination module 270 is used to determine a target cluster number model based on the stability of the low-dimensional behavior space point density
  • the original video acquisition module 280 is used to acquire the original videos of social behaviors of multiple animals;
  • the target behavior category acquisition module 290 inputs the original social behavior video into the target clustering model to obtain the corresponding target behavior category.
  • a two-dimensional representation acquisition module 300 based on each sequence in the sequence set, is used to acquire a two-dimensional representation corresponding to each sequence;
  • a discrete time segment acquisition module 310 decomposes the two-dimensional representation corresponding to each sequence using a dynamic time alignment kernelization algorithm to obtain a discrete time segment corresponding to each sequence;
  • a time segmentation point determination module 320 is used to determine corresponding target time segmentation points based on discrete time segments
  • the target time division point determination module 330 performs a merging operation on all the time division points to determine the target time division point.
  • a two-dimensional time series acquisition module 400 is used to acquire a two-dimensional time series corresponding to each sequence
  • a six-dimensional time series determination module 410 is used to determine a corresponding six-dimensional time series based on the two-dimensional time series of all sequences, wherein the six-dimensional time series is a two-dimensional time series corresponding to the motion sequence, a two-dimensional time series corresponding to the action sequence, and a two-dimensional time series corresponding to the distance sequence;
  • a training sequence acquisition module 420 used for acquiring a training six-dimensional time series based on the six-dimensional time series;
  • the manifold feature model determination module 430 obtains the corresponding training two-dimensional manifold representation based on the UMAP algorithm operation on the training six-dimensional time series, and is used to determine the manifold feature model between the training six-dimensional time series and the training two-dimensional manifold representation.
  • a matrix construction module 500 for constructing a similarity matrix based on the two-dimensional manifold representation and the time segmentation points;
  • a matrix representation data acquisition module 510 uses a UMAP algorithm to acquire matrix representation data of a similarity matrix in a two-dimensional space;
  • the low-dimensional behavior space establishment module 520 is used to establish a low-dimensional behavior space based on the matrix representation data.
  • a basis vector extraction module 600 is used to randomly extract a portion of the time segments as basis vectors from all the time segments;
  • a retrieval module 610 used to retrieve a two-dimensional manifold representation and a time segmentation point
  • the matrix construction module 620 measures the similarity of all time series segments and basis vectors using a dynamic time alignment kernelization algorithm based on the two-dimensional manifold representation and time segmentation points to construct a similarity matrix.
  • a function set acquisition module 700 which acquires a corresponding Gaussian kernel function set based on a plurality of variance parameters
  • the behavior space point density set acquisition module 710 inputs the coordinates of the low-dimensional behavior space into the Gaussian kernel function in the Gaussian kernel function set to obtain the low-dimensional behavior space point density output by each Gaussian kernel function;
  • the maximum stable value determination module 720 is used to determine the maximum stable value and the minimum stable value corresponding to the variance parameter based on the stability of all low-dimensional behavior space point densities;
  • the target cluster number model determination module 730 defines the maximum stable value and the minimum stable value as the upper bound and the lower bound of the cluster number function, and uses them to determine the target cluster number model.
  • An input module 800 is used to input the original video of social behavior into the target clustering model
  • a segmentation module 810 is used to segment the original social behavior video based on the upper bound and the lower bound of the cluster number function
  • the segment set acquisition module 820 is used to acquire the corresponding target video segment set
  • the saving module 830 is used to save each target video segment in the target video segment set into a folder named after the behavior category.
  • the multi-animal free social behavior mapping classification device provided in the above embodiment only uses the division of the above-mentioned functional modules as an example when performing multi-animal free social behavior mapping classification.
  • the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the multi-animal free social behavior mapping classification device will be divided into different functional modules to complete all or part of the functions described above.
  • multi-animal free social behavior mapping classification device and the multi-animal free social behavior mapping classification method provided in the above embodiment belong to the same concept, wherein each module performs operation
  • each module performs operation
  • the specific method has been described in detail in the method embodiment and will not be repeated here.
  • An electronic device 4000 is provided in an embodiment of the present application.
  • the electronic device 4000 may include: a desktop computer, a laptop computer, a server, etc.
  • the electronic device 4000 includes at least one processor 4001 and at least one memory 4003 .
  • the data interaction between the processor 4001 and the memory 4003 can be realized through at least one communication bus 4002.
  • the communication bus 4002 may include a path for transmitting data between the processor 4001 and the memory 4003.
  • the communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc.
  • the communication bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in FIG9, but it does not mean that there is only one bus or one type of bus.
  • the electronic device 4000 may further include a transceiver 4004, which may be used for data interaction between the electronic device and other electronic devices, such as data transmission and/or data reception.
  • a transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.
  • Processor 4001 can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 4001 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
  • the memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser
  • the present invention may be a computer program product (such as a 32-bit disk, optical disc, digital versatile disc, Blu-ray disc, etc.), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program instructions or code in the form of instructions or data structures and can be accessed by the electronic device 400, but is not limited to this.
  • Computer-readable instructions are stored in the memory 4003 , and the processor 4001 can read the computer-readable instructions stored in the memory 4003 through the communication bus 4002 .
  • the computer-readable instructions are executed by one or more processors 4001 to implement the multi-animal free social behavior mapping and classification method in the above-mentioned embodiments.
  • a storage medium is provided in an embodiment of the present application, on which computer-readable instructions are stored, and the computer-readable instructions are executed by one or more processors to implement the multi-animal free social behavior mapping and classification method as described above.
  • a computer program product is provided in an embodiment of the present application.
  • the computer program product includes computer-readable instructions, which are stored in a storage medium.
  • One or more processors of an electronic device read the computer-readable instructions from the storage medium, load and execute the computer-readable instructions, so that the electronic device implements the multi-animal free social behavior mapping and classification method as described above.
  • a social behavior shooting module it is first necessary to use a social behavior shooting module to shoot videos of multiple animals socializing freely, so as to obtain corresponding body posture data.
  • no manual labeling is required, and the training and convergence deep neural network is directly used for processing, thereby reducing the involvement of human factors; then, the corresponding motion sequence, action sequence and distance sequence can be obtained based on the body posture data.
  • the corresponding motion sequence, action sequence and distance sequence can be obtained based on the body posture data.
  • the pre-trained manifold feature model is called up.
  • all time points corresponding to all two-dimensional time series are input into the manifold feature model, and the two-dimensional manifold representation corresponding to all time points can be obtained.
  • a part of all time segments is randomly extracted as basis vectors, and the dynamic time alignment kernelization algorithm is used to calculate the similarity between the time series segments and the basis vectors, so as to construct a similarity matrix.
  • the similarity matrix is obtained using the UMAP algorithm to obtain the matrix representation data of the similarity matrix in two-dimensional space. After obtaining the matrix representation data, a low-dimensional behavior space can be established.
  • the coordinates of the low-dimensional behavior space can be obtained, and the coordinate information can be input into the Gaussian kernel function.
  • the density of the low-dimensional behavior space output by the Gaussian kernel is stable as the criterion, wherein the variance parameter of the Gaussian kernel is used as input, and the Gaussian kernel calculates the density of the low-dimensional behavior space as output.
  • Two variance parameters that can stabilize the density of the low-dimensional behavior space are selected as boundaries, wherein the large variance parameter is used as the upper bound, and the small variance parameter is used as the lower bound, and then the lower bound is compared with The variance corresponding to the upper bound is used as the minimum stable value and the maximum stable value of the cluster number model, thereby determining the target cluster number model.
  • the original videos of social behaviors of multiple animals are input into the target clustering number model.
  • a set of target video clips distinguished by the target clustering number model can be obtained, and each target video clip is saved in a corresponding file, thereby realizing the classification of social behaviors of multiple animals.
  • the social behaviors of multiple animals can be classified, and in the process of classifying the original videos of social behaviors, while ensuring the classification accuracy, the problem of inconsistent social behavior classification can be solved through the animal's body posture data and decomposition.
  • the manifold features and low-dimensional space mapping classification solve the problem of incomplete definition of social behavior.
  • the strategy of first classification and then definition of unsupervised social behavior classification solves the problem of low efficiency of social behavior classification. Without the need for human participation, the efficiency of subdivision and classification of social behaviors of multiple animals can be improved.

Landscapes

  • Engineering & Computer Science (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Theoretical Computer Science (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

本申请提供了一种多动物自由社交行为映射分类方法、装置、电子设备及存储介质,涉及社交行为分类领域。其中,该方法包括:基于多动物自由社交的视频获取对应的身体姿态数据;基于身体姿态数据获取对应的序列集合;获取每个序列对应的二维时间序列以及对应的目标时间分割点;调取流行特征模型;获取对应的二维流形表征;基于目标时间分割点以及二维流形表征确定低维行为空间;获取低维行为空间点密度;基于低维行为空间点密度的稳定性确定目标聚类数模型;获取多动物的社交行为原始视频;将社交行为原始视频输入目标聚类数模型获取对应的目标行为类别。本申请解决了相关技术中人工参与行为分类的鉴定不一致、定义不全面和分类效率低的问题。

Description

多动物自由社交行为映射分类方法、装置、电子设备及存储介质 技术领域
本申请涉及计算机技术领域,具体而言,本申请涉及一种多动物自由社交行为映射分类方法、装置、电子设备及存储介质。
背景技术
动物的行为组成具有动态、层次、高维和并行的特点,因此,人工鉴定行为首先很难判断一个行为的时间起点和终点,不同的实验者观测的行为的起点和终点往往不同,造成行为鉴定不一致的问题。
其次,一个行为由动物的多个身体部位的运动组成,难以提前设置足够的类别和参数精细的定义每一种行为,造成行为定义不全面的问题。为了提高人工鉴定的精度,实验者需要反复观看动物的行为,鉴定一个小时的单只动物的行为学视频需要约十个小时的时间,而对于多动物的社交行为鉴定的情况将更加复杂更加耗时,人工参与行为分类的鉴定不一致、定义不全面和分类效率低,对此情况有待进一步改善。
发明内容
本申请各提供了一种多动物自由社交行为映射分类方法、装置、电子设备及存储介质,可以解决相关技术中存在的人工参与行为分类的鉴定不一致、定义不全面和分类效率低的问题。所述技术方案如下:
根据本申请的一个方面,一种多动物自由社交行为映射分类方法,包括:基于多动物自由社交的视频获取对应的身体姿态数据;基于身体姿态数据获取对应的序列集合,其中,序列集合包括运动序列、动作序列以及距离序列;基于序列集合中的每个序列获取每个序列对应的二维时间序列以及对应的目标时间分割点;调取预先设置的流行特征模型;基于全部二维时间序列对应的所有时间点输入流形特征模型,获取对应的二维流形表征;基于目标时间分割点以及二维流形表征确定低维行为空间;将低维行为空间的坐标输入高斯核函数获取低维行为空间点密度;基于低维行为空间点密度的稳定性确定 目标聚类数模型;获取多动物的社交行为原始视频;将社交行为原始视频输入目标聚类数模型获取对应的目标行为类别。
根据本申请的一个方面,一种多动物自由社交行为映射分类装置,包括但不限于:
姿态数据获取模块,基于多动物自由社交的视频用于获取对应的身体姿态数据;
序列集合获取模块,基于身体姿态数据用于获取对应的序列集合,其中,序列集合包括运动序列、动作序列以及距离序列;
时间获取模块,基于序列集合中的每个序列用于获取每个序列对应的二维时间序列以及对应的目标时间分割点;
特征模型调取模块,用于调取预先设置的流行特征模型;
二维流形表征获取模块,基于全部二维时间序列对应的所有时间点输入流形特征模型,用于获取对应的二维流形表征;
行为空间确定模块,基于目标时间分割点以及二维流形表征用于确定低维行为空间;
空间点密度获取模块,将低维行为空间的坐标输入高斯核函数用于获取低维行为空间点密度;
聚类数模型确定模块,基于低维行为空间点密度的稳定性用于确定目标聚类数模型;
原始视频获取模块,用于获取多动物的社交行为原始视频;
目标行为类别获取模块,将社交行为原始视频输入目标聚类数模型用于获取对应的目标行为类别。
在一示例性实施例中,包括但不限于:
二维表征获取模块,基于序列集合中的每个序列用于获取每个序列对应的二维表征;
离散时间片段获取模块,将每个序列对应的二维表征用动态时间对齐核化算法分解用于获取每个序列对应的离散时间片段;
时间分割点确定模块,基于离散时间片段用于确定对应的目标时间分割点;
目标时间分割点确定模块,将全部的时间分割点进行合并运算用于确定目标时间分割点。
在一示例性实施例中,包括但不限于:
二维时间序列获取模块,用于获取每个序列对应的二维表征的二维时间序列;
六维时间序列确定模块,基于全部序列的二维时间序列用于确定对应的六维时间序列,其中,六维时间序列为运动序列对应的二维时间序列、动作序列对应的二维时间序列以及距离序列对应的二维时间序列;
训练序列获取模块,基于所述六维时间序列用于获取训练六维时间序列;
流形特征模型确定模块,基于将训练六维时间序列进行UMAP算法运算获取对应的训练二维流形表征,并用于确定训练六维时间序列与训练二维流形表征之间的流形特征模型。
在一示例性实施例中,包括但不限于:
矩阵构建模块,基于所述二维流形表征与所述时间分割点用于构建相似度矩阵;
矩阵表征数据获取模块,使用UMAP算法用于获取相似度矩阵在二维空间的矩阵表征数据;
低维行为空间建立模块,基于矩阵表征数据用于建立低维行为空间。
在一示例性实施例中,包括但不限于:
基向量抽取模块,从所有的时间片段随机用于抽取一部分片段作为基向量;
调取模块,用于调取二维流形表征和时间分割点;
矩阵构建模块,基于所述二维流形表征和时间分割点使用动态时间对齐核化算法度量所有时间序列片段和基向量的相似度,用于构建相似度矩阵。
在一示例性实施例中,包括但不限于:
函数集合获取模块,基于多个的方差参数获取对应的高斯核函数集合;
行为空间点密度集合获取模块,将低维行为空间的坐标输入到高斯核函数集合中的高斯核函数中用于获取每个高斯核函数输出的低维行为空间点密度;
最值稳定值确定模块,基于全部低维行为空间点密度的稳定性用于确定方差参数对应的最大稳定值与最小稳定值;
目标聚类数模型确定模块,将最大稳定值与最小稳定值定义为聚类数函数的上界与下界,并用于确定目标聚类数模型。
在一示例性实施例中,包括但不限于:
输入模块,用于将社交行为原始视频输入目标聚类数模型;
分割模块,基于聚类数函数的上界与下界用于将社交行为原始视频进行分割;
片段集合获取模块,用于获取对应的目标视频片段集合;
保存模块,用于将目标视频片段集合中的每个目标视频片段保存至以行为类别命名的文件夹中。
根据本申请的一个方面,一种电子设备,包括至少一个处理器以及至少一个存储器,其中,所述存储器上存储有计算机可读指令;所述计算机可读指令被一个或多个所述处理器执行,使得电子设备实现如上所述的多动物自由社交行为映射分类方法。
根据本申请的一个方面,一种存储介质,其上存储有计算机可读指令,所述计算机可读指令被一个或多个处理器执行,以实现如上所述的多动物自由社交行为映射分类方法。
根据本申请的一个方面,一种计算机程序产品,计算机程序产品包括计算机可读指令,计算机可读指令存储在存储介质中,电子设备的一个或多个处理器从存储介质读取计算机可读指令,加载并执行该计算机可读指令,使得电子设备实现如上所述的多动物自由社交行为映射分类方法。
本申请提供的技术方案带来的有益效果是:通过本申请的实施方案,而在进行行为分类的过程中,按照社交行为的自然结构设计无监督行为映射和分类框架,因此能全面覆盖未知的社交行为,先由算法将可能的行为类别一致地分开,然后人工定义,既有效覆盖了人工难定义的行为类别,也得到了一致的结果,并且可以对多动物的社交行为进行分类,从而可以提高对于动物社交行为分类的效率。
附图说明
为了更清楚地说明本申请实施例中的技术方案,下面将对本申请实施例描述中所需要使用的附图作简单地介绍。显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1是根据本申请所涉及的实施环境的示意图;
图2是根据一示例性实施例示出的一种多动物自由社交行为映射分类方 法的流程图;
图3是根据一示例性实施例示出的另一种多动物自由社交行为映射分类方法中S121到S123的流程图;
图4是根据一示例性实施例示出的另一种多动物自由社交行为映射分类方法中S131到S134的流程图;
图5是根据一示例性实施例示出的另一种多动物自由社交行为映射分类方法中S151到S154的流程图;
图6是根据一示例性实施例示出的另一种多动物自由社交行为映射分类方法中S171到S174的流程图;
图7是根据一示例性实施例示出的另一种多动物自由社交行为映射分类方法中S181到S184的流程图;
图8是根据一示例性实施例示出的一种多动物自由社交行为映射分类装置的结构框图;
图9是根据一示例性实施例示出的一种电子设备的结构框图。
具体实施方式
下面详细描述本申请的实施例,所述实施例的示例在附图中示出,其中自始至终相同或类似的标号表示相同或类似的元件或具有相同或类似功能的元件。下面通过参考附图描述的实施例是示例性的,仅用于解释本申请,而不能解释为对本申请的限制。
本技术领域技术人员可以理解,除非特意声明,这里使用的单数形式“一”、“一个”、“所述”和“该”也可包括复数形式。应该进一步理解的是,本公开的说明书中使用的措辞“包括”是指存在所述特征、整数、步骤、操作、元件和/或组件,但是并不排除存在或添加一个或多个其他特征、整数、步骤、操作、元件、组件和/或它们的组。应该理解,当我们称元件被“连接”或“耦接”到另一元件时,它可以直接连接或耦接到其他元件,或者也可以存在中间元件。此外,这里使用的“连接”或“耦接”可以包括无线连接或无线耦接。这里使用的措辞“和/或”包括一个或更多个相关联的列出项的全部或任一单元和全部组合。
为使本申请的目的、技术方案和优点更加清楚,下面将结合附图对本申请实施方式作进一步地详细描述。
图1为一种多动物自由社交行为映射分类方法所涉及的实施环境的示意图。该实施环境包括终端、服务器和配置了成员关联数据库的服务系统。
具体地,终端可供提供视频分类的客户端运行,可以是台式电脑、笔记本电脑、平板电脑、智能手机等等电子设备,在此不进行限定。
其中,客户端,提供视频分类功能,例如,媒体播放器、浏览器等,可以是应用程序形式,也可以是网页形式,相应地,客户端进行播放视频的用户界面则可以是程序窗口形式,还可以是网页页面形式的,此处也并未加以限定。
服务器可以是独立的物理服务器,也可以是多个物理服务器构成的服务器集群或者分布式系统,还可以是提供云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、CDN、以及大数据和人工智能平台等基础云计算服务的云服务器。此服务器是用于提供后台服务的电子设备,例如,本实施环境中,服务器为终端提供音视频数据的云存储服务。
服务器通过有线或者无线等方式预先与服务系统之间建立通信连接,并通过通信连接实现与服务系统的联动。该服务系统可以是一台服务器,也可以是由多台服务器构成的服务器集群。
通过终端与服务器的交互,运行于终端的客户端将向服务器发起资源使用邀请,请求服务器通过资源分配确定资源所在位置以及资源使用时间,并以此发出邀请。
对于服务器而言,通过资源使用邀请联动服务系统为邀请方所在客户端执行资源分配过程,并返回指示了资源所在位置以及资源使用时间的邀请结果至邀请方所在客户端,以供邀请方所在客户端进一步地确认是否向被邀请方所在客户端发出资源使用的邀请。
当然,根据实际营运的需要,服务器和服务系统也可以整合在同一服务器集群,以使资源分配由该同一服务器集群完成。
请参阅图2,本申请实施例提供了一种多动物自由社交行为映射分类方法,包括:
S100,基于多动物自由社交的视频获取对应的身体姿态数据;
其中,在使用社交行为拍摄模块拍摄多动物自由社交的视频,然后使用深度学习姿态估计模块估计多动物自由社交的身体姿态。
并且可以使用DeepLabCut工具中的多动物社交姿态估计实现该身体姿态的获取,首先提取多动物自由社交的视频中部分关键视频帧进行多动物身体姿态的人工标注,根据动物的种类不同,预定义的身体点数也不同,例如小鼠采用16点标注、鸟采用21点标注、犬类采用17点标注;在标注完成后,使用标注帧训练深度神经网络,待网络训练收敛后使用该网络估计所有视频中动物的身体姿态进行后续处理;通过上述过程,可以实现身体姿态数据的获取,并且在使用收敛后的神经网络,后续无须进行人工标注,从而可以减少人为因素的参与,从而提高标注效率。
S 110,基于身体姿态数据获取对应的序列集合。
其中,序列集合包括运动序列、动作序列以及距离序列,需要指出的是,运动序列表征社交动物每个身体点的瞬时运动速度,计算方法为相邻帧差分的绝对值乘以帧率;动作序列表征社交动物肢体的动作,计算方法为每个身体点的坐标值减身体中心点的坐标值,将每个动物对齐到笛卡尔坐标原点;距离序列表征社交动物在社交过程中肢体之间的位置变化,计算方法为每只动物和其他动物对应身体点的欧式距离。
S120,基于序列集合中的每个序列获取每个序列对应的二维时间序列以及对应的目标时间分割点;
其中,可以采用UMAP算法分解每个序列的二维时间序列,再获取到对应的二维时间序列;请参阅图3,对于目标时间分割点的获取过程具体如下:
S121,基于序列集合中的每个序列获取每个序列对应的二维表征;
其中,在本申请实施例中,包括运动序列、动作序列以及距离序列三种序列,所以对于二维表征,一共有三种序列的二维表征。
S122,将每个序列对应的二维表征用动态时间对齐核化算法分解获取每个序列对应的离散时间片段。
S123,基于离散时间片段确定对应的目标时间分割点;
其中,每个离散时间片段可以使用动态时间对齐核化算法计算对应有时间分割点,在将全部的时间分割点中确定所需要的目标时间分割点作为后续使用。
S130,调取预先设置的流行特征模型;
其中,在调取流形特征模型之前,需要将本申请所需对应的流形特征模型提前训练好。请参阅图4,对于流行特征模型的具体训练过程如下:
S131,获取每个序列对应的二维表征的二维时间序列。
S132,基于全部序列的二维时间序列确定对应的六维时间序列。
其中,六维时间序列为运动序列对应的二维时间序列、动作序列对应的二维时间序列以及距离序列对应的二维时间序列,通过三中序列对应的二维时间序列可以确定六维时间序列。
S133,在全部的六维时间序列中选择一部分作为模型训练的训练六维时间序列。
S134,基于将训练六维时间序列进行UMAP算法运算获取对应的训练二维流形表征,可以学习训练六维时间序列训练二维流形表征之间的映射关系,这里需要指出的是,这里输入的为六维时间序列对应的时间点,当输入的训练数据量足够时且以UMAP算法运算为基础的模型收敛之后,此时可以确定训练六维时间序列与训练二维流形表征之间的流形特征模型。
S140,基于全部二维时间序列对应的所有时间点输入流形特征模型,获取所有时间点对应的二维流形表征。
S150,基于目标时间分割点以及二维流形表征确定低维行为空间;
其中,为了确定低维行为空间,需要先构建对应的相似度矩阵,才能通过对应的表征数据,为了构建相似度矩阵,请参阅图5,具体的流程如下:
S151,从所有的时间片段随机抽取一部分片段作为基向量,调取二维流形表征和时间分割点;
其中,这里的时间片段前文提及离散时间片段,对于基向量的选择,可以随机抽取一部分进行确定。
S152,基于二维流形表征和时间分割点使用动态时间对齐核化算法度量所有时间序列片段和基向量的相似度,构建相似度矩阵。
S153,在构建相似度矩阵之后,可以使用UMAP算法获取相似度矩阵在二维空间的矩阵表征数据。
S154,基于矩阵表征数据建立低维行为空间。
S160,将低维行为空间的坐标输入高斯核函数获取低维行为空间点密度。
S170,基于低维行为空间点密度的稳定性确定目标聚类数模型;
其中,为了确定目标聚类数模型,请参阅图6,进行如下步骤:
S171,基于多个的方差参数获取对应的高斯核函数集合;
S172,将低维行为空间的坐标输入到高斯核函数集合中的高斯核函数中 获取每个高斯核函数输出的低维行为空间点密度;
S173,基于全部低维行为空间点密度的稳定性确定方差参数对应的最大稳定值与最小稳定值;
S174,将最大稳定值与最小稳定值定义为聚类数模型的上界与下界,并且确定目标聚类数模型。
对于S171到S173的执行步骤,是为了确定高斯核函数,是以高斯核输出的低维行为空间的密度是否稳定作为评判标准,其中,高斯核的方差参数作为输入,高斯核计算低维行为空间的密度为输出,选取能够使得低维行为空间的密度稳定的两个方差参数作为界限,其中,大的方差参数作为上界,小的方差参数为下界,然后将下界与上界对应的方差作为聚类数模型的最小稳定值与最大稳定值,从而确定聚类数模型。
通过本申请实施例中S100到S170的执行步骤,基于初始阶段的身体姿态数据,同现有技术不同的在于,不单单是针对单一动物,使用训练收敛后的深度神经网络,可以多种动物进行标注,并且不需要进行人工参与标注,减少人为因素参与的情况下,可以提高标注效率。并且从运动序列、动作序列以及距离序列等序列计算动物的动作数据,对于训练得到稳定的流行特征模型,可以输出稳定的二维流形表征,结合目标时间分割点可以建立低维空间模型,采用合适的高斯函数计算低维行为空间的密度,并且以低维行为空间点的密度稳定性选择方差参数,待选定方差参数之后,便可以使用聚类数模型确定对应的目标聚类数模型,通过目标聚类数模型将多动物社交行为的视频进行分别,便可以实现动物行为的分类,在目标聚类数模型将动物不同的社交行为进行区分之后,后续可以由人工将目标聚类数模型区分的社交行为进行命名。
S180,获取多动物的社交行为原始视频,将社交行为原始视频输入目标聚类数模型获取对应的目标行为类别。
其中,在具体将多动物社交行为分类的过程中,请参阅图7,具体的行为包括:
S181,将社交行为原始视频输入目标聚类数模型;
其中,对于输入目标聚类数模型的社交行为原始视频,可以是一个动物,也可以多个动物的视频,对于本申请实施例中的方案而言,针对视频中不同的动物,可以对不同的动物进行标注。
S182,基于聚类数模型的上界与下界将社交行为原始视频进行分割;
其中,在进行社交行为原始视频分割时,聚类数模型的上界用于区分动物社交行为全局的差异,而聚类数模型的下界用于细分动物社交行为的精细差异。
S183,获取对应的目标视频片段集合;
其中,在经过聚类数模型的上界以及下界进行精细区分后,可以得到由聚类数模型区分出来的目标视频片段集合,并且每种行为都对应有目标视频片段集合。
S184,将目标视频片段集合中的每个目标视频片段保存至以行为类别命名的文件夹中;
其中,在将动物不同社交行为区分之后,可以由人工为动物的每种社交行为进行命名,并且在将社交行为原始视频输入之后,系统可以根据目标聚类数模型的分类将每个目标视频片段保存在对应的文件中。
下述为本申请装置实施例,可以用于执行本申请所涉及的多动物自由社交行为映射分类方法。对于本申请装置实施例中未披露的细节,请参照本申请所涉及的多动物自由社交行为映射分类方法的方法实施例。
请参阅图8,本申请实施例中提供了一种多动物自由社交行为映射分类装置,包括但不限于:
姿态数据获取模块200,基于多动物自由社交的视频用于获取对应的身体姿态数据;
序列集合获取模块210,基于身体姿态数据用于获取对应的序列集合,其中,序列集合包括运动序列、动作序列以及距离序列;
时间获取模块220,基于序列集合中的每个序列用于获取每个序列对应的二维时间序列以及对应的目标时间分割点;
特征模型调取模块230,用于调取预先设置的流行特征模型;
二维流形表征获取模块240,基于全部二维时间序列对应的所有时间点输入流形特征模型,用于获取对应的二维流形表征;
行为空间确定模块250,基于目标时间分割点以及二维流形表征用于确定低维行为空间;
空间点密度获取模块260,将低维行为空间的坐标输入高斯核函数用于获取低维行为空间点密度;
聚类数模型确定模块270,基于低维行为空间点密度的稳定性用于确定目标聚类数模型;
原始视频获取模块280,用于获取多动物的社交行为原始视频;
目标行为类别获取模块290,将社交行为原始视频输入目标聚类数模型用于获取对应的目标行为类别。
在一示例性实施例中,包括但不限于:
二维表征获取模块300,基于序列集合中的每个序列用于获取每个序列对应的二维表征;
离散时间片段获取模块310,将每个序列对应的二维表征用动态时间对齐核化算法分解用于获取每个序列对应的离散时间片段;
时间分割点确定模块320,基于离散时间片段用于确定对应的目标时间分割点;
目标时间分割点确定模块330,将全部的时间分割点进行合并运算用于确定目标时间分割点。
在一示例性实施例中,包括但不限于:
二维时间序列获取模块400,用于获取每个序列对应的二维表征的二维时间序列;
六维时间序列确定模块410,基于全部序列的二维时间序列用于确定对应的六维时间序列,其中,六维时间序列为运动序列对应的二维时间序列、动作序列对应的二维时间序列以及距离序列对应的二维时间序列;
训练序列获取模块420,基于所述六维时间序列用于获取训练六维时间序列;
流形特征模型确定模块430,基于将训练六维时间序列进行UMAP算法运算获取对应的训练二维流形表征,并用于确定训练六维时间序列与训练二维流形表征之间的流形特征模型。
在一示例性实施例中,包括但不限于:
矩阵构建模块500,基于所述二维流形表征与所述时间分割点用于构建相似度矩阵;
矩阵表征数据获取模块510,使用UMAP算法用于获取相似度矩阵在二维空间的矩阵表征数据;
低维行为空间建立模块520,基于矩阵表征数据用于建立低维行为空间。
在一示例性实施例中,包括但不限于:
基向量抽取模块600,从所有的时间片段随机用于抽取一部分片段作为基向量;
调取模块610,用于调取二维流形表征和时间分割点;
矩阵构建模块620,基于所述二维流形表征和时间分割点使用动态时间对齐核化算法度量所有时间序列片段和基向量的相似度,用于构建相似度矩阵。
在一示例性实施例中,包括但不限于:
函数集合获取模块700,基于多个的方差参数获取对应的高斯核函数集合;
行为空间点密度集合获取模块710,将低维行为空间的坐标输入到高斯核函数集合中的高斯核函数中用于获取每个高斯核函数输出的低维行为空间点密度;
最值稳定值确定模块720,基于全部低维行为空间点密度的稳定性用于确定方差参数对应的最大稳定值与最小稳定值;
目标聚类数模型确定模块730,将最大稳定值与最小稳定值定义为聚类数函数的上界与下界,并用于确定目标聚类数模型。
在一示例性实施例中,包括但不限于:
输入模块800,用于将社交行为原始视频输入目标聚类数模型;
分割模块810,基于聚类数函数的上界与下界用于将社交行为原始视频进行分割;
片段集合获取模块820,用于获取对应的目标视频片段集合;
保存模块830,用于将目标视频片段集合中的每个目标视频片段保存至以行为类别命名的文件夹中。
需要说明的是,上述实施例所提供的多动物自由社交行为映射分类装置在进行多动物自由社交行为映射分类时,仅以上述各功能模块的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能模块完成,即多动物自由社交行为映射分类装置的内部结构将划分为不同的功能模块,以完成以上描述的全部或者部分功能。
另外,上述实施例所提供的多动物自由社交行为映射分类装置与多动物自由社交行为映射分类方法的实施例属于同一构思,其中各个模块执行操作 的具体方式已经在方法实施例中进行了详细描述,此处不再赘述。
请参阅图9,本申请实施例中提供了一种电子设备4000,该电子设备400可以包括:台式电脑、笔记本电脑、服务器等。
在图9中,该电子设备4000包括至少一个处理器4001以及至少一个存储器4003。
其中,处理器4001和存储器4003之间的数据交互,可以通过至少一个通信总线4002实现。该通信总线4002可包括一通路,用于在处理器4001和存储器4003之间传输数据。通信总线4002可以是PCI(Peripheral Component Interconnect,外设部件互连标准)总线或EISA(Extended Industry Standard Architecture,扩展工业标准结构)总线等。通信总线4002可以分为地址总线、数据总线、控制总线等。为便于表示,图9中仅用一条粗线表示,但并不表示仅有一根总线或一种类型的总线。
可选地,电子设备4000还可以包括收发器4004,收发器4004可以用于该电子设备与其他电子设备之间的数据交互,如数据的发送和/或数据的接收等。需要说明的是,实际应用中收发器4004不限于一个,该电子设备4000的结构并不构成对本申请实施例的限定。
处理器4001可以是CPU(Central Processing Unit,中央处理器),通用处理器,DSP(Digital Signal Processor,数据信号处理器),ASIC(Application Specific Integrated Circuit,专用集成电路),FPGA(Field Programmable Gate Array,现场可编程门阵列)或者其他可编程逻辑器件、晶体管逻辑器件、硬件部件或者其任意组合。其可以实现或执行结合本申请公开内容所描述的各种示例性的逻辑方框,模块和电路。处理器4001也可以是实现计算功能的组合,例如包含一个或多个微处理器组合,DSP和微处理器的组合等。
存储器4003可以是ROM(Read Only Memory,只读存储器)或可存储静态信息和指令的其他类型的静态存储设备,RAM(Random Access Memory,随机存取存储器)或者可存储信息和指令的其他类型的动态存储设备,也可以是EEPROM(Electrically Erasable Programmable Read Only Memory,电可擦可编程只读存储器)、CD-ROM(Compact Disc Read Only Memory,只读光盘)或其他光盘存储、光碟存储(包括压缩光碟、激光 碟、光碟、数字通用光碟、蓝光光碟等)、磁盘存储介质或者其他磁存储设备、或者能够用于携带或存储具有指令或数据结构形式的期望的程序指令或代码并能够由电子设备400存取的任何其他介质,但不限于此。
存储器4003上存储有计算机可读指令,处理器4001可以通过通信总线4002读取存储器4003中存储的计算机可读指令。
该计算机可读指令被一个或多个处理器4001执行以实现上述各实施例中的多动物自由社交行为映射分类方法。
此外,本申请实施例中提供了一种存储介质,该存储介质上存储有计算机可读指令,所述计算机可读指令被一个或多个处理器执行,以实现如上所述的多动物自由社交行为映射分类方法。
本申请实施例中提供了一种计算机程序产品,计算机程序产品包括计算机可读指令,计算机可读指令存储在存储介质中,电子设备的一个或多个处理器从存储介质读取计算机可读指令,加载并执行该计算机可读指令,使得电子设备实现如上所述的多动物自由社交行为映射分类方法。
在本申请实施例中的方案,首先需要通过社交行为拍摄模块去拍摄多动物自由社交的视频,从而获取到对应的身体姿态数据,期间不需要进行人工标注,直接由训练收敛的深度神经网络处理,从而可以减少人为因素的参与;然后可以基于身体姿态数据获取到对应的运动序列、动作序列以及距离序列,相对应的,对于每个序列都有对应的二维时间序列以及二维时间序列对应的离散时间片段,通过离散时间片段可以获取到时间分割点。
然后调取预先训练好的流形特征模特,对应的,将全部二维时间序列对应的所有时间点输入流形特征模型,可以获取所有时间点对应的二维流形表征;同时需要所有的时间片段随机抽取一部分片段作为基向量,使用动态时间对齐核化算法计算出时间序列片段与基向量的相似度,从而可以构建相似度矩阵,将相似度矩阵使用UMAP算法获取相似度矩阵在二维空间的矩阵表征数据,在获取矩阵表征数据之后,便可以建立低维行为空间。
在建立低维行为空间之后,可以获取低维行为空间的坐标,并且将坐标信息输入高斯核函数中,是以高斯核输出的低维行为空间的密度是否稳定作为评判标准,其中,高斯核的方差参数作为输入,高斯核计算低维行为空间的密度为输出,选取能够使得低维行为空间的密度稳定的两个方差参数作为界限,其中,大的方差参数作为上界,小的方差参数为下界,然后将下界与 上界对应的方差作为聚类数模型的最小稳定值与最大稳定值,从而确定目标聚类数模型。
在确定目标聚类数模型之后,将获取的多动物的社交行为原始视频输入到目标聚类数模型中,在经目标过聚类数模型的上界以及下界进行精细区分后,可以得到由目标聚类数模型区分出来的目标视频片段集合,并且将每个目标视频片段保存在对应的文件中,从而实现多动物社交行为的分类。
综上所述,在在本申请实施例中的方案,首先可以进行多动物的社交行为的分类,并且在进行社交行为原始视频分类的过程中,在保证分类精确度的情况下,在通过动物的身体姿态数据以及分解,可以解决社交行为分类不一致的问题,流形特征和低维空间映射分类解决了社交行为定义不全面的问题,无监督社交行为分类的先分类后定义的策略解决了社交行为分类效率低的问题,在不需要人为参与的情况下,可以提高多动物社交行为细分分类的效率。
应该理解的是,虽然附图的流程图中的各个步骤按照箭头的指示依次显示,但是这些步骤并不是必然按照箭头指示的顺序依次执行。除非本文中有明确的说明,这些步骤的执行并没有严格的顺序限制,其可以以其他的顺序执行。而且,附图的流程图中的至少一部分步骤可以包括多个子步骤或者多个阶段,这些子步骤或者阶段并不必然是在同一时刻执行完成,而是可以在不同的时刻执行,其执行顺序也不必然是依次进行,而是可以与其他步骤或者其他步骤的子步骤或者阶段的至少一部分轮流或者交替地执行。
以上所述仅是本申请的部分实施方式,应当指出,对于本技术领域的普通技术人员来说,在不脱离本申请原理的前提下,还可以做出若干改进和润饰,这些改进和润饰也应视为本申请的保护范围。

Claims (10)

  1. 一种多动物自由社交行为映射分类方法,其特征在于,包括:
    基于多动物自由社交的视频获取对应的身体姿态数据;
    基于身体姿态数据获取对应的序列集合,其中,序列集合包括运动序列、动作序列以及距离序列;
    基于序列集合中的每个序列获取每个序列对应的二维时间序列以及对应的目标时间分割点;
    调取预先设置的流行特征模型;
    基于全部二维时间序列对应的所有时间点输入流形特征模型,获取对应的二维流形表征;
    基于目标时间分割点以及二维流形表征确定低维行为空间;
    将低维行为空间的坐标输入高斯核函数获取低维行为空间点密度;
    基于低维行为空间点密度的稳定性确定目标聚类数模型;
    获取多动物的社交行为原始视频;
    将社交行为原始视频输入目标聚类数模型获取对应的目标行为类别。
  2. 如权利要求1所述的方法,其特征在于,在所述获取对应的目标时间分割点的过程中,所述方法还包括:
    基于序列集合中的每个序列获取每个序列对应的二维表征;
    将每个序列对应的二维表征用动态时间对齐核化算法分解获取每个序列对应的离散时间片段;
    基于离散时间片段确定对应的目标时间分割点;
    将全部的时间分割点进行合并运算确定目标时间分割点。
  3. 如权利要求2所述的方法,其特征在于,在所述调取流形特征模型前,所述方法还包括:
    获取每个序列对应的二维表征的二维时间序列;
    基于全部序列的二维时间序列确定对应的六维时间序列,其中,六维时间序列为运动序列对应的二维时间序列、动作序列对应的二维时间序列以及距离序列对应的二维时间序列;
    基于所述六维时间序列获取训练六维时间序列;
    基于将训练六维时间序列进行UMAP算法运算获取对应的训练二维流形表征,并确定训练六维时间序列与训练二维流形表征之间的流形特征模型。
  4. 如权利要求2所述的方法,其特征在于,在所述基于目标时间分割点以及二维流形表征确定低维行为空间的过程中,所述方法还包括:
    基于所述二维流形表征与所述时间分割点构建相似度矩阵;
    使用UMAP算法获取相似度矩阵在二维空间的矩阵表征数据;
    基于矩阵表征数据建立低维行为空间。
  5. 如权利要求1所述的方法,其特征在于,在所述基于所述二维流形表征与所述时间分割点构建相似度矩阵的过程中,所述方法还包括:
    从所有的时间片段随机抽取一部分片段作为基向量,调取二维流形表征和时间分割点;
    基于所述二维流形表征和时间分割点使用动态时间对齐核化算法度量所有时间序列片段和基向量的相似度,构建相似度矩阵。
  6. 如权利要求1所述的方法,其特征在于,在所述基于低维行为空间点密度的稳定性确定目标聚类数模型的过程中,所述方法还包括:
    基于多个的方差参数获取对应的高斯核函数集合;
    将低维行为空间的坐标输入到高斯核函数集合中的高斯核函数中获取每个高斯核函数输出的低维行为空间点密度;
    基于全部低维行为空间点密度的稳定性确定方差参数对应的最大稳定值与最小稳定值;
    将最大稳定值与最小稳定值定义为聚类数函数的上界与下界,并且确定目标聚类数模型。
  7. 如权利要求1所述的方法,其特征在于,在所述将社交行为原始视频基于目标聚类数模型进行分割的过程中,所述方法还包括:
    将社交行为原始视频输入目标聚类数模型;
    基于聚类数函数的上界与下界将社交行为原始视频进行分割;
    获取对应的目标视频片段集合;
    将目标视频片段集合中的每个目标视频片段保存至以行为类别命名的文件夹中。
  8. 一种多动物自由社交行为映射分类装置,其特征在于,包括:
    姿态数据获取模块,基于多动物自由社交的视频用于获取对应的身体姿 态数据;
    序列集合获取模块,基于身体姿态数据用于获取对应的序列集合,其中,序列集合包括运动序列、动作序列以及距离序列;
    时间获取模块,基于序列集合中的每个序列用于获取每个序列对应的二维时间序列以及对应的目标时间分割点;
    特征模型调取模块,用于调取预先设置的流行特征模型;
    二维流形表征获取模块,基于全部二维时间序列对应的所有时间点输入流形特征模型,用于获取对应的二维流形表征;
    行为空间确定模块,基于目标时间分割点以及二维流形表征用于确定低维行为空间;
    空间点密度获取模块,将低维行为空间的坐标输入高斯核函数用于获取低维行为空间点密度;
    聚类数模型确定模块,基于低维行为空间点密度的稳定性用于确定目标聚类数模型;
    原始视频获取模块,用于获取多动物的社交行为原始视频;
    目标行为类别获取模块,将社交行为原始视频输入目标聚类数模型用于获取对应的目标行为类别。
  9. 一种电子设备,其特征在于,包括:至少一个处理器以及至少一个存储器,其中,
    所述存储器上存储有计算机可读指令;
    所述计算机可读指令被一个或多个所述处理器执行,使得电子设备实现如权利要求1至7中任一项所述的多动物自由社交行为映射分类方法。
  10. 一种存储介质,其上存储有计算机可读指令,其特征在于,所述计算机可读指令被一个或多个处理器执行,以实现如权利要求1至7中任一项所述的多动物自由社交行为映射分类方法。
PCT/CN2023/135942 2023-12-01 2023-12-01 多动物自由社交行为映射分类方法、装置、电子设备及存储介质 Pending WO2025112052A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/CN2023/135942 WO2025112052A1 (zh) 2023-12-01 2023-12-01 多动物自由社交行为映射分类方法、装置、电子设备及存储介质

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2023/135942 WO2025112052A1 (zh) 2023-12-01 2023-12-01 多动物自由社交行为映射分类方法、装置、电子设备及存储介质

Publications (1)

Publication Number Publication Date
WO2025112052A1 true WO2025112052A1 (zh) 2025-06-05

Family

ID=95896180

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2023/135942 Pending WO2025112052A1 (zh) 2023-12-01 2023-12-01 多动物自由社交行为映射分类方法、装置、电子设备及存储介质

Country Status (1)

Country Link
WO (1) WO2025112052A1 (zh)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN1975779A (zh) * 2006-09-14 2007-06-06 浙江大学 三维人体运动数据分割方法
CN112057079A (zh) * 2020-08-07 2020-12-11 中国科学院深圳先进技术研究院 一种基于状态与图谱的行为量化方法和终端
WO2022027590A1 (zh) * 2020-08-07 2022-02-10 中国科学院深圳先进技术研究院 一种基于状态与图谱的行为量化方法和终端
US20220076003A1 (en) * 2020-09-04 2022-03-10 Hitachi, Ltd. Action recognition apparatus, learning apparatus, and action recognition method

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN1975779A (zh) * 2006-09-14 2007-06-06 浙江大学 三维人体运动数据分割方法
CN112057079A (zh) * 2020-08-07 2020-12-11 中国科学院深圳先进技术研究院 一种基于状态与图谱的行为量化方法和终端
WO2022027590A1 (zh) * 2020-08-07 2022-02-10 中国科学院深圳先进技术研究院 一种基于状态与图谱的行为量化方法和终端
US20220076003A1 (en) * 2020-09-04 2022-03-10 Hitachi, Ltd. Action recognition apparatus, learning apparatus, and action recognition method

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
HAN YANING, CHEN KE, WANG YUNKE, LIU WENHAO, WANG XIAOJING, LIAO JIAHUI, HUANG YITING, HAN CHUANLIANG, HUANG KANG, ZHANG JIAJIA, C: "Social Behavior Atlas: A computational framework for tracking and mapping 3D close interactions of free-moving animals", BIORXIV, 6 March 2023 (2023-03-06), XP093317386, Retrieved from the Internet <URL:https://www.biorxiv.org/content/10.1101/2023.03.05.531235v1.full.pdf> DOI: 10.1101/2023.03.05.531235 *

Similar Documents

Publication Publication Date Title
CN108460338B (zh) 人体姿态估计方法和装置、电子设备、存储介质、程序
JP7601730B2 (ja) 学習方法、情報処理装置、学習プログラム
CN110431560B (zh) 目标人物的搜索方法和装置、设备和介质
CN115098732B (zh) 数据处理方法及相关装置
WO2019154262A1 (zh) 一种图像分类方法及服务器、用户终端、存储介质
CN113255713B (zh) 用于跨对象变化的数字图像选择的机器学习
CN110363210A (zh) 一种图像语义分割模型的训练方法和服务器
CN110019896A (zh) 一种图像检索方法、装置及电子设备
CN105468596A (zh) 图片检索方法和装置
CN113704617A (zh) 物品推荐方法、系统、电子设备及存储介质
CN113569888A (zh) 图像标注方法、装置、设备及介质
CN115439922A (zh) 对象行为识别方法、装置、设备及介质
CN102999635A (zh) 语义可视搜索引擎
CN108537270A (zh) 基于多标签学习的图像标注方法、终端设备及存储介质
CN117216362A (zh) 内容推荐方法、装置、设备、介质和程序产品
Lu et al. Content-based search for deep generative models
CN106407281B (zh) 图像检索方法及装置
Li et al. MVHANet: multi-view hierarchical aggregation network for skeleton-based hand gesture recognition
CN117475340A (zh) 视频数据处理方法、装置、计算机设备和存储介质
WO2025112052A1 (zh) 多动物自由社交行为映射分类方法、装置、电子设备及存储介质
CN117636395A (zh) 多动物自由社交行为映射分类方法、装置、电子设备及存储介质
CN114708449B (zh) 相似视频的确定方法、实例表征模型的训练方法及设备
CN116522271B (zh) 特征融合模型处理、样本检索方法、装置和计算机设备
CN116863184B (zh) 图像分类模型训练方法、图像分类方法、装置、设备
CN114139711B (zh) 分布式推断方法、数据处理方法、装置、终端及计算设备

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23959958

Country of ref document: EP

Kind code of ref document: A1