WO2020187300A1 - 监控系统、方法、装置、服务器及存储介质 - Google Patents

监控系统、方法、装置、服务器及存储介质 Download PDF

Info

Publication number
WO2020187300A1
WO2020187300A1 PCT/CN2020/080256 CN2020080256W WO2020187300A1 WO 2020187300 A1 WO2020187300 A1 WO 2020187300A1 CN 2020080256 W CN2020080256 W CN 2020080256W WO 2020187300 A1 WO2020187300 A1 WO 2020187300A1
Authority
WO
WIPO (PCT)
Prior art keywords
emotion recognition
user
state information
emotion
voice
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2020/080256
Other languages
English (en)
French (fr)
Inventor
李婉瑜
陈展
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hangzhou Hikvision Digital Technology Co Ltd
Original Assignee
Hangzhou Hikvision Digital Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hangzhou Hikvision Digital Technology Co Ltd filed Critical Hangzhou Hikvision Digital Technology Co Ltd
Publication of WO2020187300A1 publication Critical patent/WO2020187300A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/51Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
    • G10L25/63Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for estimating an emotional state
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L17/00Speaker identification or verification techniques
    • G10L17/04Training, enrolment or model building

Definitions

  • the present disclosure relates to the field of monitoring technology, in particular to a monitoring system, method, device, server, and storage medium.
  • monitoring systems are widely used in various fields. For example, they can be used to monitor and manage targets such as users and animals. Illustratively, they can be used to monitor and manage vulnerable groups such as left-behind children or the elderly.
  • the monitoring system generally includes a security bracelet and a data management server.
  • the security bracelet is usually worn on the user's wrist and can be used to obtain the user's location information and report the user to the data management server Location information.
  • the data management server can be used to store the basic information of each user and the location information reported by the safety bracelet worn by each user, so that the management personnel can monitor and manage related personnel based on the information.
  • the embodiments of the present disclosure provide a monitoring system, method, device, server, and storage medium, which can solve the problem of poor management performance of the monitoring system in related technologies.
  • the technical solution is as follows:
  • a monitoring system in a first aspect, includes: a voice collection device, an emotion recognition server, and a data management server, the emotion recognition server establishes a communication connection with the voice collection device and the data management server, respectively ;
  • the voice collection device is configured to collect voice signals, and send the voice signal and an identity identifier to the emotion recognition server, where the identity identifier is an identifier of a user who has an associated relationship with the voice collection device;
  • the emotion recognition server is configured to call all emotion recognition models of the user based on the identity identifier, and each emotion recognition model corresponds to a kind of emotional state information of the user; based on the voice signal, through the emotions that are called
  • the recognition model determines the current emotional state information of the user; sends the emotional state information to the data management server;
  • the data management server is used to manage the emotional state information.
  • the emotion recognition server is configured to call all emotion recognition models of the user based on the identity identifier, including:
  • the emotion recognition server is configured to determine all corresponding emotion recognition model identifiers from the stored reference correspondence relationship based on the identity identifier, and the reference correspondence relationship is used to store the identity identifier of each user in at least one user and all Describe the correspondence between all the emotion recognition model identifiers of each user; call the emotion recognition model corresponding to all the determined emotion recognition model identifiers.
  • the emotion recognition server is configured to determine the current emotion state information of the user through the invoked emotion recognition model based on the voice signal, including:
  • the emotion recognition server is used to extract the voiceprint features of the speech signal; respectively input the voiceprint features to each emotion recognition model of all the emotion recognition models called, and each emotion recognition model pairs
  • the voiceprint feature is recognized and processed and the emotion similarity is output; based on all the output emotion similarities, the emotion recognition model corresponding to the maximum emotion similarity is determined; the emotion state information corresponding to the determined emotion recognition model is determined as the current user Emotional state information.
  • the emotion recognition server is used for invoking all emotion recognition models of the user based on the identity, and is also used for:
  • the emotion recognition server is configured to call a voice verification model corresponding to the user based on the identity; perform voice verification on the voice signal through the voice verification model; correspondingly, when the voice signal verification is passed When, execute the operation of invoking all emotion recognition models of the user based on the identity identifier.
  • the emotion recognition server is configured to determine the current emotional state information of the user through the invoked emotion recognition model based on the voice signal, and then to:
  • the emotion recognition server is configured to save the voice signal and the emotion state information as training samples, and the training samples are used to continue training the emotion recognition model corresponding to the emotion state information.
  • the emotion recognition server configured to send the emotion state information to the data management server, includes:
  • the emotion recognition server is configured to send the voice signal and the emotion state information to the data management server;
  • the data management server is used to manage the emotional state information, including:
  • the data management server is used to manage the voice signal and the emotional state information.
  • a monitoring method is provided.
  • the method is applied to an emotion recognition server of a monitoring system.
  • the monitoring system further includes a voice collection device and a data management server, and the emotion recognition server is respectively connected to the voice collection device.
  • a communication connection is established with the data management server; the method includes:
  • each emotion recognition model corresponds to a kind of emotional state information of the user
  • the determined emotional state information is sent to the data management server for management.
  • the invoking all emotion recognition models of the user based on the identity identifier includes:
  • all corresponding emotion recognition model identifiers are determined from the stored reference correspondence relationship, and the reference correspondence relationship is used to store the identity identifier of each user in at least one user and all the emotion recognition of each user Correspondence between model identifiers;
  • the determining the current emotional state information of the user through the invoked emotional recognition model based on the voice signal includes:
  • each emotion recognition model Respectively input the voiceprint feature to each emotion recognition model of all the called emotion recognition models, and each emotion recognition model will recognize the voiceprint feature and output the emotion similarity;
  • the emotion state information corresponding to the determined emotion recognition model is determined as the current emotion state information of the user.
  • the method before the invoking all emotion recognition models of the user based on the identity, the method further includes:
  • the method further includes:
  • the voice signal and the emotional state information are saved as training samples, and the training samples are used to continue training the emotional recognition model corresponding to the emotional state information.
  • the sending the determined emotional state information to the data management server includes:
  • the voice signal and the determined emotional state information are sent to the data management server for management.
  • a monitoring device configured in an emotion recognition server of a monitoring system.
  • the monitoring system further includes a voice collection device and a data management server.
  • the emotion recognition server is connected to the voice collection device and the The data management server establishes a communication connection; the device includes:
  • the receiving module is configured to receive a voice signal and an identity identifier sent by the voice collection device, the voice signal is collected by the voice collection device, and the identity identifier is that of a user who is associated with the voice collection device logo
  • a calling module configured to call all emotion recognition models of the user based on the identity, each emotion recognition model corresponds to a kind of emotional state information of the user;
  • the determining module is configured to determine the current emotional state information of the user through the called emotion recognition model based on the voice signal;
  • the sending module is used to send the determined emotional state information to the data management server for management.
  • the calling module is used to:
  • all corresponding emotion recognition model identifiers are determined from the stored reference correspondence relationship, and the reference correspondence relationship is used to store the identity identifier of each user in at least one user and all the emotion recognition of each user Correspondence between model identifiers;
  • the determining module is used to:
  • each emotion recognition model Respectively input the voiceprint feature to each emotion recognition model of all the called emotion recognition models, and each emotion recognition model will recognize the voiceprint feature and output the emotion similarity;
  • the calling module is also used to:
  • the device further includes:
  • the storage module is configured to save the voice signal and the emotional state information as training samples, and the training samples are used to continue training the emotional recognition model corresponding to the emotional state information.
  • the sending module is used to:
  • the voice signal and the determined emotional state information are sent to the data management server for management.
  • an emotion recognition server including:
  • a memory for storing processor executable instructions
  • the processor is configured to implement the monitoring method described in the second aspect.
  • a computer-readable storage medium is provided, and instructions are stored on the computer-readable storage medium, and when the instructions are executed by a processor, the monitoring method described in the second aspect is implemented.
  • a computer program product containing instructions which when running on a computer, causes the computer to execute the monitoring method described in the second aspect.
  • the voice collection device collects a voice signal, and sends the voice signal and an identity identifier to the emotion server, where the identity identifier is an identifier of a user who has an associated relationship with the voice collection device.
  • the emotion server calls all the emotion recognition models corresponding to the identity, that is, calls all the emotion recognition models of the user. Then, based on the voice signal, the current emotional state information of the user is determined through all the emotion recognition models called, and sent to the data management server for management. That is, the monitoring system can monitor the user's emotional state, which increases the management performance of the monitoring system.
  • Fig. 1 is a frame diagram of a monitoring system according to an exemplary embodiment
  • Fig. 2 is a flowchart showing a monitoring method according to an exemplary embodiment
  • Fig. 3 is a schematic diagram showing a principle of voice verification according to another exemplary embodiment
  • Fig. 4 shows a schematic diagram of the basic principle of emotion recognition according to another exemplary embodiment
  • Fig. 5 is a schematic structural diagram showing a monitoring device according to an exemplary embodiment
  • Fig. 6 is a schematic structural diagram showing a monitoring device according to another exemplary embodiment
  • Fig. 7 is a schematic structural diagram showing a server 700 according to an exemplary embodiment.
  • the current monitoring system only has the basic information management of left-behind children (such as reminding which basic information is incomplete, etc.), GPS (Global Positioning System, global positioning system) positioning and other functions, and the safety bracelet worn by the left-behind children
  • the function is limited and can only be used to locate and store basic information, but it is difficult to make more targeted dynamic monitoring of the psychological and physical health of left-behind children, and it is impossible to monitor children’s emotional movements in a timely and effective manner.
  • the embodiments of the present disclosure provide a monitoring system, which can monitor and manage children's emotional trends, and increase the management performance of the monitoring system.
  • FIG. 1 is a framework diagram of a monitoring system according to an exemplary embodiment.
  • the monitoring system mainly includes: a voice collection device 110, an emotion recognition server 120, and a data management server 130.
  • the emotion recognition server 120 establishes a communication connection with the voice collection device 110 and the data management server 130 respectively.
  • the voice collection device 110 has a voice collection function, which can be used to collect a user's voice signal and send the voice signal to the emotion recognition server 120.
  • the voice collection device 110 may be configured with a wearable component, so that the user can use the wearable component to wear it on the body.
  • the voice collection device 110 may also be a wearable device configured with a voice collector, for example, it may be a bracelet, a watch, etc., configured with a voice collector, which is not limited in the embodiment of the present disclosure.
  • the voice collection device 110 may also have functions such as positioning in addition to the voice collection function.
  • the emotion recognition server 120 may be used to perform emotion recognition on the user based on the user's voice signal to determine the current emotional state of the user.
  • the emotion recognition server 120 may be one server, or the emotion recognition server 120 may also be a server cluster composed of multiple servers.
  • the emotion recognition server 120 may It includes a voiceprint authentication algorithm server 120a, an emotion recognition algorithm server 120b, and an emotion management library 120c.
  • the voiceprint authentication algorithm server 120a can be used to perform voice verification on the user's voice signal.
  • the emotion management database 120c stores all the emotion recognition models of each of the multiple users.
  • the voiceprint authentication algorithm server 120a After the voice verification is passed, the emotion management library 120c is triggered to obtain the emotion recognition model of the user.
  • the emotion management library 120c shares the acquired emotion recognition model with the emotion recognition algorithm server 120b, so that the emotion recognition algorithm server 120b uses the emotion recognition model shared by the emotion management library 120c to recognize the emotional state of the user, and recognize The sent emotional state information is sent to the data management server 130.
  • the data management server 130 can be used to dynamically manage the emotional state information to obtain management information, so that the management personnel can timely monitor the user's emotional trend according to the management information in the data management server 130.
  • the data management server 130 may be one server, or may also be a server cluster composed of multiple servers, which is not limited in the embodiment of the present disclosure.
  • the monitoring system may further include a virtual server 140, which is respectively connected to the voice collection device 110 and the emotion recognition server 120 to transfer the voice signal transmitted by the voice collection device 110 to the emotion recognition Server 120.
  • the virtual server 140 may be referred to as a switch.
  • the monitoring system may further include a remote monitoring server 150 which is respectively connected to the emotion recognition server 120 and the data management server 130 to transfer data transmitted by the emotion recognition server 120 to the data management server 130.
  • a remote monitoring server 150 which is respectively connected to the emotion recognition server 120 and the data management server 130 to transfer data transmitted by the emotion recognition server 120 to the data management server 130.
  • the monitoring system can realize longer-distance monitoring through the data transmission between the virtual server 140 and the remote monitoring server 150, that is, the monitoring system can realize a wider range of monitoring.
  • Fig. 2 is a flowchart of a monitoring method according to an exemplary embodiment.
  • the monitoring method may include the following steps:
  • Step 201 The voice collection device collects a voice signal, and sends the voice signal and an identity identifier to the emotion recognition server, where the identity identifier is an identifier of a user who has an associated relationship with the voice collection device.
  • the user having an association relationship with the voice collection device may refer to the user who uses the voice collection device, or may also refer to the owner of the voice collection device, etc.
  • the identity can be used to uniquely identify a user.
  • the manager can issue a voice collection device for each left-behind child.
  • the voice collection device can be embedded with A security bracelet of a voice collector, etc., to collect voice signals through the voice collection device.
  • the voice collection device can perform the collection operation in real time. In another possible implementation manner, the voice collection device can also perform the collection operation every reference time length.
  • the reference duration can be set by the user according to actual needs, or can also be set by default by the voice collection device, which is not limited in the embodiment of the present disclosure.
  • the voice collection device collects the voice signal
  • the voice signal and the identity identifier are sent to the emotion recognition server, so that the emotion recognition server can recognize the emotional state of the user.
  • the monitoring system further includes a virtual server
  • the voice collection device sends the collected voice signal and the identity to the virtual server
  • the virtual server forwards the voice signal and the identity to the virtual server.
  • the emotion recognition server
  • the virtual server can determine the spectral energy of the voice signal, and when the spectral energy is greater than or equal to the spectral energy threshold, forward the voice signal and the identity to the emotion recognition server; otherwise, when the spectral energy is less than the spectral energy
  • the threshold is set, the voice signal and the identity identifier may not be forwarded to the emotion recognition server.
  • the spectrum energy threshold can be customized by the user according to actual needs, or can be set by default by the virtual server, which is not limited in the embodiment of the present disclosure.
  • the virtual server can decide whether to send the voice signal and the identity collected by the voice collection device to the emotion recognition server.
  • speech signals expressed by different emotions also have different structural characteristics and distribution laws in their time structure, amplitude structure, fundamental frequency structure and formant structure.
  • the spectral energy of the speech signal is greater than or equal to the spectral energy threshold, it can generally indicate that the user's speech is not stable, which can indicate that the user may be emotionally agitated. In this case, the user may be abnormal. Therefore, the virtual The server sends the voice signal and the identification to the emotion recognition server for further emotion recognition.
  • the spectral energy of the speech signal When the spectral energy of the speech signal is less than the spectral energy threshold, it can generally indicate that the user’s speaking mood is relatively stable, which can indicate that the user’s mood may be relatively stable. In this case, it indicates that the user is not abnormal. Forward the voice signal and the identification to the emotion recognition server, that is, the voice signal can be discarded, and continue to wait or process the next voice signal and identification data sent by the voice collection device, thus reducing the calculation of the emotion recognition server the amount.
  • Step 202 The emotion recognition server invokes a voice verification model corresponding to the user based on the identity, and performs voice verification on the voice signal through the voice verification model.
  • the voice signal sent by the voice collection device to the emotion recognition server may not be from a user who has an associated relationship with the voice collection device.
  • a voice collection device worn by a left-behind child A is sent to emotion
  • the voice signal of the recognition server may come from left-behind child B who has a dispute with left-behind child A.
  • the emotion recognition server after the emotion recognition server receives the voice signal and the identity sent by the voice collection device, it can verify the voice signal based on the identity, that is, verify whether the voice signal belongs to the user. .
  • the emotion recognition server may pre-store the correspondence between the identity of each user and the voice verification model, that is, each user may correspond to a voice verification model. In this way, the emotion recognition server can call the voice verification model corresponding to the user based on the identity, and use the voice verification model to perform voice verification.
  • the emotion recognition server may extract the voiceprint feature of the voice signal, input the voiceprint feature into the voice verification model, and the voice verification model will perform verification processing to output voice similarity.
  • the voice similarity is greater than or equal to the voice similarity threshold, it may be determined that the voice signal verification is passed; otherwise, when the voice similarity is less than the voice similarity threshold, it may be determined that the voice signal verification fails.
  • the voice similarity threshold may be customized by the user according to actual needs, or may be set by default by the emotion recognition server, which is not limited in the embodiment of the present disclosure.
  • the voice similarity when the voice similarity is greater than or equal to the voice similarity threshold, it means that the voice signal is from the user, that is, belongs to the user, so it can be determined that the voice signal has passed the verification.
  • the voice similarity is less than the voice similarity threshold, it indicates that the voice signal does not come from the user, that is, it does not belong to the user. In this case, it can be determined that the voice signal verification fails.
  • the emotion recognition server after receiving the voice signal, the emotion recognition server first performs voice verification on the voice signal to determine whether the voice signal really belongs to the user, which can improve the accuracy of the management of the monitoring system.
  • the voice verification model of each user may be obtained through training in advance. Specifically, the voice verification model of each user may be obtained by training the network model to be trained based on a large number of training samples of each user. For example, please refer to Figure 3.
  • a voice verification model to be trained is established, multiple voice segments of any user are obtained, and the voiceprint feature of each voice segment in the multiple voice segments is extracted , Input the extracted voiceprint features as training samples into the voice verification model to be trained for deep learning, and obtain the trained voice verification model corresponding to any user.
  • the trained voice verification model corresponding to each user can be determined.
  • test samples of each user can be used to evaluate the performance of the trained speech verification model.
  • the voiceprint feature of the test sample of any user can be extracted, and the voiceprint feature can be input into the trained voice verification model corresponding to any user.
  • the voice verification output result When greater than or equal to the first performance threshold, it is determined that the trained voice verification model meets the actual verification requirements. In this case, the correspondence between the trained voice verification model and the identity of any user can be stored . Conversely, when the voice verification output result is less than the first performance threshold, it means that the trained voice verification model does not meet the actual verification requirements. In this case, you can continue to obtain training samples of any user after verification.
  • the voice verification model for deep learning. In this way, according to this implementation, the voice verification model of each user can be determined.
  • the first performance threshold may be customized by the user according to actual needs, or may be set by default by the emotion recognition server, which is not limited in the embodiment of the present disclosure.
  • Step 203 When the voice signal is verified, the emotion recognition server calls all the emotion recognition models of the user based on the identity identifier, and each emotion recognition model corresponds to an emotional state of the user.
  • the emotion recognition server calls all the emotion recognition models of the user based on the identity.
  • the specific implementation of invoking all emotion recognition models of the user based on the identity may include: based on the identity, determining all corresponding emotion recognition model identities from the stored reference correspondence relationship, the reference The corresponding relationship is used to store the corresponding relationship between the identity identifier of each user in the multiple users and all the emotion recognition model identifiers of each user; and call the emotion recognition model corresponding to all the determined emotion recognition model identifiers.
  • each emotion recognition model identifier can be used to uniquely identify an emotion recognition model, and each user can correspond to one or more emotion recognition models, and each emotion recognition model corresponds to a kind of emotional state information of the user.
  • the emotion recognition model of each user can include a first emotion recognition model, a second emotion recognition model, and a third emotion recognition model
  • the emotion state information corresponding to the first emotion recognition model can be "fear”
  • the second emotion recognition model The emotional state information corresponding to the emotion recognition model may be "crying”
  • the emotional state information corresponding to the third emotion recognition model may be "grief and anger” and so on.
  • the emotion recognition server can pre-store all the emotion recognition models of each user, and store the reference correspondence between the identity of each user and all the emotion recognition model identifiers corresponding to the user. In this way, the emotion recognition server can Based on the user's identity and the reference corresponding relationship, all the user's emotion recognition models are called. For example, all the emotion recognition models of the user called include the first emotion recognition model, the second emotion recognition model, and the third emotion recognition model.
  • each emotion recognition model of each user may be obtained through training in advance. Specifically, each emotion recognition model of each user may be obtained by training the network model to be trained based on a large number of training samples. For example, please refer to Figure 4.
  • any user for any emotional state of any user, establish an emotional recognition model to be trained, and obtain the emotional state of any user To obtain the voiceprint features of each voice segment, and then perform digitization and preprocessing, endpoint detection processing, and feature extraction processing on each of the multiple voice segments in sequence, and then extract the voiceprint features Input as training data into the emotion recognition model to be trained for deep learning, and obtain the emotion recognition model corresponding to any emotion state, and then record the emotion state information of any emotion state and the emotion recognition model. relationship. In this way, according to this implementation manner, the emotion recognition model corresponding to each emotion state information of any user can be determined.
  • the test sample of any user can be used to evaluate the performance of the trained emotion recognition model.
  • test samples in the emotional state corresponding to the certain trained emotion recognition model can be obtained, and the obtained test samples are sequentially digitized and preprocessed, Endpoint detection processing and feature extraction processing to obtain the voiceprint feature of the test sample, and input the voiceprint feature into the trained emotion recognition model.
  • the output result of emotion recognition is greater than or equal to the second performance threshold, It is determined that the trained emotion recognition model meets the actual emotion recognition needs.
  • the corresponding relationship between the trained emotion recognition model and the identity of any user can be stored and recorded The emotional state information corresponding to the emotional recognition model.
  • the output result of emotion recognition is less than the second performance threshold, it indicates that the certain trained emotion recognition model does not meet the actual emotion recognition requirements.
  • training samples can be continuously obtained for deep learning.
  • the emotion recognition server calls the voice verification model corresponding to the user based on the identity before calling all the emotion recognition models of the user based on the identity, and the voice verification model uses the voice verification model to Take the voice verification of the signal as an example.
  • voice verification may not be performed on the voice signal, that is, after the emotion recognition server receives the voice signal and the identity, it can directly call all the emotion recognition models of the user based on the identity. Not limited.
  • Step 204 The emotion recognition server determines the current emotional state information of the user through the invoked emotion recognition model based on the voice signal.
  • the emotion recognition server can use the called emotion recognition model of the user to perform emotion recognition on the voice signal of the user to determine the current emotional state of the user.
  • the emotion recognition server determines the current emotional state information of the user through the invoked emotion recognition model based on the voice signal.
  • the specific implementation may include: extracting the voiceprint features of the voice signal, and separately The pattern feature is input to each emotion recognition model in all the called emotion recognition models, and each emotion recognition model recognizes the voiceprint feature and outputs the emotion similarity. Based on all the output emotion similarities, the maximum emotion is determined
  • the emotion recognition model corresponding to the similarity determines the emotion state information corresponding to the determined emotion recognition model as the current emotion state information of the user.
  • each emotion recognition model corresponds to a kind of emotion state information
  • the current emotion state information of the user can be determined according to the recognition results output by each emotion recognition model.
  • the recognition result output by each emotion recognition model is the emotion similarity, that is, it can be judged through which emotion recognition model the speech signal has the greatest emotion similarity.
  • the greater the emotion similarity output by the emotion recognition model the greater the emotion similarity of the speech signal
  • the emotion expressed is closer to the emotion state corresponding to the emotion recognition model, therefore, the emotion recognition model corresponding to the maximum emotion similarity is determined, and the emotion state information corresponding to the determined emotion recognition model is determined as the current emotion state information of the user.
  • the voice signal can be digitized, preprocessed, and endpoint detection processed in sequence, and then the voiceprint feature extraction operation is performed.
  • the emotional recognition server saves the voice signal and the emotional state information as training samples, and the training samples are used to continue training the emotional recognition model corresponding to the emotional state information.
  • the speech signal and recognition result used this time can be used to continue to update the emotion recognition model corresponding to the emotion state information, so that the emotion recognition performance of the emotion recognition model is better. The more accurate.
  • the emotion recognition server when the emotion recognition server includes a voiceprint authentication algorithm server, an emotion recognition algorithm server, and an emotion management library, the voice verification model of each user can be stored in the voiceprint authentication algorithm server, And all the emotion recognition models of each user can be stored in the emotion management library.
  • the voice signal can be verified by the voiceprint authentication algorithm server, and after the verification is passed, the voiceprint authentication algorithm server sends a verification success message to the emotion management database, and further, the verification success message can carry The identity of the user.
  • the emotion management library After the emotion management library receives the verification success message, it obtains all the emotion recognition models of the user based on the identity, and shares them with the emotion recognition algorithm server, which uses all the emotion recognition models shared by the emotion management library, Perform emotion recognition on the user's voice signal.
  • the voice signal and the emotional state information can be stored as training samples in the corresponding emotion management database, so as to continuously collect different users in different emotional states.
  • the voice signal of each user continuously improves the emotion recognition model of each user, thereby increasing the accuracy of emotion recognition.
  • Step 205 The emotion recognition server sends the determined emotion state information to the data management server.
  • the emotion recognition server may send the voice signal and the determined emotion state information to the data management server. That is to say, in addition to sending the determined emotional state information to the data management server, the emotion recognition server can also send the voice signal to the data management server, so that the manager can know the actual status of the monitored person. happening.
  • the emotion recognition server can also send the identification to the data management server.
  • the voice collection device also reports location information, time of occurrence and other information
  • the emotion recognition server can also determine the current location of the user. The information, time of occurrence, and other information are forwarded to the data management server, so that managers can learn more about the user from the data management server.
  • the emotion recognition server may first send the data that needs to be sent to the data management server to the remote monitoring server, and then the remote monitoring server will send the data to the remote monitoring server. Forward to the data management server. In this way, by adding the remote monitoring server, the monitoring area of the monitoring system can be made wider.
  • Step 206 The data management server manages the emotional state information.
  • the specific implementation of the data management server for managing the emotional state information may include: the data management server can determine whether the emotional state corresponding to the emotional state information belongs to an abnormal emotional state, and when it is determined to belong to the abnormal emotional state, An early warning reminder is generated, and the early warning reminder is used to remind the user that the emotional state of the user is abnormal, so that the manager can find the abnormal person in time.
  • the emotion recognition server sends the voice signal and the determined emotional state information to the data management server
  • the data management server manages the voice signal and the emotional state information.
  • the emotion recognition server also sends the voice signal to the data management server
  • the data management server also manages the voice signal, for example, the voice signal can be played.
  • the data management server can also update the user’s emotional state occurrence time, frequency, location and other elements according to the received data.
  • Managers use big data analysis technology to regularly analyze the above The data related to left-behind children on the data management server is separately marked for some children with frequent emotional fluctuations in the short term and fed back to parents and guardians in time.
  • big data analysis is used for the locations or places where the left-behind children are emotionally extreme Increase surveillance measures such as cameras to protect the behavioral safety and mental health of left-behind children from the source.
  • the above description is only an example in which the monitoring system is applied to the monitoring of the emotional state of left-behind children.
  • the monitoring system can also be applied to any scenario that requires emotional monitoring.
  • the implementation of the present disclosure The example does not limit this.
  • the voice collection device collects a voice signal, and sends the voice signal and an identity identifier to the emotion server.
  • the identity identifier is an identifier of a user who has an associated relationship with the voice collection device.
  • the emotion server calls all the emotion recognition models corresponding to the identity, that is, calls all the emotion recognition models of the user. Then, based on the voice signal, the current emotional state information of the user is determined through all the emotion recognition models called, and sent to the data management server for management. That is, the monitoring system can monitor the user's emotional state, which increases the management performance of the monitoring system.
  • Fig. 5 is a schematic structural diagram showing a monitoring device according to an exemplary embodiment.
  • the monitoring device may be configured in an emotion recognition server.
  • the monitoring device may include:
  • the receiving module 510 is configured to receive a voice signal and an identity identifier sent by the voice collection device, the voice signal is collected by the voice collection device, and the identity identifier is a user who has an associated relationship with the voice collection device The logo;
  • the calling module 520 is configured to call all emotion recognition models of the user based on the identity identifier, and each emotion recognition model corresponds to a kind of emotional state information of the user;
  • the determining module 530 is configured to determine the current emotional state information of the user through the invoked emotion recognition model based on the voice signal;
  • the sending module 540 is configured to send the determined emotional state information to the data management server for management.
  • the calling module 520 is used to:
  • all corresponding emotion recognition model identifiers are determined from the stored reference correspondence relationship, and the reference correspondence relationship is used to store the identity identifier of each user in at least one user and all the emotion recognition of each user Correspondence between model identifiers;
  • the determining module 530 is configured to:
  • each emotion recognition model Respectively input the voiceprint feature to each emotion recognition model of all the called emotion recognition models, and each emotion recognition model will recognize the voiceprint feature and output the emotion similarity;
  • the invoking module 520 is also used to:
  • the device further includes:
  • the storage module 550 is configured to save the voice signal and the emotional state information as training samples, and the training samples are used to continue training the emotional recognition model corresponding to the emotional state information.
  • the sending module 540 is used to:
  • the voice signal and the determined emotional state information are sent to the data management server for management.
  • the voice collection device collects a voice signal, and sends the voice signal and an identity identifier to the emotion server.
  • the identity identifier is an identifier of a user who has an associated relationship with the voice collection device.
  • the emotion server calls all the emotion recognition models corresponding to the identity, that is, calls all the emotion recognition models of the user. Then, based on the voice signal, the current emotional state information of the user is determined through all the emotion recognition models called, and sent to the data management server for management. That is, the monitoring system can monitor the user's emotional state, which increases the management performance of the monitoring system.
  • the monitoring device provided in the above embodiment implements the monitoring method
  • only the division of the above functional modules is used as an example for illustration.
  • the above functions can be allocated by different functional modules as needed, namely The internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
  • the monitoring device provided in the foregoing embodiment and the monitoring method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
  • FIG. 7 is a schematic structural diagram of a server 700 provided by an embodiment of the present disclosure.
  • the server 700 may have relatively large differences due to different configurations or performance, and may include one or more processors (central processing units, CPU) 701 and One or more memories 702, wherein at least one instruction is stored in the memory 702, and the at least one instruction is loaded and executed by the processor 701 to implement the monitoring methods provided by the foregoing method embodiments.
  • processors central processing units, CPU
  • memories 702 wherein at least one instruction is stored in the memory 702, and the at least one instruction is loaded and executed by the processor 701 to implement the monitoring methods provided by the foregoing method embodiments.
  • the server 700 may also have components such as a wired or wireless network interface, a keyboard, an input and output interface for input and output, and the server 700 may also include other components for implementing device functions, which will not be repeated here.
  • the embodiments of the present disclosure also provide a non-transitory computer-readable storage medium.
  • the instructions in the storage medium are executed by the processor of the mobile terminal, the mobile terminal can execute the monitoring methods provided by the above-mentioned various embodiments.
  • the embodiments of the present disclosure also provide a computer program product containing instructions, which when run on a computer, cause the computer to execute the monitoring methods provided by the foregoing various embodiments.

Landscapes

  • Engineering & Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Multimedia (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Child & Adolescent Psychology (AREA)
  • General Health & Medical Sciences (AREA)
  • Hospice & Palliative Care (AREA)
  • Psychiatry (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Telephonic Communication Services (AREA)

Abstract

一种监控系统、方法、装置、服务器(700)及存储介质,属于监控技术领域。监控系统包括:语音采集设备(110)、情感识别服务器(120)和数据管理服务器(130),语音采集设备(110),用于采集语音信号,将语音信号和身份标识发送给情感识别服务器(120),身份标识为与语音采集设备(110)具有关联关系的用户的标识;情感识别服务器(120),用于基于身份标识调用用户的所有情感识别模型,每个情感识别模型与用户的一种情感状态信息对应;基于语音信号,通过调用的情感识别模型确定用户当前的情感状态信息;将确定的情感状态信息发送给数据管理服务器;数据管理服务器(130),用于对情感状态信息进行管理。监控系统可以对用户的情感状态进行监控,增加了监控系统的管理性能。

Description

监控系统、方法、装置、服务器及存储介质
本公开要求于2019年03月21日提交的申请号为201910219098.2、发明名称为“监控系统、方法、装置、服务器及存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本公开中。
技术领域
本公开涉及监控技术领域,特别涉及一种监控系统、方法、装置、服务器及存储介质。
背景技术
目前,监控系统在各个领域得到广泛应用,如可以用于对用户、动物等目标进行监控管理,示例性的,可以用于对留守儿童或老年人等弱势群体的监控管理。
在针对弱势群体的应用场景中,监控系统一般包括安全手环和数据管理服务器,该安全手环通常佩戴在用户的手腕上,可以用于获取用户的位置信息,并向该数据管理服务器上报用户的位置信息。该数据管理服务器可以用于存储各个用户的基本信息和各个用户佩戴的安全手环上报的位置信息,以便于管理人员基于这些信息对相关人员进行监护管理。
然而,在上述实现方式中,由于监控系统具备的功能仅仅在于对用户的位置信息和基本信息进行管理,所以,该监控系统的管理性能较差。
发明内容
本公开实施例提供了一种监控系统、方法、装置、服务器及存储介质,可以解决相关技术中监控系统的管理性能较差的问题。所述技术方案如下:
第一方面,提供了一种监控系统,所述系统包括:语音采集设备、情感识别服务器和数据管理服务器,所述情感识别服务器分别与所述语音采集设备和所述数据管理服务器建立有通信连接;
所述语音采集设备,用于采集语音信号,将所述语音信号和身份标识发送 给所述情感识别服务器,所述身份标识为与所述语音采集设备具有关联关系的用户的标识;
所述情感识别服务器,用于基于所述身份标识调用所述用户的所有情感识别模型,每个情感识别模型与所述用户的一种情感状态信息对应;基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息;将所述情感状态信息发送给所述数据管理服务器;
所述数据管理服务器,用于对所述情感状态信息进行管理。
在本公开一种可能的实现方式中,所述情感识别服务器,用于基于所述身份标识调用所述用户的所有情感识别模型,包括:
所述情感识别服务器,用于基于所述身份标识,从存储的参考对应关系中确定对应的所有情感识别模型标识,所述参考对应关系用于存储至少一个用户中每个用户的身份标识与所述每个用户的所有情感识别模型标识之间的对应关系;调用所确定的所有情感识别模型标识对应的情感识别模型。
在本公开一种可能的实现方式中,所述情感识别服务器,用于基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息,包括:
所述情感识别服务器,用于提取所述语音信号的声纹特征;分别将所述声纹特征输入至调用的所有情感识别模型中的每个情感识别模型,由所述每个情感识别模型对所述声纹特征进行识别处理并输出情感相似度;基于输出的所有情感相似度,确定最大情感相似度对应的情感识别模型;将确定的情感识别模型对应的情感状态信息确定为所述用户当前的情感状态信息。
在本公开一种可能的实现方式中,所述情感识别服务器,用于基于所述身份标识调用所述用户的所有情感识别模型之前,还用于:
所述情感识别服务器,用于基于所述身份标识,调用所述用户对应的语音验证模型;通过所述语音验证模型对所述语音信号进行语音验证;对应的,当对所述语音信号验证通过时,执行所述基于所述身份标识调用所述用户的所有情感识别模型的操作。
在本公开一种可能的实现方式中,所述情感识别服务器,用于基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息之后,还用于:
所述情感识别服务器,用于将所述语音信号和所述情感状态信息保存为训练样本,所述训练样本用于对所述情感状态信息对应的情感识别模型继续训练。
在本公开一种可能的实现方式中,所述情感识别服务器,用于将所述情感状态信息发送给所述数据管理服务器,包括:
所述情感识别服务器,用于将所述语音信号和所述情感状态信息发送给所述数据管理服务器;
对应的,所述数据管理服务器,用于对所述情感状态信息进行管理,包括:
所述数据管理服务器,用于对所述语音信号和所述情感状态信息进行管理。
第二方面,提供了一种监控方法,所述方法应用于监控系统的情感识别服务器中,所述监控系统还包括语音采集设备和数据管理服务器,所述情感识别服务器分别与所述语音采集设备和所述数据管理服务器建立有通信连接;所述方法包括:
接收所述语音采集设备发送的语音信号和身份标识,所述语音信号是由所述语音采集设备采集得到,所述身份标识为与所述语音采集设备具有关联关系的用户的标识;
基于所述身份标识调用所述用户的所有情感识别模型,每个情感识别模型与所述用户的一种情感状态信息对应;
基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息;
将确定的情感状态信息发送给所述数据管理服务器进行管理。
在本公开一种可能的实现方式中,所述基于所述身份标识调用所述用户的所有情感识别模型,包括:
基于所述身份标识,从存储的参考对应关系中确定对应的所有情感识别模型标识,所述参考对应关系用于存储至少一个用户中每个用户的身份标识与所述每个用户的所有情感识别模型标识之间的对应关系;
调用所确定的所有情感识别模型标识对应的情感识别模型。
在本公开一种可能的实现方式中,所述基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息,包括:
提取所述语音信号的声纹特征;
分别将所述声纹特征输入至调用的所有情感识别模型中的每个情感识别模型,由所述每个情感识别模型对所述声纹特征进行识别处理并输出情感相似度;
基于输出的所有情感相似度,确定最大情感相似度对应的情感识别模型;
将确定的情感识别模型对应的情感状态信息确定为所述用户当前的情感状 态信息。
在本公开一种可能的实现方式中,所述基于所述身份标识调用所述用户的所有情感识别模型之前,还包括:
基于所述身份标识,调用所述用户对应的语音验证模型;
通过所述语音验证模型对所述语音信号进行语音验证;
对应的,当对所述语音信号验证通过时,执行所述基于所述身份标识调用所述用户的所有情感识别模型的操作。
在本公开一种可能的实现方式中,所述基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息之后,还包括:
将所述语音信号和所述情感状态信息保存为训练样本,所述训练样本用于对所述情感状态信息对应的情感识别模型继续训练。
在本公开一种可能的实现方式中,所述将确定的情感状态信息发送给所述数据管理服务器,包括:
将所述语音信号和所确定的情感状态信息发送给所述数据管理服务器进行管理。
第三方面,提供了一种监控装置,配置于监控系统的情感识别服务器中,所述监控系统还包括语音采集设备和数据管理服务器,所述情感识别服务器分别与所述语音采集设备和所述数据管理服务器建立通信连接;所述装置包括:
接收模块,用于接收所述语音采集设备发送的语音信号和身份标识,所述语音信号是由所述语音采集设备采集得到,所述身份标识为与所述语音采集设备具有关联关系的用户的标识;
调用模块,用于基于所述身份标识调用所述用户的所有情感识别模型,每个情感识别模型与所述用户的一种情感状态信息对应;
确定模块,用于基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息;
发送模块,用于将确定的情感状态信息发送给所述数据管理服务器进行管理。
在本公开一种可能的实现方式中,所述调用模块用于:
基于所述身份标识,从存储的参考对应关系中确定对应的所有情感识别模型标识,所述参考对应关系用于存储至少一个用户中每个用户的身份标识与所述每个用户的所有情感识别模型标识之间的对应关系;
调用所确定的所有情感识别模型标识对应的情感识别模型。
在本公开一种可能的实现方式中,所述确定模块用于:
提取所述语音信号的声纹特征;
分别将所述声纹特征输入至调用的所有情感识别模型中的每个情感识别模型,由所述每个情感识别模型对所述声纹特征进行识别处理并输出情感相似度;
基于输出的所有情感相似度,确定最大情感相似度对应的情感识别模型;
将确定的情感识别模型对应的情感状态信息确定为所述用户当前的情感状态信息。
在本公开一种可能的实现方式中,所述调用模块还用于:
基于所述身份标识,调用所述用户对应的语音验证模型;
通过所述语音验证模型对所述语音信号进行语音验证;
当对所述语音信号验证通过时,基于所述身份标识调用所述用户的所有情感识别模型。
在本公开一种可能的实现方式中,所述装置还包括:
存储模块,用于将所述语音信号和所述情感状态信息保存为训练样本,所述训练样本用于对所述情感状态信息对应的情感识别模型继续训练。
在本公开一种可能的实现方式中,所述发送模块用于:
将所述语音信号和所确定的情感状态信息发送给所述数据管理服务器进行管理。
第四方面,提供了一种情感识别服务器,包括:
处理器;
用于存储处理器可执行指令的存储器;
其中,所述处理器被配置为实现上述第二方面所述的监控方法。
第五方面,提供了一种计算机可读存储介质,所述计算机可读存储介质上存储有指令,所述指令被处理器执行时实现上述第二方面所述的监控方法。
第六方面,提供了一种包含指令的计算机程序产品,当其在计算机上运行时,使得计算机执行上述第二方面所述的监控方法。
本公开实施例提供的技术方案带来的有益效果是:
语音采集设备采集语音信号,并将该语音信号和身份标识发送给情感服务器,该身份标识为与该语音采集设备具有关联关系的用户的标识。该情感服务 器调用该身份标识对应的所有情感识别模型,即调用该用户的所有情感识别模型。然后基于该语音信号,通过调用的所有情感识别模型确定用户当前的情感状态信息,并发给数据管理服务器进行管理。即该监控系统可以对用户的情感状态进行监控,增加了监控系统的管理性能。
附图说明
为了更清楚地说明本公开实施例中的技术方案,下面将对实施例描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本公开的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。
图1是根据一示例性实施例示出的一种监控系统的框架图;
图2是根据一示例性实施例示出的一种监控方法的流程图;
图3是根据另一示例性实施例示出的一种语音验证的原理示意图;
图4根据另一示例性实施例示出的一种情感识别的基本原理示意图;
图5是根据一示例性实施例示出的一种监控装置的结构示意图;
图6是根据另一示例性实施例示出的一种监控装置的结构示意图;
图7是根据一示例性实施例示出的一种服务器700的结构示意图。
具体实施方式
为使本公开的目的、技术方案和优点更加清楚,下面将结合附图对本公开实施方式作进一步地详细描述。
首先,对本公开实施例提供的应用场景进行简单介绍。
近几年,随着大数据和互联网等信息技术的发展创新以及政府对留守儿童的心理和人身安全的逐渐重视,部分地区引入“留守儿童工作大数据平台”,该“留守儿童工作大数据平台”实际也是一种监控系统,一般包括有安全手环和数据管理服务器,该数据管理服务器可以归属于管理平台。在实施中,管理人员可以为中小学阶段的留守儿童配发安全手环,以通过安全手环向该数据管理服务器上报留守儿童的位置信息,从而实现儿童监控信息与该管理平台的无缝对接。然而,目前的监控系统仅仅具备对留守儿童的基本信息管理(如对哪些基本信息不完整进行提醒等)、GPS(Global Positioning System,全球定位系统)定位等功能,且留守儿童佩戴的安全手环的功能有限,仅仅可用于定位、存储 基本信息,但很难对留守儿童的心理和人身健康作更具针对性的动态监测,更无法及时有效地监测儿童的情感动向。为此,本公开实施例提供了一种监控系统,该监控系统可以对儿童的情感动向进行监测和管理,增加了监控系统的管理性能。其具体实现请参见如下各个实施例。
接下来,请参考图1,该图1是根据一示例性实施例示出的一种监控系统的框架图,该监控系统主要包括:语音采集设备110、情感识别服务器120和数据管理服务器130,该情感识别服务器120分别与语音采集设备110和数据管理服务器130建立有通信连接。
其中,该语音采集设备110具有语音采集功能,可以用于采集用户的语音信号,并将该语音信号发送给情感识别服务器120。在一些实施例中,该语音采集设备110可以配置有可佩戴部件,以便于用户可以利用该可佩戴部件将其佩戴在身上。或者,该语音采集设备110也可以为配置有语音采集器的可穿戴设备,譬如,可以为配置有语音采集器的手环、手表等,本公开实施例对此不限定。另外,该语音采集设备110除了具有语音采集功能外,还可以具有定位等功能。
其中,该情感识别服务器120可以用于基于用户的语音信号对该用户进行情感识别,以确定用户当前的情感状态。在一些实施例中,该情感识别服务器120可以为一台服务器,或者,该情感识别服务器120还可以为由多台服务器组成的服务器集群,比如,请继续参考图1,该情感识别服务器120可以包括声纹认证算法服务器120a、情感识别算法服务器120b和情感管理库120c。其中,该声纹认证算法服务器120a可以用于对用户的语音信号进行语音验证,该情感管理库120c存储有多个用户中每个用户的所有情感识别模型,如此,该声纹认证算法服务器120a在语音验证通过后,触发该情感管理库120c获取用户的情感识别模型。该情感管理库120c将获取的情感识别模型共享给该情感识别算法服务器120b,以便于该情感识别算法服务器120b使用情感管理库120c共享的情感识别模型,对用户的情感状态进行识别,并将识别出的情感状态信息发送给数据管理服务器130。
其中,该数据管理服务器130可以用于对情感状态信息进行动态管理,得到管理信息,以便于管理人员可以根据数据管理服务器130中的管理信息,及时对用户的情感动向进行监控。在一些实施例中,该数据管理服务器130可以为一台服务器,或者,也可以为由多台服务器组成的服务器集群,本公开实施 例对此不做限定。
进一步地,请参考图1,该监控系统还可以包括虚拟服务器140,该虚拟服务器140分别与语音采集设备110和情感识别服务器120连接,以将该语音采集设备110传输的语音信号传递给情感识别服务器120。在一些实施例中,该虚拟服务器140可以称为交换机。
另外,该监控系统还可以包括远程监控服务器150,该远程监控服务器150分别与情感识别服务器120和数据管理服务器130连接,以将该情感识别服务器120传输的数据传递给数据管理服务器130。
如此,该监控系统通过该虚拟服务器140和该远程监控服务器150的数据传递,可以实现更远距离的监控,即可以使得监控系统能够实现更广范围的监控。
接下来将结合图1所示的监控系统,对监控系统的监控过程进行详细介绍。请参考图2,该图2是根据一示例性实施例示出的一种监控方法的流程图,该监控方法可以包括如下几个步骤:
步骤201:语音采集设备采集语音信号,将该语音信号和身份标识发送给该情感识别服务器,该身份标识为与语音采集设备具有关联关系的用户的标识。
其中,与语音采集设备具有关联关系的用户可以是指使用该语音采集设备的用户,或者,也可以是指该语音采集设备的拥有者等。另外,该身份标识可以用于唯一的标识一个用户。
以该监控系统应用于对某个村的留守儿童进行监控管理为例,在该种应用场景中,管理人员可以为每个留守儿童配发语音采集设备,比如,该语音采集设备可以为嵌有语音采集器的安全手环等,以通过该语音采集设备采集语音信号。
在一种可能的实现方式中,该语音采集设备可以实时执行采集操作,在另一种可能的实现方式中,该语音采集设备也可以每隔参考时长进行一次采集操作。其中,该参考时长可以由用户根据实际需求进行设置,或者,也可以由该语音采集设备默认设置,本公开实施例对此不作限定。
该语音采集设备采集语音信号后,将该语音信号和该身份标识发送给该情感识别服务器,以便于该情感识别服务器对该用户的情感状态进行识别。
进一步地,请参考图1,当该监控系统还包括虚拟服务器时,该语音采集设 备将采集的语音信号和该身份标识发送给虚拟服务器,由该虚拟服务器将该语音信号和该身份标识转发给该情感识别服务器。
进一步地,该虚拟服务器可以确定该语音信号的频谱能量,当该频谱能量大于或等于频谱能量阈值时,将该语音信号和该身份标识转发给情感识别服务器,否则,当该频谱能量小于频谱能量阈值时,可以不将该语音信号和该身份标识转发给情感识别服务器。
其中,该频谱能量阈值可以由用户根据实际需求自定义设置,也可以由该虚拟服务器默认设置,本公开实施例对此不做限定。
也就是说,当该监控系统还包括虚拟服务器时,可以由该虚拟服务器抉择是否将该语音采集设备采集的语音信号和身份标识发送给情感识别服务器。一般来说,不同情感表达的语音信号在其时间构造、振幅构造、基频构造和共振峰构造等特征上也有着不同的构造特点和分布规律。当该语音信号的频谱能量大于或等于频谱能量阈值时,一般可以说明该用户说话语气不平稳,从而可以说明该用户情绪可能比较激动,在该种情况下说明该用户可能存在异常,所以,虚拟服务器将该语音信号和身份标识发送给情感识别服务器作进一步情感识别。而当该语音信号的频谱能量小于频谱能量阈值时,一般可以说明该用户说话语气比较平稳,从而可以说明该用户情绪可能比较稳定,在该种情况下说明该用户不存在异常,因此,可以不将该语音信号和该身份标识转发给情感识别服务器,即可以丢弃该语音信号,并继续等待或处理语音采集设备发送的下一个语音信号和身份标识等数据,如此可以减小情感识别服务器的运算量。
步骤202:情感识别服务器基于该身份标识,调用该用户对应的语音验证模型,通过该语音验证模型对该语音信号进行语音验证。
在一种可能的实现方式中,语音采集设备发送至情感识别服务器的语音信号可能并不是与该语音采集设备具有关联关系的用户的,譬如,某个留守儿童A佩戴的语音采集设备发送给情感识别服务器的语音信号可能来自与留守儿童A发生争执的留守儿童B。针对该种情况,为了避免监控管理出错,该情感识别服务器接收该语音采集设备发送的语音信号和身份标识后,可以基于该身份标识对语音信号进行验证,即验证该语音信号是否属于该用户的。
在一种可能的实现方式中,该情感识别服务器可以预先存储有各个用户的身份标识与语音验证模型之间的对应关系,即每个用户可以对应一个语音验证模型。如此,情感识别服务器可以基于该身份标识,调用该用户对应的语音验 证模型,并使用该语音验证模型进行语音验证。在实施中,该情感识别服务器可以提取该语音信号的声纹特征,将该声纹特征输入至该语音验证模型中,由该语音验证模型进行验证处理,输出语音相似度。当该语音相似度大于或等于语音相似度阈值时,可以确定该语音信号验证通过,否则,当该语音相似度小于该语音相似度阈值时,可以确定该语音信号验证未通过。
其中,语音相似度阈值可以由用户根据实际需求自定义设置,也可以由情感识别服务器默认设置,本公开实施例对此不做限定。
其中,当语音相似度大于或等于语音相似度阈值时,说明该语音信号就是来自该用户的,即属于该用户,所以,可以确定该语音信号验证通过。当该语音相似度小于该语音相似度阈值时,说明该语音信号不是来自该用户的,即不属于该用户,在该种情况下,可以确定该语音信号验证未通过。
值得一提的是,在接收到语音信号后,该情感识别服务器先对该语音信号进行语音验证,以确定该语音信号是否真正属于该用户,如此可以提高监控系统的管理的准确性。
进一步地,每个用户的语音验证模型可以预先通过训练得到,具体地,每个用户的语音验证模型可以基于每个用户的大量的训练样本对待训练的网络模型进行训练得到。譬如,请参考图3,在实施中,针对任一用户,建立待训练的语音验证模型,获取该任一用户的多个语音片段,提取该多个语音片段中每个语音片段的声纹特征,将提取的声纹特征作为训练样本输入至待训练的语音验证模型中进行深度学习,得到该任一用户对应的训练后的语音验证模型。按照该种实现方式,可以确定每个用户对应的训练后的语音验证模型。
进一步地,可以使用该每个用户的测试样本对训练后的语音验证模型进行性能评估。在实施中,对于任一用户来说,可以提取该任一用户的测试样本的声纹特征,将该声纹特征输入至任一用户对应的训练后的语音验证模型中,当语音验证输出结果大于或等于第一性能阈值时,确定该训练后的语音验证模型符合实际验证需求,在该种情况下,可以存储该训练后的语音验证模型与该任一用户的身份标识之间的对应关系。反之,当该语音验证输出结果小于该第一性能阈值时,说明该训练后的语音验证模型不符合实际验证需求,在该种情况下,可以继续获取该任一用户的训练样本对该验证后的语音验证模型进行深度学习。如此,按照该种实现方式,可以确定每个用户的语音验证模型。
其中,该第一性能阈值可以由用户根据实际需求自定义设置,也可以由该 情感识别服务器默认设置,本公开实施例对此不做限定。
步骤203:当对该语音信号验证通过时,情感识别服务器基于该身份标识调用该用户的所有情感识别模型,每个情感识别模型与该用户的一种情感状态对应。
当对该语音信号验证通过时,说明该语音信号确实是来自于与该语音采集设备具有关联关系的用户,此时该情感识别服务器基于该身份标识调用该用户的所有情感识别模型。
在一种可能的实现方式中,基于该身份标识调用该用户的所有情感识别模型的具体实现可以包括:基于该身份标识,从存储的参考对应关系中确定对应的所有情感识别模型标识,该参考对应关系用于存储多个用户中每个用户的身份标识与该每个用户的所有情感识别模型标识之间的对应关系;调用所确定的所有情感识别模型标识对应的情感识别模型。
其中,每个情感识别模型标识可以用于唯一的标识一种情感识别模型,每个用户可以对应有一个或者多个情感识别模型,每种情感识别模型对应该用户的一种情感状态信息。譬如,假设每个用户的情感识别模型可以包括第一情感识别模型、第二情感识别模型和第三情感识别模型,该第一情感识别模型对应的情感状态信息可以为“恐惧”,该第二情感识别模型对应的情感状态信息可以为“恸哭”,该第三情感识别模型对应的情感状态信息可以为“悲愤”等。
该情感识别服务器可以预先存储有每个用户的所有情感识别模型,并存储有每个用户的身份标识与该用户对应的所有情感识别模型标识之间的参考对应关系,如此,情感识别服务器即可基于用户的身份标识和参考对应关系,调用用户的所有情感识别模型。譬如,调用的该用户的所有情感识别模型包括第一情感识别模型、第二情感识别模型和第三情感识别模型。
进一步地,每个用户的每种情感识别模型均可以预先通过训练得到,具体地,每个用户的每种情感识别模型可以基于大量的训练样本对待训练的网络模型进行训练得到。譬如,请参考图4,在实施中,对于任一用户来说,针对该任一用户的任一种情感状态,建立待训练的情感识别模型,获取该任一用户在该任一种情感状态下的多个语音片段,对该多个语音片段中的每个语音片段依次进行数字化及预处理、端点检测处理、特征提取处理,得到每个语音片段的声纹特征,将提取的声纹特征作为训练数据输入至待训练的情感识别模型中进行深度学习,得到该任一种情感状态对应的情感识别模型,之后,可以记录该任 一种情感状态的情感状态信息与该情感识别模型的对应关系。如此,按照该种实现方式,可以确定该任一用户的每种情感状态信息对应的情感识别模型。
进一步地,对于任一用户来说,可以使用该任一用户的测试样本对训练后的情感识别模型进行性能评估。譬如,针对该任一用户对应的某个训练后的情感识别模型,可以获取该某个训练后的情感识别模型对应的情感状态下的测试样本,对获取的测试样本依次进行数字化及预处理、端点检测处理、特征提取处理,得到该测试样本的声纹特征,将该声纹特征输入至该某个训练后的情感识别模型中,当情感识别的输出结果大于或等于第二性能阈值时,确定该某个训练后的情感识别模型符合实际的情感识别需求,在该种情况下,可以存储该某个训练后的情感识别模型与该任一用户的身份标识之间的对应关系,并记录该情感识别模型对应的情感状态信息。反之,当情感识别的输出结果小于该第二性能阈值时,说明该某个训练后的情感识别模型不符合实际的情感识别需求,在该种情况下,可以继续获取训练样本进行深度学习。
需要说明的是,本公开实施例是以情感识别服务器在基于该身份标识调用该用户的所有情感识别模型之前,基于该身份标识调用该用户对应的语音验证模型,通过该语音验证模型对该语音信号进行语音验证为例进行说明。在另一实施例中,也可以不对语音信号进行语音验证,即情感识别服务器接收到语音信号和身份标识后,可以直接基于该身份标识调用该用户的所有情感识别模型,本公开实施例对此不作限定。
步骤204:情感识别服务器基于该语音信号,通过调用的情感识别模型确定该用户当前的情感状态信息。
也即是,情感识别服务器可以使用所调用的该用户的情感识别模型,对该用户的语音信号进行情感识别,以确定该用户当前是怎样的情感状态。
在一种可能的实现方式中,情感识别服务器基于该语音信号,通过调用的情感识别模型确定该用户当前的情感状态信息的具体实现可以包括:提取该语音信号的声纹特征,分别将该声纹特征输入至调用的所有情感识别模型中的每个情感识别模型,由该每个情感识别模型对该声纹特征进行识别处理并输出情感相似度,基于输出的所有情感相似度,确定最大情感相似度对应的情感识别模型,将确定的情感识别模型对应的情感状态信息确定为该用户当前的情感状态信息。
由于每种情感识别模型对应一种情感状态信息,因此,可以根据每个情感 识别模型输出的识别结果来确定该用户当前的情感状态信息。其中,该每个情感识别模型输出的识别结果为情感相似度,即可以判断该语音信号通过哪个情感识别模型输出的情感相似度最大,情感识别模型输出的情感相似度越大,说明该语音信号所表达的情感与该情感识别模型对应的情感状态越接近,因此,确定最大情感相似度对应的情感识别模型,将确定的情感识别模型对应的情感状态信息确定为该用户当前的情感状态信息。
进一步地,在提取该语音信号的声纹特征之前,还可以对该语音信号依次进行数字化及预处理、端点检测处理,然后执行提取声纹特征的操作。
进一步地,确定该用户当前的情感状态信息之后,情感识别服务器将该语音信号和该情感状态信息保存为训练样本,该训练样本用于对该情感状态信息对应的情感识别模型继续训练。
也就是说,在每次情感识别结束后,还可以利用本次使用的语音信号和识别结果,继续对该情感状态信息对应的情感识别模型进行更新,以使得该情感识别模型的情感识别性能越来越准确。
在本公开的一种可能实现方式中,当该情感识别服务器包括声纹认证算法服务器、情感识别算法服务器和情感管理库时,每个用户的语音验证模型可以存储在声纹认证算法服务器中,且每个用户的所有情感识别模型可以存储在该情感管理库中。在该种情况下,可以通过声纹认证算法服务器对语音信号进行验证,并在验证通过后,声纹认证算法服务器向该情感管理库发送验证成功消息,进一步,该验证成功消息中可以携带有该用户的身份标识。该情感管理库接收到该验证成功消息后,基于该身份标识获取该用户的所有情感识别模型,并分享给情感识别算法服务器,该情感识别算法服务器通过该情感管理库分享的所有情感识别模型,对该用户的语音信号进行情感识别。
进一步地,情感识别算法服务器确定该用户当前的情感状态信息之后,可以将该语音信号和该情感状态信息作为训练样本存储至对应的情感管理库中,以通过不断收集不同用户在不同情感状态下的语音信号,不断完善每个用户的情感识别模型,从而增加情感识别的准确率。
步骤205:情感识别服务器将确定的情感状态信息发送给该数据管理服务器。
进一步地,该情感识别服务器可以将该语音信号和确定的情感状态信息发送给该数据管理服务器。也就是说,该情感识别服务器除了将确定的情感状态 信息发送给该数据管理服务器外,还可以将该语音信号也一同发送给该数据管理服务器,以便于管理人员可以获知所监控的人员的实际情况。
进一步地,情感识别服务器还可以将该身份标识发送给该数据管理服务器,另外,当该语音采集设备还上报了位置信息、发生时间等信息时,该情感识别服务器还可以将该用户当前的位置信息、发生时间等信息转发给该数据管理服务器,以便于管理人员可以从该数据管理服务器中获知更多关于该用户的信息。
在一种可能的实现方式中,当该监控系统包括远程监控服务器时,该情感识别服务器可以将需要发送给数据管理服务器的数据先发送至远程监控服务器,然后,再由该远程监控服务器将数据转发给数据管理服务器。如此,通过增加该远程监控服务器,可以使得该监控系统的监控区域更广泛。
步骤206:数据管理服务器对该情感状态信息进行管理。
作为一种示例,数据管理服务器对该情感状态信息进行管理的具体实现可以包括:该数据管理服务器可以判断该情感状态信息对应的情感状态是否属于异常情感状态,当确定属于异常情感状态时,可以生成预警提示,该预警提示用于提示该用户的情感状态异常,以便于管理人员及时发现异常人员。
进一步地,当情感识别服务器将该语音信号和该确定的情感状态信息发送给该数据管理服务器时,该数据管理服务器对该语音信号和该情感状态信息进行管理。
也即是,当该情感识别服务器还将该语音信号发送给该数据管理服务器时,该数据管理服务器还对该语音信号进行管理,譬如,可以播放该语音信号等。
在一种可能的实现方式中,数据管理服务器还可以根据所接收的数据,更新该用户的情绪状态发生时间、频度及发生地点等要素的记录,管理人员利用大数据分析技术定期梳理分析以上数据管理服务器上的留守儿童相关数据,对于一些短期内情绪波动较为频繁的儿童信息单独标记并及时反馈给家长及监护人,当然对于留守儿童情绪极端所在的位置或场所通过大数据分析后,可适当增加摄像头等监管措施,从源头上保护留守儿童的行为安全和心理健康。
需要说明的是,上述仅是以该监控系统应用于对留守儿童的情感状态进行监控的场景中为例进行说明,此外,该监控系统还可以应用于任何需要情感监控的场景中,本公开实施例对此不做限定。
在本公开实施例中,语音采集设备采集语音信号,并将该语音信号和身份标识发送给情感服务器,该身份标识为与该语音采集设备具有关联关系的用户 的标识。该情感服务器调用该身份标识对应的所有情感识别模型,即调用该用户的所有情感识别模型。然后基于该语音信号,通过调用的所有情感识别模型确定用户当前的情感状态信息,并发给数据管理服务器进行管理。即该监控系统可以对用户的情感状态进行监控,增加了监控系统的管理性能。
图5是根据一示例性实施例示出的一种监控装置的结构示意图,该监控装置可以配置于情感识别服务器中。该监控装置可以包括:
接收模块510,用于接收所述语音采集设备发送的语音信号和身份标识,所述语音信号是由所述语音采集设备采集得到,所述身份标识为与所述语音采集设备具有关联关系的用户的标识;
调用模块520,用于基于所述身份标识调用所述用户的所有情感识别模型,每个情感识别模型与所述用户的一种情感状态信息对应;
确定模块530,用于基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息;
发送模块540,用于将确定的情感状态信息发送给所述数据管理服务器进行管理。
在本公开一种可能的实现方式中,所述调用模块520用于:
基于所述身份标识,从存储的参考对应关系中确定对应的所有情感识别模型标识,所述参考对应关系用于存储至少一个用户中每个用户的身份标识与所述每个用户的所有情感识别模型标识之间的对应关系;
调用所确定的所有情感识别模型标识对应的情感识别模型。
在本公开一种可能的实现方式中,所述确定模块530用于:
提取所述语音信号的声纹特征;
分别将所述声纹特征输入至调用的所有情感识别模型中的每个情感识别模型,由所述每个情感识别模型对所述声纹特征进行识别处理并输出情感相似度;
基于输出的所有情感相似度,确定最大情感相似度对应的情感识别模型;
将确定的情感识别模型对应的情感状态信息确定为所述用户当前的情感状态信息。
在本公开一种可能的实现方式中,所述调用模块520还用于:
基于所述身份标识,调用所述用户对应的语音验证模型;
通过所述语音验证模型对所述语音信号进行语音验证;
当对所述语音信号验证通过时,基于所述身份标识调用所述用户的所有情感识别模型。
在本公开一种可能的实现方式中,请参考图6,所述装置还包括:
存储模块550,用于将所述语音信号和所述情感状态信息保存为训练样本,所述训练样本用于对所述情感状态信息对应的情感识别模型继续训练。
在本公开一种可能的实现方式中,所述发送模块540用于:
将所述语音信号和所确定的情感状态信息发送给所述数据管理服务器进行管理。
在本公开实施例中,语音采集设备采集语音信号,并将该语音信号和身份标识发送给情感服务器,该身份标识为与该语音采集设备具有关联关系的用户的标识。该情感服务器调用该身份标识对应的所有情感识别模型,即调用该用户的所有情感识别模型。然后基于该语音信号,通过调用的所有情感识别模型确定用户当前的情感状态信息,并发给数据管理服务器进行管理。即该监控系统可以对用户的情感状态进行监控,增加了监控系统的管理性能。
需要说明的是:上述实施例提供的监控装置在实现监控方法时,仅以上述各功能模块的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能模块完成,即将设备的内部结构划分成不同的功能模块,以完成以上描述的全部或者部分功能。另外,上述实施例提供的监控装置与监控方法实施例属于同一构思,其具体实现过程详见方法实施例,这里不再赘述。
图7是本公开实施例提供的一种服务器700的结构示意图,该服务器700可因配置或性能不同而产生比较大的差异,可以包括一个或一个以上处理器(central processing units,CPU)701和一个或一个以上的存储器702,其中,所述存储器702中存储有至少一条指令,所述至少一条指令由所述处理器701加载并执行以实现上述各个方法实施例提供的监控方法。
当然,该服务器700还可以具有有线或无线网络接口、键盘以及输入输出接口等部件,以便进行输入输出,该服务器700还可以包括其他用于实现设备功能的部件,在此不做赘述。
本公开实施例还提供了一种非临时性计算机可读存储介质,当所述存储介质中的指令由移动终端的处理器执行时,使得移动终端能够执行上述各个实施 例提供的监控方法。
本公开实施例还提供了一种包含指令的计算机程序产品,当其在计算机上运行时,使得计算机执行上述各个实施例提供的监控方法。
本领域普通技术人员可以理解实现上述实施例的全部或部分步骤可以通过硬件来完成,也可以通过程序来指令相关的硬件完成,所述的程序可以存储于一种计算机可读存储介质中,上述提到的存储介质可以是只读存储器,磁盘或光盘等。
以上所述仅为本公开的可选实施例,并不用以限制本公开,凡在本公开的精神和原则之内,所作的任何修改、等同替换、改进等,均应包含在本公开的保护范围之内。

Claims (20)

  1. 一种监控系统,其特征在于,所述系统包括:语音采集设备、情感识别服务器和数据管理服务器,所述情感识别服务器分别与所述语音采集设备和所述数据管理服务器建立有通信连接;
    所述语音采集设备,用于采集语音信号,将所述语音信号和身份标识发送给所述情感识别服务器,所述身份标识为与所述语音采集设备具有关联关系的用户的标识;
    所述情感识别服务器,用于基于所述身份标识调用所述用户的所有情感识别模型,每个情感识别模型与所述用户的一种情感状态信息对应;基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息;将所述情感状态信息发送给所述数据管理服务器;
    所述数据管理服务器,用于对所述情感状态信息进行管理。
  2. 如权利要求1所述的系统,其特征在于,所述情感识别服务器,用于基于所述身份标识调用所述用户的所有情感识别模型,包括:
    所述情感识别服务器,用于基于所述身份标识,从存储的参考对应关系中确定对应的所有情感识别模型标识,所述参考对应关系用于存储至少一个用户中每个用户的身份标识与所述每个用户的所有情感识别模型标识之间的对应关系;调用所确定的所有情感识别模型标识对应的情感识别模型。
  3. 如权利要求1或2所述的系统,其特征在于,所述情感识别服务器,用于基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息,包括:
    所述情感识别服务器,用于提取所述语音信号的声纹特征;分别将所述声纹特征输入至调用的所有情感识别模型中的每个情感识别模型,由所述每个情感识别模型对所述声纹特征进行识别处理并输出情感相似度;基于输出的所有情感相似度,确定最大情感相似度对应的情感识别模型;将确定的情感识别模型对应的情感状态信息确定为所述用户当前的情感状态信息。
  4. 如权利要求1所述的系统,其特征在于,所述情感识别服务器,用于基 于所述身份标识调用所述用户的所有情感识别模型之前,还用于:
    所述情感识别服务器,用于基于所述身份标识,调用所述用户对应的语音验证模型;通过所述语音验证模型对所述语音信号进行语音验证;对应的,当对所述语音信号验证通过时,执行所述基于所述身份标识调用所述用户的所有情感识别模型的操作。
  5. 如权利要求1所述的系统,其特征在于,所述情感识别服务器,用于基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息之后,还用于:
    所述情感识别服务器,用于将所述语音信号和所述情感状态信息保存为训练样本,所述训练样本用于对所述情感状态信息对应的情感识别模型继续训练。
  6. 如权利要求1所述的系统,其特征在于,所述情感识别服务器,用于将所述情感状态信息发送给所述数据管理服务器,包括:
    所述情感识别服务器,用于将所述语音信号和所述情感状态信息发送给所述数据管理服务器;
    对应的,所述数据管理服务器,用于对所述情感状态信息进行管理,包括:
    所述数据管理服务器,用于对所述语音信号和所述情感状态信息进行管理。
  7. 一种监控方法,其特征在于,所述方法应用于监控系统的情感识别服务器中,所述监控系统还包括语音采集设备和数据管理服务器,所述情感识别服务器分别与所述语音采集设备和所述数据管理服务器建立有通信连接,所述方法包括:
    接收所述语音采集设备发送的语音信号和身份标识,所述语音信号是由所述语音采集设备采集得到,所述身份标识为与所述语音采集设备具有关联关系的用户的标识;
    基于所述身份标识调用所述用户的所有情感识别模型,每个情感识别模型与所述用户的一种情感状态信息对应;
    基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息;
    将确定的情感状态信息发送给所述数据管理服务器进行管理。
  8. 如权利要求7所述的方法,其特征在于,所述基于所述身份标识调用所述用户的所有情感识别模型,包括:
    基于所述身份标识,从存储的参考对应关系中确定对应的所有情感识别模型标识,所述参考对应关系用于存储至少一个用户中每个用户的身份标识与所述每个用户的所有情感识别模型标识之间的对应关系;
    调用所确定的所有情感识别模型标识对应的情感识别模型。
  9. 如权利要求7或8所述的方法,其特征在于,所述基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息,包括:
    提取所述语音信号的声纹特征;
    分别将所述声纹特征输入至调用的所有情感识别模型中的每个情感识别模型,由所述每个情感识别模型对所述声纹特征进行识别处理并输出情感相似度;
    基于输出的所有情感相似度,确定最大情感相似度对应的情感识别模型;
    将确定的情感识别模型对应的情感状态信息确定为所述用户当前的情感状态信息。
  10. 如权利要求7所述的方法,其特征在于,所述基于所述身份标识调用所述用户的所有情感识别模型之前,还包括:
    基于所述身份标识,调用所述用户对应的语音验证模型;
    通过所述语音验证模型对所述语音信号进行语音验证;
    对应的,当对所述语音信号验证通过时,执行所述基于所述身份标识调用所述用户的所有情感识别模型的操作。
  11. 如权利要求7所述的方法,其特征在于,所述基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息之后,还包括:
    将所述语音信号和所述情感状态信息保存为训练样本,所述训练样本用于对所述情感状态信息对应的情感识别模型继续训练。
  12. 如权利要求7所述的方法,其特征在于,所述将确定的情感状态信息发送给所述数据管理服务器,包括:
    将所述语音信号和所确定的情感状态信息发送给所述数据管理服务器进行管理。
  13. 一种监控装置,其特征在于,配置于监控系统的情感识别服务器中,所述监控系统还包括语音采集设备和数据管理服务器,所述情感识别服务器分别与所述语音采集设备和所述数据管理服务器建立通信连接;所述装置包括:
    接收模块,用于接收所述语音采集设备发送的语音信号和身份标识,所述语音信号是由所述语音采集设备采集得到,所述身份标识为与所述语音采集设备具有关联关系的用户的标识;
    调用模块,用于基于所述身份标识调用所述用户的所有情感识别模型,每个情感识别模型与所述用户的一种情感状态信息对应;
    确定模块,用于基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息;
    发送模块,用于将确定的情感状态信息发送给所述数据管理服务器进行管理。
  14. 如权利要求13所述的装置,其特征在于,所述调用模块用于:
    基于所述身份标识,从存储的参考对应关系中确定对应的所有情感识别模型标识,所述参考对应关系用于存储至少一个用户中每个用户的身份标识与所述每个用户的所有情感识别模型标识之间的对应关系;
    调用所确定的所有情感识别模型标识对应的情感识别模型。
  15. 如权利要求13或14所述的装置,其特征在于,所述确定模块用于:
    提取所述语音信号的声纹特征;
    分别将所述声纹特征输入至调用的所有情感识别模型中的每个情感识别模型,由所述每个情感识别模型对所述声纹特征进行识别处理并输出情感相似度;
    基于输出的所有情感相似度,确定最大情感相似度对应的情感识别模型;
    将确定的情感识别模型对应的情感状态信息确定为所述用户当前的情感状 态信息。
  16. 如权利要求13所述的装置,其特征在于,所述调用模块还用于:
    基于所述身份标识,调用所述用户对应的语音验证模型;
    通过所述语音验证模型对所述语音信号进行语音验证;
    当对所述语音信号验证通过时,基于所述身份标识调用所述用户的所有情感识别模型。
  17. 如权利要求13所述的装置,其特征在于,所述装置还包括:
    存储模块,用于将所述语音信号和所述情感状态信息保存为训练样本,所述训练样本用于对所述情感状态信息对应的情感识别模型继续训练。
  18. 如权利要求13所述的装置,其特征在于,所述发送模块用于:
    将所述语音信号和所确定的情感状态信息发送给所述数据管理服务器进行管理。
  19. 一种情感识别服务器,其特征在于,包括:
    处理器;
    用于存储处理器可执行指令的存储器;
    其中,所述处理器被配置为:
    接收所述语音采集设备采集的语音信号和与所述语音采集设备具有关联关系的用户的身份标识;
    基于所述身份标识调用所述用户的所有情感识别模型,每个情感识别模型与所述用户的一种情感状态对应;
    基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息;
    将确定的情感状态信息发送给所述数据管理服务器进行管理。
  20. 一种计算机可读存储介质,所述计算机可读存储介质上存储有指令,其特征在于,所述指令被处理器执行时实现如下方法:
    接收所述语音采集设备采集的语音信号和与所述语音采集设备具有关联关系的用户的身份标识;
    基于所述身份标识调用所述用户的所有情感识别模型,每个情感识别模型与所述用户的一种情感状态对应;
    基于所述语音信号,通过调用的情感识别模型确定所述用户当前的情感状态信息;
    将确定的情感状态信息发送给所述数据管理服务器进行管理。
PCT/CN2020/080256 2019-03-21 2020-03-19 监控系统、方法、装置、服务器及存储介质 Ceased WO2020187300A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201910219098.2A CN111739558B (zh) 2019-03-21 2019-03-21 监控系统、方法、装置、服务器及存储介质
CN201910219098.2 2019-03-21

Publications (1)

Publication Number Publication Date
WO2020187300A1 true WO2020187300A1 (zh) 2020-09-24

Family

ID=72519629

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2020/080256 Ceased WO2020187300A1 (zh) 2019-03-21 2020-03-19 监控系统、方法、装置、服务器及存储介质

Country Status (2)

Country Link
CN (1) CN111739558B (zh)
WO (1) WO2020187300A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN115394304A (zh) * 2021-03-30 2022-11-25 北京百度网讯科技有限公司 声纹判定方法、装置、系统、设备和存储介质
US12027242B2 (en) * 2018-06-29 2024-07-02 Signant Health Global Llc Continuous user identity verification in clinical trials via voice-based user interface

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111968679B (zh) * 2020-10-22 2021-01-29 深圳追一科技有限公司 情感识别方法、装置、电子设备及存储介质
CN112767946A (zh) * 2021-01-15 2021-05-07 北京嘀嘀无限科技发展有限公司 确定用户状态的方法、装置、设备、存储介质和程序产品
CN113988155A (zh) * 2021-09-27 2022-01-28 北京智象信息技术有限公司 一种基于智能语音的电子相框图片展示方法及系统
CN117524262A (zh) * 2023-12-20 2024-02-06 广州易风健康科技股份有限公司 基于ai的语音情绪识别模型的训练方法

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2005283647A (ja) * 2004-03-26 2005-10-13 Matsushita Electric Ind Co Ltd 感情認識装置
WO2008092473A1 (en) * 2007-01-31 2008-08-07 Telecom Italia S.P.A. Customizable method and system for emotional recognition
CN101930735A (zh) * 2009-06-23 2010-12-29 富士通株式会社 语音情感识别设备和进行语音情感识别的方法
CN107452385A (zh) * 2017-08-16 2017-12-08 北京世纪好未来教育科技有限公司 一种基于语音的数据评价方法及装置
CN107452405A (zh) * 2017-08-16 2017-12-08 北京易真学思教育科技有限公司 一种根据语音内容进行数据评价的方法及装置
CN107705807A (zh) * 2017-08-24 2018-02-16 平安科技(深圳)有限公司 基于情绪识别的语音质检方法、装置、设备及存储介质

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104036776A (zh) * 2014-05-22 2014-09-10 毛峡 一种应用于移动终端的语音情感识别方法
CN105895101A (zh) * 2016-06-08 2016-08-24 国网上海市电力公司 用于电力智能辅助服务系统的语音处理设备及处理方法
CN106128475A (zh) * 2016-07-12 2016-11-16 华南理工大学 基于异常情绪语音辨识的可穿戴智能安全设备及控制方法

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2005283647A (ja) * 2004-03-26 2005-10-13 Matsushita Electric Ind Co Ltd 感情認識装置
WO2008092473A1 (en) * 2007-01-31 2008-08-07 Telecom Italia S.P.A. Customizable method and system for emotional recognition
CN101930735A (zh) * 2009-06-23 2010-12-29 富士通株式会社 语音情感识别设备和进行语音情感识别的方法
CN107452385A (zh) * 2017-08-16 2017-12-08 北京世纪好未来教育科技有限公司 一种基于语音的数据评价方法及装置
CN107452405A (zh) * 2017-08-16 2017-12-08 北京易真学思教育科技有限公司 一种根据语音内容进行数据评价的方法及装置
CN107705807A (zh) * 2017-08-24 2018-02-16 平安科技(深圳)有限公司 基于情绪识别的语音质检方法、装置、设备及存储介质

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US12027242B2 (en) * 2018-06-29 2024-07-02 Signant Health Global Llc Continuous user identity verification in clinical trials via voice-based user interface
CN115394304A (zh) * 2021-03-30 2022-11-25 北京百度网讯科技有限公司 声纹判定方法、装置、系统、设备和存储介质

Also Published As

Publication number Publication date
CN111739558B (zh) 2023-03-28
CN111739558A (zh) 2020-10-02

Similar Documents

Publication Publication Date Title
CN111739558B (zh) 监控系统、方法、装置、服务器及存储介质
CN107645562A (zh) 数据传输处理方法、装置、设备及系统
Feng et al. Enhancing privacy through domain adaptive noise injection for speech emotion recognition
CN107832720B (zh) 基于人工智能的信息处理方法和装置
WO2016115835A1 (zh) 人体特征数据的处理方法及装置
CN113760674A (zh) 信息生成方法、装置、电子设备和计算机可读介质
CN111241883A (zh) 防止远程被测人员作弊的方法和装置
CN111312243B (zh) 设备交互方法和装置
US20180365779A1 (en) Administering pre-trial judicial services
CN119718640B (zh) 数据分析方法、电子设备、可穿戴设备系统及存储介质
US12405986B2 (en) Efficient content extraction from unstructured dialog text
CN110393539B (zh) 心理异常检测方法、装置、存储介质及电子设备
Setiawan et al. A framework for real time emotion recognition based on human ans using pervasive device
US20180349688A1 (en) Method and system for determining an intent of a subject using behavioural pattern
CN116304629A (zh) 机械设备故障智能诊断方法、系统、电子设备以及存储介质
CN112364285B (zh) 基于ueba建立异常侦测模型的方法、装置及相关产品
CN113380224A (zh) 语种确定方法、装置、电子设备及存储介质
CN113241070A (zh) 热词召回及更新方法、装置、存储介质和热词系统
Olanrewaju An enhanced web-based examination system using automated proctoring and background activity detection
US9785711B2 (en) Online location sharing through an internet service search engine
US20240037458A1 (en) Systems and methods for reducing network traffic associated with a service
CN113784215B (zh) 基于智能电视的性格特征的检测方法和装置
CN117271177A (zh) 基于链路数据的根因定位方法、装置、电子设备及存储介质
CN116610790A (zh) 应答数据的获取方法、装置、设备和介质
CN111949858B (zh) 用于推送信息、呈现信息的方法和设备

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20773387

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 20773387

Country of ref document: EP

Kind code of ref document: A1

122 Ep: pct application non-entry in european phase

Ref document number: 20773387

Country of ref document: EP

Kind code of ref document: A1