WO2017191669A1 - 行動認識装置、行動認識方法、及び行動認識プログラム - Google Patents

行動認識装置、行動認識方法、及び行動認識プログラム Download PDF

Info

Publication number
WO2017191669A1
WO2017191669A1 PCT/JP2016/063538 JP2016063538W WO2017191669A1 WO 2017191669 A1 WO2017191669 A1 WO 2017191669A1 JP 2016063538 W JP2016063538 W JP 2016063538W WO 2017191669 A1 WO2017191669 A1 WO 2017191669A1
Authority
WO
WIPO (PCT)
Prior art keywords
weight
acceleration
feature amount
mobile terminal
behavior
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2016/063538
Other languages
English (en)
French (fr)
Inventor
私市一宏
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Fujitsu Ltd
Original Assignee
Fujitsu Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Fujitsu Ltd filed Critical Fujitsu Ltd
Priority to PCT/JP2016/063538 priority Critical patent/WO2017191669A1/ja
Publication of WO2017191669A1 publication Critical patent/WO2017191669A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H40/00ICT specially adapted for the management or administration of healthcare resources or facilities; ICT specially adapted for the management or operation of medical equipment or devices
    • G16H40/60ICT specially adapted for the management or administration of healthcare resources or facilities; ICT specially adapted for the management or operation of medical equipment or devices for the operation of medical equipment or devices
    • G16H40/67ICT specially adapted for the management or administration of healthcare resources or facilities; ICT specially adapted for the management or operation of medical equipment or devices for the operation of medical equipment or devices for remote operation

Definitions

  • the present invention relates to an action recognition device, an action recognition method, and an action recognition program.
  • the mobile terminal can provide various services such as a navigation service and a health management service to the user using such a sensor.
  • FIG. 11 is a graph showing an example of an experimental result with respect to a recognition rate when sensing data regarding sound and acceleration is used.
  • the vertical axis represents the recognition rate.
  • the recognition rate represents, for example, the probability that the daily behavior automatically recognized by the mobile terminal matches the daily behavior actually performed by the user.
  • the horizontal axis represents the weight given to the acceleration when the sound is constant.
  • the horizontal axis represents an example in which the weight given to the sound is “1.0” and the weight given to the acceleration is changed in the range of “0.0” to “1.0”.
  • the horizontal axis is “0.0”, the weight given to the sound is “1.0”, the weight given to the acceleration is “0.0”, and the horizontal axis “1.0” is the weight given to the sound “1”. .0 ”, and the weight given to the acceleration is“ 1.0 ”.
  • the recognition rate is “77%”, which is the highest recognition rate. Got the rate. That is, when the weight is set to “0.3”, it is possible to improve the recognition rate of living activities compared to the case where the weight is set to another.
  • Examples of technologies related to action recognition using terminals include the following. That is, the sensor data obtained from the sensor is divided into time sections and the same identifier is given to similar sections, and the user gives the mobile terminal a relationship based on the relationship between each identifier and the operation content for the mobile terminal. There is a mobile terminal that estimates the operation content to be performed. According to this technique, it is possible to provide a mobile terminal capable of estimating various operations performed by a user with higher accuracy.
  • a user device that recognizes the current situation of the user. According to this technology, it is said that a user's behavior using a user device such as a smartphone can be analyzed in real time, and a service suitable for the user's situation can be provided according to the analysis result.
  • a user interface that estimates the end time of the action by the user based on the measured acceleration and presents the dialog information associated with the estimated action type at the time based on the estimated end time of the action.
  • a device According to this technology, it is possible to provide a user interface device capable of presenting information corresponding to user behavior in a timely manner.
  • the recognition rate in FIG. 11 is a recognition rate that summarizes various living behaviors, and is not a recognition rate for a specific living behavior. Therefore, in some living behavior, the recognition rate may be maximized when the weight is set to “0.3”. In another living behavior, the weight is set to other than “0.3”. The recognition rate may be maximized. Therefore, when action recognition is performed with the weight fixed at “0.3” at the terminal, the recognition rate for recognizing living action is not necessarily the maximum, and the living action of the user who uses the terminal is accurately recognized. You may not be able to.
  • one disclosure is to provide an action recognition device, an action recognition method, and an action recognition program that improve the accuracy of action recognition.
  • the first weight for sound and the second weight for acceleration are distributed, and the first feature value related to sound measured by the mobile terminal is compared with the first feature amount.
  • a weight control unit weighted with a weight and weighted with the second weight with respect to a second feature with respect to the acceleration measured by the mobile terminal; and the first feature with a weight with the first weight
  • a behavior recognition unit that recognizes the behavior of the user who uses the mobile terminal based on the second feature value weighted by the second weight.
  • an action recognition device an action recognition method, and an action recognition program that improve the accuracy of action recognition can be provided.
  • FIG. 1 is a diagram illustrating a configuration example of an action recognition device.
  • FIG. 2 is a diagram illustrating a configuration example of the behavior recognition system.
  • FIG. 3 is a flowchart showing an operation example.
  • FIG. 4 is a flowchart showing an operation example of the action recognition process.
  • FIG. 5A shows an example of the acoustic learning result
  • FIG. 5B shows an example of the acceleration learning result.
  • 6A and 6B are graphs showing examples of recognition rates.
  • FIG. 7 is a diagram illustrating a configuration example of the behavior recognition system.
  • FIG. 8 is a diagram illustrating an example of the recognition rate.
  • FIG. 9 is a flowchart showing an operation example.
  • FIG. 10 is a diagram illustrating a configuration example of the behavior recognition system.
  • FIG. 11 is a diagram illustrating an example of the recognition rate.
  • FIG. 1 shows a configuration example of the action recognition apparatus 400 in the first embodiment.
  • the behavior recognition device 400 can automatically recognize the behavior of a user who uses a mobile terminal device (hereinafter, also referred to as “terminal”) 100, for example.
  • terminal mobile terminal device
  • the behavior recognition apparatus 400 includes a weight control unit 410 and a behavior recognition unit 420.
  • the weight control unit 410 distributes a first weight for sound and a second weight for acceleration. Then, the weight control unit 410 weights the first feature quantity related to sound with the first weight, and weights the second feature quantity related to acceleration with the second weight.
  • the behavior recognition unit 216 recognizes the behavior of the user who uses the terminal 100 based on the first feature value weighted with the first weight and the second feature value weighted with the second weight.
  • the first weight for sound and the second weight for acceleration are distributed, and the first and second feature values are weighted. Therefore, compared with the case where the first weight and the second weight are fixedly set, the first weight and the second weight can be distributed flexibly.
  • the first weight for sound and the second weight for acceleration are fixedly set to be “0.3”.
  • the feature for sound is more likely to be output than the feature for acceleration.
  • the feature amount is more likely to appear in acceleration than in sound. Therefore, when the first weight and the second weight are fixed, it is impossible to accurately recognize the living behavior “watching TV” and the living behavior “working in the office”. There is.
  • the first embodiment it is possible to perform weighting corresponding to such living behavior by flexibly distributing the first weight and the second weight. And in the action recognition part 216, compared with the case where the 1st weight and the 2nd weight are fixed, possibility that the above living action will be recognized automatically becomes high.
  • the action recognition apparatus 400 can improve the accuracy of action recognition of the user who uses the terminal 100.
  • FIG. 2 is a diagram illustrating a configuration example of the action recognition system 10 in the second embodiment.
  • the behavior recognition system 10 includes a mobile terminal device (hereinafter sometimes referred to as “terminal”) 100 and a cloud server (hereinafter sometimes referred to as “server”) 200. Terminal 100 and server 200 are connected via network 300.
  • terminal mobile terminal device
  • server cloud server
  • the terminal 100 is a communication device such as a smartphone, a feature phone, a tablet terminal, a personal computer, or a game device.
  • the terminal 100 includes various sensors, can convert data acquired by the sensor into a file, and transmit the filed data (hereinafter also referred to as “file data”) to the server 200.
  • the server 200 can acquire the file data transmitted from the terminal 100 and recognize the living behavior of the user who uses the terminal 100 based on the data included in the file data.
  • the server 200 can receive file data transmitted from a plurality of terminals 100.
  • the server 200 can also recognize the living behavior of each user who uses each of the plurality of terminals 100 based on the data included in each file data.
  • the network 300 may be connected to, for example, a base station device or a gateway device. In FIG. 2, such a device is represented as a network 300.
  • the terminal 100 includes a control unit 110, a storage unit 120, a sound pressure level reading unit 130, and a communication unit 140.
  • the control unit 110 controls the storage unit 120, the sound pressure level reading unit 130, and the communication unit 140, for example.
  • the control unit 110 stores data or the like in the storage unit 120 or reads the stored data from the storage unit 120.
  • the control unit 110 can exchange data and the like with the server 200 via the communication unit 140, for example.
  • the control unit 110 includes a measurement unit 111 and a transmission unit 112.
  • the measurement unit 111 includes, for example, a sensor and a sensor function.
  • the measurement unit 111 includes a microphone 1110 and an acceleration sensor 1111.
  • the microphone 1110 collects sound (or sound; hereinafter referred to as “sound”) around the terminal 100 and collects sound data (or sound data; hereinafter referred to as “acoustic data”). Is generated.
  • the microphone 1110 outputs the generated acoustic data to the transmission unit 112 and the storage unit 120.
  • the sound is, for example, an environmental sound, and includes a mechanical sound emitted from a machine, a voice emitted from a human being, and the like.
  • the acceleration sensor 1111 measures acceleration of the mobile terminal 100 and generates acceleration data, for example.
  • the acceleration sensor 1111 outputs the generated acceleration data to the transmission unit 112 and the storage unit 120.
  • the transmission unit 112 outputs the acoustic data and acceleration data acquired from the measurement unit 111 and the sound pressure level data acquired from the sound pressure level reading unit 130 to the storage unit 120 and the communication unit 140.
  • the control unit 110 may be a processor such as a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a DSP (Digital Signal Processor), or an FPGA (Field Programmable Gate Array).
  • the control unit 110 reads the program stored in the storage unit 120 and executes the program, thereby realizing the processing and functions of the control unit 110 described in the second embodiment. May be.
  • the storage unit 120 is a memory such as a RAM (Random Access Memory).
  • the storage unit 120 stores acoustic data and acceleration data output from the control unit 110.
  • the storage unit 120 also stores sound pressure level data output from the sound pressure level reading unit 130.
  • the storage unit 120 may store these data in a sensor DB (Data ⁇ Base) 121.
  • the storage unit 120 may store acoustic data, acceleration data, and sound pressure level data as a file. Data filed in this way may be referred to as acoustic file data, for example.
  • the control unit 110 can read the acoustic file data from the storage unit 120.
  • the sound pressure level reading unit 130 acquires sound pressure level data for the sound, for example, when the sound data is acquired by the microphone 1110.
  • the sound pressure level data is data representing sound pressure, for example, and is represented by [dB], for example.
  • the sound pressure level reading unit 130 outputs the read sound pressure level data to the transmission unit 112 and the storage unit 120.
  • the communication unit 140 transmits, for example, acoustic file data received from the control unit 110 to the server 200 via the network 300. Further, the communication unit 140 outputs data received from the server 200 via the network 300 to the control unit 110.
  • the server 200 includes a control unit 210, a storage unit 220, and a communication unit 240.
  • the control unit 210 controls the storage unit 220 and the communication unit 240, for example. For example, when the control unit 210 receives the acoustic file data transmitted from the terminal 100 from the communication unit 240, the control unit 210 performs processing such as action recognition. Then, the control unit 210 stores the recognition result and the measurement result in the storage unit 220. In addition, the control unit 210 can appropriately read the recognition result stored in the storage unit 220 and transmit it to the terminal 100 via the communication unit 240.
  • the control unit 210 includes a reception unit 211, a feature amount calculation unit 212, a weight control unit 213, and an action recognition unit 216.
  • corresponds to the server 200, for example.
  • the weight control unit 410 in the first embodiment corresponds to the weight control unit 213, for example.
  • the action recognition unit 420 in the first embodiment corresponds to the action recognition unit 216, for example.
  • the receiving unit 211 receives, for example, the acoustic file data output from the communication unit 240, and extracts acoustic data, acceleration data, and sound pressure level data from the received acoustic file data. Then, the reception unit 211 outputs acoustic data and acceleration data to the feature amount calculation unit 212. In addition, the reception unit 211 outputs sound pressure level data to the weight control unit 213.
  • the feature amount calculation unit 212 calculates an acoustic feature amount and an acceleration feature amount based on the acoustic data and the acceleration data.
  • an acoustic feature amount for example, there is MFCC (Mell-Frequency Cepstrum Coefficients).
  • the MFCC represents, for example, acoustic feature quantities based on filter banks arranged at equal intervals on the mel frequency axis.
  • the feature amount calculation unit 212 performs, for example, the following processing when calculating the feature amount by MFCC. That is, the feature amount calculation unit 212 emphasizes high-frequency components of the acoustic data by pre-emphasis processing, and converts the acoustic data from time components to frequency components by FFT (Fast Fourier ⁇ Transform) processing. . Further, the feature amount calculation unit 212 obtains the MFCC by performing DCT (Discrete Cosine Transform) processing on the converted acoustic data.
  • DCT Discrete Cosine Transform
  • the MFCC is an example of an acoustic feature amount, and for example, an LPC (Linear Predictive Coding) coefficient or a PACOR (Parcial Corelated) coefficient may be used.
  • the LPC coefficient is, for example, a coefficient that minimizes the mean square error between the value predicted from the past M input signals and the actual input signal.
  • the PACOR coefficient is, for example, a PACOR (Parcial Corelated) coefficient that is a correlation coefficient between prediction errors of forward prediction and backward prediction.
  • the acceleration feature quantity may be, for example, an average value, a variance value, a value related to a correlation, or the like.
  • the feature amount calculation unit 212 may read acceleration data acquired in the past from the measurement result DB 221 and calculate a feature amount such as an average value or a variance value.
  • the feature amount calculation unit 212 outputs the acoustic feature amount and the acceleration feature amount to the action recognition unit 216.
  • the weight control unit 213 allocates the acoustic weight ws for the acoustic learning result 214 and the acceleration weight wa for the acceleration learning result 215 based on the sound pressure level, and sets the weights ws and wa.
  • the weight control unit 213 sets the sound weight ws to be equal to or higher than the sound weight threshold (or sets the acceleration weight wa). Sort to be smaller than the acceleration weight threshold.
  • the weight control unit 213 makes the acoustic weight ws smaller than the acoustic weight threshold (or makes the acceleration weight wa greater than the acceleration weight threshold). Distribute. The reason why the acoustic weight ws and the acceleration weight wa are distributed in this way will be described in an operation example.
  • the weight control unit 213 weights the acoustic learning result 214 with the distributed acoustic weight ws. Also, the weight control unit 213 weights the acceleration learning result 215 with the distributed acceleration weight wa.
  • the two learning results 214 and 215 will be described.
  • FIG. 5A shows an example of the acoustic learning result 214
  • FIG. 5B shows an example of the acceleration learning result 215.
  • the acoustic learning result 214 and the acceleration learning result 215 include, for example, “TV” (life behavior of watching TV) or “Office” (office). There is a learning result for each life behavior of the user, such as (life behavior of working at (workplace)).
  • the acoustic learning result 214 and the acceleration learning result 215 are, for example, the same unit as the feature amount of the acoustic data and the feature amount of the acceleration data calculated by the feature amount calculation unit 212, and the representative value of the feature amount of the acoustic data.
  • a representative value or a representative value (or representative amount) of a feature amount of acceleration data.
  • the representative value may be, for example, an average value, a median value, a variance value, or the like.
  • “Office sound ” represents an acoustic learning result in the living behavior “Office”.
  • “TV acceleration ” represents the acceleration learning result (or representative value) in the living behavior “TV”.
  • the learning results 214 and 215 are, for example, those calculated based on the acoustic and acceleration feature amounts calculated by the feature amount calculation unit 212 in the past.
  • the weight control unit 213 weights, for example, the acoustic feature amount and the acceleration feature amount calculated in this way with the assigned weights ws and wa, respectively.
  • each learning result exists for each living behavior. Therefore, for example, the weight control unit 213 weights the plurality of acoustic learning results 214 corresponding to the living behavior with the same weight ws, and the same weight wa to the plurality of acceleration learning results 215 corresponding to the living behavior. Weight with.
  • the acoustic learning result 214 and the acceleration learning result 215 may be stored in the storage unit 220, for example, and the weight control unit 213 may appropriately read from the storage unit 220 and perform processing.
  • the weight control unit 213 outputs the weighted acoustic learning result 214 and the acceleration learning result 215 to the action recognition unit 216.
  • the behavior recognition unit 216 recognizes (or estimates) the living behavior of the user who uses the terminal 100 based on the feature amount output from the feature amount calculation unit 212 and the weighted learning result output from the weight control unit 213. )
  • the behavior recognition unit 216 performs behavior recognition as follows, for example. That is, the action recognition unit 216 selects a first learning result closest to the acoustic feature amount acquired from the feature amount calculation unit 212 from the plurality of weighted acoustic learning results 214. In addition, the behavior recognition unit 216 selects a second learning result closest to the acceleration feature amount acquired from the feature amount calculation unit 212 from the plurality of weighted acceleration learning results 215. However, the behavior recognition unit 216 has the same lifestyle behavior as the lifestyle behavior corresponding to the first learning result (for example, “TV”) and the lifestyle behavior corresponding to the second learning result (for example, “TV”). . When the living behavior is different, the behavior recognition unit 216 selects a learning result close to the feature amount acquired from the feature amount calculation unit 212 next to the first or second learning result, and repeats this sequentially. The learning result may be selected until the same living behavior is obtained.
  • behavior recognition is an example, and the behavior recognition unit 216 may perform behavior recognition using various selection algorithms.
  • the behavior recognition unit 216 recognizes (or estimates) the lifestyle behavior corresponding to the selected learning result as the lifestyle behavior of the user.
  • the behavior recognition unit 216 stores data indicating the recognition result in the storage unit 220 or transmits the data to the terminal 100 via the communication unit 240.
  • the control unit 210 may also be a processor such as a CPU, MPU, DSP, or FPGA, for example. In this case, for example, the control unit 210 reads the program stored in the storage unit 220 and executes the program, thereby realizing the processing and functions of the control unit 210 described in the second embodiment. It's okay.
  • the storage unit 220 stores data indicating the recognition result recognized by the behavior recognition unit 216, the acoustic and acceleration feature amounts calculated by the feature amount calculation unit 212, and the like.
  • the storage unit 220 may store data indicating the recognition result in the recognition result DB 222, the acoustic or acceleration feature amount calculated by the feature amount calculation unit 212, and the like in the measurement result DB 221.
  • the communication unit 240 receives, for example, acoustic file data transmitted from the terminal 100 via the network 300, and outputs the received acoustic file data to the control unit 210.
  • the communication unit 240 receives data indicating the recognition result from the control unit 210 and transmits the received data to the terminal 100 via the network 300.
  • FIG. 3 is a flowchart showing an operation example of the action recognition system 10 as a whole.
  • S10 to S14 are processes performed by the control unit 110 of the terminal 100
  • S15 to S20 are processes performed by the control unit 210 of the server 200 or the like.
  • the terminal 100 When the terminal 100 starts processing (S10), it performs sensing (S11). For example, the control unit 110 acquires acoustic data and acceleration data from the microphone 1110 and the acceleration sensor 1111, respectively.
  • the terminal 100 reads the sound pressure level (S12).
  • the sound pressure level reading unit 130 reads the sound pressure level simultaneously with sensing, and acquires sound pressure level data corresponding to the sound data.
  • the terminal 100 stores the sensing data and sound pressure level data in the sensor DB 121 (S13).
  • the control unit 110 stores sensing data in the sensor DB 121
  • the sound pressure level reading unit 130 stores sound pressure level data in the sensor DB 121.
  • these data are stored in the sensor DB 121 as acoustic file data, for example.
  • the terminal 100 transmits acoustic file data to the server 200 (S14).
  • the control unit 110 reads the acoustic file data from the sensor DB 121 and transmits the acoustic file data to the server 200.
  • the server 200 When the server 200 receives the acoustic file data (S15), the server 200 extracts the sound pressure level data from the acoustic file data (S16).
  • the server 200 extracts the feature amount of each file (S17).
  • the feature amount calculation unit 212 calculates an acoustic feature amount and an acceleration feature amount for each acoustic file data.
  • the server 200 performs action recognition processing (S18).
  • FIG. 4 is a flowchart showing an operation example of the action recognition process. The process illustrated in FIG. 4 is performed by, for example, the weight control unit 213 and the action recognition unit 216 of the server 200.
  • the server 200 refers to the sound pressure level data (S181).
  • the weight control unit 213 refers to the sound pressure level data received from the receiving unit 211.
  • the server 200 sets the acoustic weight ws and the acceleration weight wa based on the sound pressure level (S182 to S185).
  • the weight control unit 213 sets the acoustic weight ws to be equal to or higher than the acoustic weight threshold (or the acceleration weight wa from the acceleration weight threshold). Sort to make it smaller). For example, when the sound pressure level is smaller than the sound pressure level threshold, the weight control unit 213 makes the acoustic weight ws smaller than the acoustic weight threshold (or makes the acceleration weight wa equal to or greater than the acceleration weight threshold). To).
  • FIG. 6A shows an example of an experimental result with respect to the recognition rate in the case of “TV” (while watching TV).
  • TV represents an example of living behavior that should be recognized by the behavior recognition unit 216.
  • the vertical axis represents the recognition rate
  • the horizontal axis represents the acceleration weight wa when the acoustic weight ws is constant.
  • the horizontal axis represents, for example, the weight when the acoustic weight ws is fixed to “1.0” and the acceleration weight wa is changed from “0.0” to “1.0”.
  • the horizontal axis “0.0” represents the weight when the acoustic weight ws is “1.0” and the acceleration weight wa is “0.0”, and the horizontal axis “1.0” This represents the weight when the acoustic weight ws is “1.0” and the acceleration weight wa is “1.0”.
  • the sound weight ws becomes relatively gradually larger and the acceleration weight wa becomes relatively smaller as it goes to the left side in the drawing.
  • the acoustic weight ws becomes relatively smaller and the acceleration weight wa becomes relatively larger gradually.
  • the recognition rate is highest when the weight is “0.0”.
  • the sound tends to be larger than the acceleration and features more easily. That is, since the user is sitting, the acoustic feature is larger than the acceleration feature.
  • the terminal 100 collects sound from the TV when the user views the TV, the characteristics of the sound are larger than the acceleration. Therefore, by increasing the acoustic weight ws relatively (as it goes to the left side in FIG. 6A), the recognition rate of the living behavior “TV” can be gradually increased.
  • the sound weight ws becomes relatively smaller toward the right side. Therefore, the closer the weight is to “1.0”, the more difficult it becomes to recognize the characteristics of the sound from TV viewing.
  • the weight control unit 213 sets the acoustic weight ws to be equal to or greater than the acoustic weight threshold (or the acceleration weight wa is smaller than the acceleration weight threshold). ) Sort and weight. As a result, for example, the server 200 can accurately recognize the living behavior “TV”.
  • FIG. 6B shows an example of the experimental result for the recognition rate in the case of “Office” (working in the office). “Office” also represents an example of living behavior that should be recognized by the behavior recognition unit 216.
  • the recognition rate is highest when the weight is “0.7”.
  • the acceleration weight wa becomes relatively larger than the acoustic weight ws as it goes to the right side in the drawing, and the acceleration feature can be easily recognized. Therefore, it is possible to accurately recognize the daily behavior “Office”.
  • the weight control unit 213 makes the acoustic weight ws smaller than the acoustic weight threshold (or makes the acceleration weight wa greater than the acceleration weight threshold). To) and assign weights.
  • the server 200 can accurately recognize the living behavior “Office”.
  • the recognition rate of “TV” is “48%” in the example of FIG.
  • the recognition rate of “Office” is “85%” in the example of FIG.
  • the average of these two recognition rates is “66.5%”.
  • the weight is set to “0.1” when the sound pressure level is equal to or higher than the sound pressure level threshold, and the weight is set to “0.7” when the sound pressure level is smaller than the sound pressure level threshold.
  • the recognition rate of “TV” is “50%”.
  • the recognition rate of “Offce” is “90%”.
  • the average of the two recognition rates is “70.0%”, and it is possible to improve the recognition rate of the user's action recognition as compared with the case where the weight is fixed to “0.3”.
  • the weight control unit 213 sets the acoustic weight ws to “0.0” and the acceleration weight wa. Is set to “1.0” (S183).
  • the weight control unit 213 sets the acoustic weight ws to “0.5” and the acceleration. Is set to "0.5" (S184).
  • the weight control unit 213 sets the acoustic weight ws to “1.0” and the acceleration weight wa to “0. 0 "is set (S185).
  • the server 200 executes the recognition module using the set weight (S186).
  • the recognition module may be executed by the action recognition unit 216 performing the process described above.
  • the server 200 ends the action recognition process (S187).
  • the server 200 when the server 200 performs the action recognition process (S18), the server 200 stores the recognition result in the recognition result DB 222 (S19).
  • the behavior recognition unit 216 stores data indicating the recognition result in the recognition result DB 222.
  • the recognition rate of the living behavior “TV” is “52%” as shown in FIG.
  • the recognition rate of the living behavior “TV” is “32%” with respect to the weight “0.0” as shown in FIG.
  • the server 200 selects the learning results 214 and 215 corresponding to the location where the terminal 100 is located and performs action recognition on the server 200 based on the position information acquired by the terminal 100. .
  • FIG. 7 shows a configuration example of the action recognition system 10 in the third embodiment
  • FIG. 9 shows a flowchart of an operation example in the third embodiment.
  • the measurement unit 111 of the terminal 100 further includes a GPS (Global Positioning System) sensor 1112.
  • GPS Global Positioning System
  • the GPS sensor 1112 acquires position data indicating the position of the terminal 100 through communication with an artificial satellite.
  • the GPS sensor 1112 outputs position data to the storage unit 120.
  • the position data acquired by the GPS sensor 1112 is included in the acoustic file data and transmitted to the server 200.
  • the receiving unit 211 of the server 200 extracts position data and sound pressure level from the acoustic file data, and outputs them to the weight control unit 213.
  • the acoustic learning results 214-1 to 214-n include acoustic learning results for each place (or position).
  • n is an integer of 2 or more
  • examples of the acoustic learning results 214-1 to 214-n for each place there are an acoustic learning result for home and an acoustic learning result for office.
  • the acceleration learning results 215-1 to 215-m include the acceleration learning results for each place.
  • Examples of the acceleration learning results 215-1 to 215-m for each place include an acceleration learning result for home use and an acceleration learning result for office use.
  • FIG. 8 shows an example of the recognition rate when the action recognition is performed using the learning result (A) at a certain place X and when the action recognition is performed using the learning result (B) at a different place Y. Yes.
  • the recognition rate may vary depending on the location where the terminal 100 is located. This is considered to be because, for example, the environmental sound around the terminal 100 differs depending on the location and affects the learning result.
  • the server 200 stores a plurality of acoustic learning results 214-1 to 214-n and a plurality of acceleration learning results 215-1 to 215-m in the storage unit 220 according to the location.
  • the server 200 selects any one of the plurality of acoustic learning results 214-1 to 214-n based on the position data acquired by the terminal 100, and selects the plurality of acceleration learning results 215-1 to 215-m. Select one from among.
  • the server 200 performs behavior recognition using the selected learning results 214 and 215 as in the second embodiment.
  • FIG. 9 is a diagram illustrating an operation example in the third embodiment.
  • the terminal 100 acquires position data with the GPS sensor 1112 (S31, S32), and stores it in the sensor DB 121 together with acoustic data, acceleration data, and the like (S33). And the terminal 100 transmits the acoustic file data containing these data to the server 200 (S34).
  • the server 200 receives the acoustic file data (S35), and determines the location of the terminal 100 based on the position data extracted from the acoustic file data (S36) (S38). In the example of FIG. 9, “Home”, “Office”, “Restaurant”, and “Other” are displayed. For example, the weight control unit 213 may select “home” or “office” based on the position data and the map data.
  • the server 200 weights the selected learning results 214 and 215 with weights ws and wa.
  • the server 200 performs action recognition using the weighted learning results 214 and 215 (S39 to S42).
  • the server 200 can perform highly accurate action recognition according to the position of the terminal 100, for example.
  • the server 200 has described an example in which the acoustic weight ws and the acceleration weight wa are variable based on the sound pressure level data.
  • the server 200 may change the acoustic weight ws and the acceleration weight wa based on the acceleration level (or acceleration data level).
  • FIG. 10 shows a configuration example of the action recognition system 10 when the weights ws and wa are variable based on the acceleration level.
  • the weight control unit 213 distributes the acoustic weight ws so as to be equal to or larger than the acoustic weight threshold (or the acceleration weight wa is smaller than the acceleration weight threshold).
  • the weight control unit 213 distributes the acoustic weight ws so as to be smaller than the acoustic weight threshold (or so that the acceleration weight wa is equal to or higher than the acceleration weight threshold).
  • the weight control unit 213 weights the learning results 214 and 215 with the weights ws and wa thus distributed, and the behavior recognition unit 216 performs behavior recognition based on the weighted learning results 214 and 215.
  • acceleration level information may be extracted instead of the sound pressure level information in S16 of FIG.
  • the acceleration level is used instead of the sound pressure level (S181 to S185), and the weights ws and wa are set by comparing the acceleration level with the acceleration level threshold value. Good.
  • the weight control unit 213 distributes the acoustic weight ws so as to be smaller than the acoustic weight threshold (or so that the acceleration weight wa is equal to or higher than the acceleration weight threshold).
  • the acoustic weight ws is relatively smaller than the acceleration weight wa. Therefore, in the server 200, two weights ws and wa are fixed (for example, “0. Compared with the case of 3)), the recognition rate can be improved.
  • the acoustic weight ws is relatively smaller than the acceleration weight wa. Therefore, in the server 200, two weights ws and wa are fixed (for example, “0. Compared with the case of 3)), the recognition rate can be improved.
  • the server 200 has been described as performing action recognition.
  • the control unit 110 of the terminal 100 includes a feature amount calculation unit 212, a weight control unit 213, and an action recognition unit 216, and stores the acoustic learning result 214 and the acceleration learning result 215 in the storage unit 120. Also good.
  • the action recognition device may be the server 200 or the terminal 100.
  • the feature amount calculation unit 212 may be provided in the terminal 100, and the weight control unit 213 and the action recognition unit 216 may be provided in the server 200.
  • the feature amount calculation unit 212 and the weight control unit 213 may be provided in the terminal 100, and the action recognition unit 216 may be provided in the server 200.
  • the feature amount may be transmitted from the terminal 100 to the server 200, or the learning results 214 and 215 weighted with the feature amount may be transmitted.
  • the sound pressure level reading unit 130 may be included in the measurement unit 111.
  • the sound pressure level reading unit 130 may be a part of a sensor.
  • Action recognition system 100 Mobile terminal device (terminal) 110: Control unit 111: Measurement unit 1110: Microphone 1111: Acceleration sensor 1112: GPS sensor 120: Storage unit 121: Sensor DB 130: Sound pressure level reading unit 140: Communication unit 200: Cloud server device (server) 210: Control unit 211: Reception unit 212: Feature amount calculation unit 213: Weight control unit 214 (214-1 to 214-n): Acoustic learning result 215 (215-1 to 215-m): Acceleration learning result 216: Behavior Recognition unit 220: Storage unit 222: Recognition result DB 300: Network

Landscapes

  • Health & Medical Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Biomedical Technology (AREA)
  • Business, Economics & Management (AREA)
  • General Business, Economics & Management (AREA)
  • Epidemiology (AREA)
  • General Health & Medical Sciences (AREA)
  • Medical Informatics (AREA)
  • Primary Health Care (AREA)
  • Public Health (AREA)
  • Telephone Function (AREA)
  • Telephonic Communication Services (AREA)

Abstract

音響についての第1の重みと加速度についての第2の重みを振り分けて、移動端末で測定された音響に関する第1の特徴量に対して前記第1の重みで重み付けし、前記移動端末で測定された加速度に関する第2の特徴量に対して前記第2の重みで重み付けする重み制御部と、前記第1の重みで重み付けられた前記第1の特徴量と前記第2の重みで重み付けられた前記第2の特徴量に基づいて、前記移動端末を使用するユーザの行動を認識する行動認識部とを備える行動認識装置。

Description

行動認識装置、行動認識方法、及び行動認識プログラム
 本発明は、行動認識装置、行動認識方法、及び行動認識プログラムに関する。
 現在、スマートフォンなどの移動端末においては種々のセンサが備えられている。移動端末は、このようなセンサを利用して、ナビゲーションサービスや健康管理サービスなど種々のサービスをユーザに提供することが可能である。
 また、移動端末におけるセンサを利用してユーザの生活行動を自動認識する研究も行われている。例えば、移動端末の加速度センサなどを利用して、“食事をしている”、“ショッピングをしている”、“TVを視聴している”などの生活行動の自動認識を行う、などである。生活行動の自動認識によって、例えば、高齢者の安否確認や行動に合致したコンテンツの提供など、ユーザに対して、様々なアプリケーションを提供することが可能となる。
 図11は、音響と加速度に関するセンシングデータを利用したときの認識率に対する実験結果の例を示すグラフである。図11において、縦軸は認識率を表している。認識率は、例えば、移動端末において自動認識した生活行動とユーザが実際に行った生活行動が一致する確率を表している。一方、横軸は音響を一定にした場合において加速度に与える重みを表している。例えば、横軸は、音響に与える重みを「1.0」とし、加速度に与える重みを「0.0」から「1.0」の範囲で変化させた場合の例を表している。横軸が「0.0」は、音響に与える重みを「1.0」、加速度に与える重みを「0.0」とし、横軸が「1.0」は、音響に与える重みを「1.0」、加速度に与える重みを「1.0」としていることをそれぞれ表している。
 図11に示すように、重みが「0.3」(音響に対して「1.0」、加速度に対して「0.3」)のとき、認識率は「77%」となり、最も高い認識率を得た。すなわち、重みを「0.3」とした場合、重みを他にした場合と比較して、生活行動の認識率を向上させることが可能である。
 端末を利用した行動認識に関する技術として、例えば、以下がある。すなわち、センサから得られたセンサデータなどについて時間区間に分割して相互に類似する区間に同一識別子を付与し、各識別子と携帯端末に対する操作内容との関係に基づき、利用者が携帯端末に対して行う操作内容を推定するようにした携帯端末がある。この技術によれば、利用者が行う各種操作内容をより高い精度で推定することが可能な携帯端末を提供できる、とされる。
 また、ユーザ装置の複数センサのうち少なくとも一つのセンサにより獲得される信号を分析して、客体(object)により順次発生される少なくとも一つの単位行動を認識し、認識される単位行動のパターンを分析して、ユーザの現在状況を認識するユーザ装置がある。この技術によれば、スマートフォンのようなユーザ装置を使用するユーザの行動をリアルタイムで分析して、分析結果に応じてユーザの状況に適したサービスを提供できる、とされる。
 さらに、測定された加速度に基づいてユーザによる行動の終了時点を推定し、推定された行動の種類に対応付けられた対話情報を、推定された行動の終了時点に基づいた時点で提示させるユーザインタフェース装置がある。この技術によれば、ユーザの行動に対応した情報を適時に提示可能なユーザインタフェース装置を提供できる、とされる。
特開2008-204040号公報 特開2012-95300号公報 特開2014-128459号公報
 しかしながら、ユーザの生活行動の種類は様々である。図11における認識率は、様々ある生活行動をまとめた認識率であり、ある特定の生活行動に対する認識率ではない。従って、ある生活行動では、重みを「0.3」とした場合にその認識率が最大となる場合もあれば、別の生活行動では、重みを「0.3」以外とした方が、その認識率は最大となる場合もある。よって、端末において、重みを「0.3」に固定して行動認識が行われる場合、生活行動を認識する認識率は必ずしも最大とはならず、端末を利用するユーザの生活行動を正確に認識することができない場合がある。
 上述した、利用者が携帯端末に対して行う操作内容を推定する技術や、認識される単位行動のパターンを分析してユーザの現在状況を認識する技術などは、音響と加速度の重みを利用した場合の認識率については何も開示されていない。従って、これらの技術も、重みを「0.3」に固定して行動認識が行われる場合に生活行動を正確に認識できない場合がある、ということに対して何ら解決策を示すものではない。
 そこで、一開示は、行動認識の精度を向上させるようにした行動認識装置、行動認識方法、及び行動認識プログラムを提供することにある。
 一態様によれば、行動認識装置において、音響についての第1の重みと加速度についての第2の重みを振り分けて、移動端末で測定された音響に関する第1の特徴量に対して前記第1の重みで重み付けし、前記移動端末で測定された加速度に関する第2の特徴量に対して前記第2の重みで重み付けする重み制御部と、前記第1の重みで重み付けられた前記第1の特徴量と前記第2の重みで重み付けられた前記第2の特徴量に基づいて、前記移動端末を使用するユーザの行動を認識する行動認識部とを備える。
 一開示によれば、行動認識の精度を向上させるようにした行動認識装置、行動認識方法、及び行動認識プログラムを提供することができる。
図1は行動認識装置の構成例を表す図である。 図2は行動認識システムの構成例を表す図である。 図3は動作例を表すフローチャートである。 図4は行動認識処理の動作例を表すフローチャートである。 図5(A)は音響学習結果の例、図5(B)は加速度学習結果の例をそれぞれ表す図である。 図6(A)と図6(B)は認識率の例を表すグラフである。 図7は行動認識システムの構成例を表す図である。 図8は認識率の例を表す図である。 図9は動作例を表すフローチャートである。 図10は行動認識システムの構成例を表す図である。 図11は認識率の例を表す図である。
 以下、本実施の形態について図面を参照して詳細に説明する。本明細書における課題及び実施例は一例であり、本願の権利範囲を限定するものではない。特に、記載の表現が異なっていたとしても技術的に同等であれば、異なる表現であっても本願の技術を適用可能であり、権利範囲を限定するものではない。そして、各実施の形態は、処理内容を矛盾させない範囲で適宜組み合わせることが可能である。
 [第1の実施の形態]
 第1の実施の形態について説明する。図1は第1の実施の形態における行動認識装置400の構成例を表している。行動認識装置400は、例えば、移動端末装置(以下、「端末」と称する場合がある)100を利用するユーザの行動を自動認識することが可能である。なお、端末100では、例えば、センサなどにより、周囲の音響と自局の加速度を測定することが可能である。
 行動認識装置400は、重み制御部410と行動認識部420を備える。
 重み制御部410は、音響についての第1の重みと加速度についての第2の重みを振り分ける。そして、重み制御部410は、音響に関する第1の特徴量に対して第1の重みで重み付けし、加速度に関する第2の特徴量に対して前記第2の重みで重み付けする。
 行動認識部216は、第1の重みで重み付けられた第1の特徴量と第2の重みで重み付けられた第2の特徴量に基づいて、端末100を使用するユーザの行動を認識する。
 このように本第1の実施の形態においては、音響についての第1の重みと加速度についての第2の重みを振り分けて、第1及び第2の特徴量に対して重み付けしている。従って、第1の重みと第2の重みが固定で設定された場合と比較して、第1の重みと第2の重みを柔軟に振り分けることが可能である。
 例えば、図11の例では、音響についての第1の重みと加速度についての第2の重みが「0.3」となるように固定で設定されている。しかし、例えば、TVを視聴している場合、音響についての特徴が加速度についての特徴よりも特徴量が出やすい。また、例えば、オフィス(職場)で作業をしている場合、音響よりも加速度の方がその特徴量は出やすい。従って、第1の重みと第2の重みを固定にすると、“TVを視聴している”という生活行動や、“オフィスで作業をしている”という生活行動を精度よく認識することができない場合がある。
 本第1の実施の形態においては、第1の重みと第2の重みを柔軟に振り分けることで、このような生活行動に対応した重み付けを行うことが可能となる。そして、行動認識部216では、第1の重みと第2の重みを固定にした場合と比較して、上記のような生活行動を自動認識する可能性は高くなる。
 従って、本行動認識装置400は、端末100を利用するユーザの行動認識の精度を向上させることが可能となる。
 [第2の実施の形態]
 次に第2の実施の形態について説明する。最初に、行動認識システムの全体構成例について説明する。
 <行動認識システム>
 図2は本第2の実施の形態における行動認識システム10の構成例を表す図である。
 行動認識システム10は、移動端末装置(以下、「端末」と称する場合がある)100とクラウドサーバ(以下、「サーバ」と称する場合がある)200を備える。端末100とサーバ200はネットワーク300を介して接続される。
 端末100は、例えば、スマートフォン、フィーチャーフォン、タブレット端末、パーソナルコンピュータ、ゲーム装置などの通信装置である。端末100は、種々のセンサを備え、センサにより取得したデータなどをファイル化し、ファイル化したデータ(以下、「ファイルデータ」と称する場合がある)をサーバ200へ送信することが可能である。
 サーバ200は、例えば、端末100から送信されたファイルデータを取得し、当該ファイルデータに含まれるデータに基づいて、端末100を利用するユーザの生活行動を認識することが可能である。サーバ200は、例えば、複数の端末100から送信されたファイルデータを受信することも可能である。この場合、サーバ200は、各ファイルデータに含まれるデータに基づいて、複数の端末100の各々を利用する各ユーザの生活行動を認識することも可能である。
 ネットワーク300は、例えば、基地局装置やゲートウェイ装置などが接続されてもよい。図2においては、そのような装置を含めて、ネットワーク300として表されている。
 次に、端末100とサーバ200の各構成例について説明する。
 <移動端末装置の構成例>
 図2に示すように、端末100は、制御部110、記憶部120、音圧レベル読取部130、及び通信部140を備える。
 制御部110は、例えば、記憶部120や音圧レベル読取部130、及び通信部140を制御する。例えば、制御部110は、データなどを記憶部120に記憶したり、記憶したデータを記憶部120から読み出したりする。また、制御部110は、例えば、通信部140を介してサーバ200との間でデータなどを交換することも可能である。
 制御部110は、測定部111と送信部112を備える。
 測定部111は、例えば、センサやセンサ機能などを備える。図2の例では、測定部111は、マイク1110と加速度センサ1111を備える。
 マイク1110は、例えば、端末100周囲の音響(又は音声。以下、「音響」と称する場合がある)を集音し、音響データ(又は音声データ。以下、「音響データ」と称する場合がある)を生成する。マイク1110は、生成した音響データを送信部112や記憶部120へ出力する。音響は、例えば、環境音などであって、機械から発せられた機械音や、人間から発せられた音声などが含まれる。
 加速度センサ1111は、例えば、移動端末100の加速度を測定し、加速度データを生成する。加速度センサ1111は、生成した加速度データを送信部112や記憶部120へ出力する。
 送信部112は、測定部111から取得した音響データと加速度データ、さらに、音圧レベル読取部130から取得した音圧レベルデータを、記憶部120や通信部140へ出力する。
 なお、制御部110は、例えば、CPU(Central Processing Unit)やMPU(Micro Processing Unit)、DSP(Digital Signal Processor)、FPGA(Field Programmable Gate Array)などのプロセッサでもよい。この場合、例えば、制御部110は、記憶部120に記憶されたプログラムを読み出して、当該プログラムを実行することで、本第2の実施の形態で説明する制御部110の処理や機能を実現してもよい。
 記憶部120は、例えば、RAM(Random Access Memory)などのメモリである。記憶部120は、制御部110から出力された音響データと加速度データなどを記憶する。また、記憶部120は、音圧レベル読取部130から出力された音圧レベルデータを記憶する。記憶部120は、これらのデータをセンサDB(Data Base)121に記憶してもよい。
 なお、記憶部120は、音響データ、加速度データ、及び音圧レベルデータをファイル化して記憶するようにしてもよい。このようにファイル化されたデータのことを、例えば、音響ファイルデータと称する場合がある。制御部110は、音響ファイルデータを記憶部120から読み出すことが可能である。
 音圧レベル読取部130は、例えば、マイク1110で音響データを取得するときに、当該音響に対する音圧レベルデータを取得する。音圧レベルデータは、例えば、音圧を表すデータであって、例えば、[dB]により表される。音圧レベル読取部130は、読み取った音圧レベルデータを送信部112や記憶部120へ出力する。
 通信部140は、例えば、制御部110から受け取った音響ファイルデータなどを、ネットワーク300を介してサーバ200へ送信する。また、通信部140は、ネットワーク300を介してサーバ200から受け取ったデータなどを制御部110へ出力する。
 <クラウドサーバ装置>
 図2に示すように、サーバ200は、制御部210、記憶部220、及び通信部240を備える。
 制御部210は、例えば、記憶部220や通信部240を制御する。例えば、制御部210は、端末100から送信された音響ファイルデータなどを、通信部240から受け取ると、行動認識などの処理を行う。そして、制御部210は、その認識結果や測定結果などを、記憶部220に記憶する。また、制御部210は、記憶部220に記憶された認識結果などを適宜読み出して、通信部240を介して端末100へ送信することも可能である。
 制御部210は、受信部211、特徴量算出部212、重み制御部213、及び行動認識部216を備える。
 なお、第1の実施の形態における行動認識装置400は、例えば、サーバ200に対応する。また、第1の実施の形態における重み制御部410は、例えば、重み制御部213に対応する。さらに、第1の実施の形態における行動認識部420は、例えば、行動認識部216に対応する。
 受信部211は、例えば、通信部240から出力された音響ファイルデータを受け取り、受け取った音響ファイルデータから、音響データ、加速度データ、及び音圧レベルデータを抽出する。そして、受信部211は、音響データと加速度データを特徴量算出部212へ出力する。また、受信部211は、音圧レベルデータを重み制御部213へ出力する。
 特徴量算出部212は、音響データと加速度データに基づいて、音響の特徴量と加速度の特徴量をそれぞれ算出する。音響の特徴量としては、例えば、MFCC(Mell-Frequency Cepstrum Coefficients:メル周波数ケプストラム係数)がある。MFCCは、例えば、メル周波数軸上で等間隔に配置されたフィルタバンクに基づいた音響の特徴量を表している。
 特徴量算出部212は、MFCCによる特徴量を算出する場合、例えば、以下の処理を行う。すなわち、特徴量算出部212は、プリエンファシス(pre-emphasis)処理により音響データの高周波成分を強調させ、FFT(Fast Fourier Transform:高速フーリエ変換)処理により音響データを時間成分から周波数成分に変換する。さらに、特徴量算出部212は、変換後の音響データに対して、DCT(Discrete Cosine Transform:離散コサイン変換)処理を施すことでMFCCを取得する。
 MFCCは、音響の特徴量の一例であって、例えば、LPC(Linear Predictive Coding:線形予測)係数や、PACOR(Parcial Corelated:編自己相関)係数が用いられてもよい。LPC係数は、例えば、過去のM個の入力信号から予測した値と実際の入力信号の二乗平均誤差が最小となる係数である。また、PACOR係数は、例えば、前方予測と後方予測の予測誤差の相関係数であるPACOR(Parcial Corelated:編自己相関)係数である。
 加速度の特徴量としては、例えば、平均値や分散値、相関関係に関する値などであってもよい。例えば、特徴量算出部212は、過去に取得した加速度データを、測定結果DB221から読み出して、平均値や分散値などの特徴量を算出してもよい。
 特徴量算出部212は、音響の特徴量と加速度の特徴量を行動認識部216へ出力する。
 重み制御部213は、音圧レベルに基づいて、音響学習結果214に対する音響の重みwsと加速度学習結果215に対する加速度の重みwaを振り分けて、各重みws,waを設定する。
 本第2の実施の形態では、重み制御部213は、例えば、音圧レベルが音圧レベル閾値以上のときは、音響の重みwsを音響重み閾値以上になるように(あるいは加速度の重みwaを加速度重み閾値より小さくなるように)振り分ける。また、重み制御部213は、音圧レベルが音圧レベル閾値より小さいときは、音響の重みwsを音響重み閾値より小さくなるように(或いは加速度の重みwaを加速度重み閾値以上になるように)振り分ける。音響の重みwsと加速度の重みwaがこのように振り分けられる理由は動作例で説明する。
 重み制御部213は、音響学習結果214に対して、振り分けた音響の重みwsで重み付けする。また、重み制御部213は、加速度学習結果215に対して、振り分けた加速度の重みwaで重み付けする。ここで、2つの学習結果214,215の例について説明する。
 図5(A)は音響学習結果214、図5(B)は加速度学習結果215の例をそれぞれ表している。図5(A)と図5(B)に示すように、音響学習結果214と加速度学習結果215は、例えば、“TV”(TVを視聴している、という生活行動)や“Office”(オフィス(職場)で作業をしている、という生活行動)など、ユーザの生活行動ごとにその学習結果がある。また、音響学習結果214や加速度学習結果215は、例えば、特徴量算出部212により算出された音響データの特徴量や加速度データの特徴量と同じ単位であって、音響データの特徴量の代表値(又は代表量)や、加速度データの特徴量の代表値(又は代表量)であってもよい。代表値としては、例えば、平均値、中央値、分散値などでもよい。図5(A)の例では、“Office音響”が“Office”という生活行動における音響学習結果を表している。また、図5(B)の例では、“TV加速度”が“TV”という生活行動における加速度学習結果(又は代表値)を表している。
 この場合、各学習結果214,215は、例えば、過去に特徴量算出部212で算出された音響と加速度の特徴量に基づいて算出されたもの、ということもできる。重み制御部213では、例えば、このように算出された音響の特徴量と加速度の特徴量に対して、振り分けた重みws,waでそれぞれ重み付けしている。
 また、音響学習結果214についても、加速度学習結果215についても、生活行動ごとに各学習結果が存在する。そのため、重み制御部213は、例えば、生活行動に対応した複数の音響学習結果214に対して同一の重みwsで重み付けし、生活行動に対応した複数の加速度学習結果215に対して同一の重みwaで重み付けする。
 なお、音響学習結果214と加速度学習結果215は、例えば、記憶部220に記憶されてもよく、重み制御部213は記憶部220から適宜読み出して処理を行ってもよい。
 図2に戻り、重み制御部213は、重み付けした音響学習結果214と加速度学習結果215を行動認識部216へ出力する。
 行動認識部216は、特徴量算出部212から出力された特徴量と、重み制御部213から出力された重み付けされた学習結果に基づいて、端末100を利用するユーザの生活行動を認識(又は推定)する。
 行動認識部216は、例えば、以下のようにして行動認識を行う。すなわち、行動認識部216は、重み付けされた複数の音響学習結果214のうち、特徴量算出部212から取得した音響の特徴量と最も近い第1の学習結果を選択する。また、行動認識部216は、重み付けされた複数の加速度学習結果215のうち、特徴量算出部212から取得した加速度の特徴量と最も近い第2の学習結果を選択する。ただし、行動認識部216は、第1の学習結果に対応する生活行動(例えば、“TV”)と、第2の学習結果に対応する生活行動(例えば、“TV”)は同じ生活行動である。生活行動が異なるときは、行動認識部216は、第1又は第2の学習結果の次に、特徴量算出部212から取得した特徴量に近い学習結果を選択して、これを順次繰り返して、同じ生活行動を得るまで学習結果を選択してもよい。
 このような行動認識は一例であって、行動認識部216は、種々の選択アルゴリズムにより、行動認識を行うようにしてもよい。
 行動認識部216は、選択した学習結果に対応する生活行動を、ユーザの生活行動であると認識(又は推定)する。行動認識部216は、認識結果を示すデータを、記憶部220に記憶したり、通信部240を介して端末100へ送信したりする。
 制御部210も、例えば、CPUやMPU、DSP、FPGAなどのプロセッサであってもよい。この場合、例えば、制御部210は、記憶部220に記憶されたプログラムを読み出して、当該プログラムを実行することで、本第2の実施の形態で説明する制御部210の処理や機能を実現してよい。
 記憶部220は、行動認識部216で認識された認識結果を示すデータや、特徴量算出部212で算出された音響や加速度の特徴量などを記憶する。記憶部220は、認識結果を示すデータを認識結果DB222、特徴量算出部212などで算出された音響や加速度の特徴量などを測定結果DB221に記憶してもよい。
 通信部240は、例えば、ネットワーク300を介して端末100から送信された音響ファイルデータなどを受け取り、受け取った音響ファイルデータなどを制御部210へ出力する。また、通信部240は、例えば、制御部210から認識結果を示すデータなどを受け取り、受け取ったデータなどを、ネットワーク300を介して端末100へ送信する。
 <動作例>
 次に動作例について説明する。図3は行動認識システム10全体の動作例を表すフローチャートである。例えば、S10からS14は端末100の制御部110で行われ、S15からS20はサーバ200の制御部210などで行われる処理である。
 端末100は処理を開始すると(S10)、センシングを行う(S11)。例えば、制御部110はマイク1110や加速度センサ1111から音響データや加速度データをそれぞれ取得する。
 次に、端末100は、音圧レベルを読み取る(S12)。例えば、音圧レベル読取部130はセンシングと同時に音圧レベルを読み取り、音響データに対応する音圧レベルデータを取得する。
 次に、端末100は、センシングデータ及び音圧レベルデータをセンサDB121に格納する(S13)。例えば、制御部110はセンシングデータをセンサDB121に記憶し、音圧レベル読取部130は音圧レベルデータをセンサDB121に記憶する。上述したように、これらのデータは、例えば、音響ファイルデータとしてセンサDB121に記憶される。
 次に、端末100は、音響ファイルデータをサーバ200へ送信する(S14)。例えば、制御部110は、センサDB121から音響ファイルデータを読み出して、当該音響ファイルデータをサーバ200へ送信する。
 サーバ200は、音響ファイルデータを受信すると(S15)、音響ファイルデータから音圧レベルデータを抽出する(S16)。
 次に、サーバ200は各ファイルの特徴量を抽出する(S17)。例えば、特徴量算出部212は、音響ファイルデータごとに、音響の特徴量と加速度の特徴量をそれぞれ算出する。
 次に、サーバ200は、行動認識処理を行う(S18)。
 図4は行動認識処理の動作例を表すフローチャートである。図4に示す処理は、例えば、サーバ200の重み制御部213や行動認識部216で行われる。
 サーバ200は、行動認識処理を開始すると(S180)、音圧レベルデータを参照する(S181)。例えば、重み制御部213は受信部211から受け取った音圧レベルデータを参照する。
 次に、サーバ200は、音圧レベルに基づいて、音響の重みwsと加速度の重みwaを設定する(S182~S185)。上述したように、例えば、重み制御部213は、音圧レベルが音圧レベル閾値以上のときは、音響の重みwsを音響重み閾値以上になるように(あるいは加速度の重みwaを加速度重み閾値より小さくなるように)振り分ける。また、例えば、重み制御部213は、音圧レベルが音圧レベル閾値より小さいときは、音響の重みwsを音響重み閾値より小さくなるように(或いは加速度の重みwaを加速度重み閾値以上となるように)振り分ける。
 <音響の重みwsと加速度の重みwaの振り分け>
 ここで、音響の重みwsと加速度の重みwaが上記のように振り分けられる理由について説明する。
 図6(A)は“TV”(TV視聴中)の場合の認識率に対する実験結果の例を表している。“TV”は行動認識部216において認識すべき生活行動の例を表している。
 図6(A)において、縦軸は認識率、横軸は音響の重みwsを一定にした場合の加速度の重みwaを表している。横軸に関しては、例えば、音響の重みwsを「1.0」に固定して、加速度の重みwaを「0.0」から「1.0」に変えたときの重みを表している。
 この場合、横軸「0.0」は、音響の重みwsを「1.0」、加速度の重みwaを「0.0」としたときの重みを表し、横軸「1.0」は、音響の重みwsを「1.0」、加速度の重みwaを「1.0」にしたときの重みを表している。
 従って、図6(A)において、図面上、左側へ行けば行くほど、音響の重みwsは相対的に除々に大きくなり、加速度の重みwaは相対的に除々に小さくなる。一方、図面上、右側へ行けば行くほど、音響の重みwsは相対的に除々に小さくなり、加速度の重みwaは相対的に除々に大きくなる。
 図6(A)に示すように、生活行動が“TV”の場合、認識率が最も高いのは、重みが「0.0」のときである。例えば、ユーザが座ってTVを視聴する場合、音響の方が加速度よりも大きく特徴が出やすい。つまり、ユーザは座っているため、音響の特徴は加速度の特徴よりも大きい。また、ユーザはTVを視聴することで端末100はTVからの音響を集音するため、音響の特徴が加速度よりも大きく出る。従って、音響の重みwsを相対的に大きくすることで(図6(A)において左側へ行くほど)、“TV”という生活行動の認識率を除々に高くすることが可能となる。
 一方、図6(A)においては、右側へ行くほど、音響の重みwsは相対的に小さくなる。従って、重みが「1.0」に近づくほど、TV視聴による音響の特徴を認識することが除々に困難になる。
 以上から、音圧レベルデータが音圧レベル閾値以上のとき、重み制御部213は、音響の重みwsを音響重み閾値以上になるように(あるいは加速度の重みwaを加速度重み閾値より小さくなるように)振り分けて重み付けする。これにより、例えば、サーバ200は“TV”という生活行動を精度よく認識することが可能となる。
 図6(B)は、“Office”(オフィスで作業中)の場合の認識率に対する実験結果の例を表している。“Office”も行動認識部216において認識すべき生活行動の例を表している。
 図6(B)に示すように、生活行動が“Office”の場合、認識率が最も高くなるのは重みが「0.7」のときである。例えば、ユーザがオフィスで作業をしているとき、オフィスのように静かな場所では、加速度の方が音響よりも大きく特徴が出やすい。つまり、図6(B)において、図面上、右側に行くほど、相対的に、加速度の重みwaが音響の重みwsより大きくなり、加速度の特徴を認識しやすくすることが可能となる。従って、“Office”という生活行動を精度よく認識することが可能となる。
 他方、重みが「0.0」に近づけば近づくほど、相対的に、加速度の重みwaは小さくなる。このような場合、加速度の特徴をもつ生活行動を認識する認識率も除々に小さくなる。従って、重みが「0.0」に近いほど、“Office”という生活行動を認識することが困難となる。
 以上から、音圧レベルデータが音圧レベル閾値よりも小さいとき、重み制御部213は、音響の重みwsを音響重み閾値より小さくなるように(或いは加速度の重みwaを加速度重み閾値以上となるように)振り分けて重み付けする。これにより、例えば、サーバ200は“Office”という生活行動を精度よく認識することが可能となる。
 例えば、重みを「0.3」に固定した場合を考える。この場合、“TV”の認識率は、図6(A)の例では、「48%」となる。また、“Office”の認識率は、図6(B)の例では、「85%」となる。この2つの認識率の平均を取ると、「66.5%」となる。
 一方、音圧レベルが音圧レベル閾値以上のときに重みを「0.1」、音圧レベルが音圧レベル閾値より小さいときの重みを「0.7」に設定した場合を考える。重みが「0.1」の場合、“TV”の認識率は「50%」となる。また、重みが「0.7」の場合、“Offce”の認識率は「90%」となる。2つの認識率の平均は、「70.0%」となり、重みを「0.3」に固定にした場合と比較して、ユーザの行動認識の認識率を向上させることが可能である。
 図4の例では、音圧レベルPが「40dB」より小さいとき(S182で「P<40dB」のとき)、重み制御部213は、音響の重みwsを「0.0」、加速度の重みwaを「1.0」に設定する(S183)。
 また、音圧レベルPが「40dB」以上、「70dB」より小さいとき(S182で「40dB≦P<70dB」のとき)、重み制御部213は、音響の重みwsを「0.5」、加速度の重みwaを「0.5」に設定する(S184)。
 さらに、音圧レベルPが「70dB」以上のとき(S182で「P≧70dB」のとき)、重み制御部213は、音響の重みwsを「1.0」、加速度の重みwaを「0.0」に設定する(S185)。
 このように、サーバ200において音圧レベルに応じて重みを可変にすることで、重みを「0.3」に固定にした場合と比較して、例えば、“TV”や“Office”などの生活行動を認識する認識率を向上させ、行動認識の精度も向上させることが可能となる。
 サーバ200は、重みを設定すると(S183~S185)、設定した重みを利用して、認識モジュールを実行する(S186)。例えば、行動認識部216が上述した処理を行うことで、認識モジュールを実行してもよい。
 そして、サーバ200は行動認識処理を終了する(S187)。
 図3に戻り、サーバ200は、行動認識処理を行うと(S18)、認識結果を認識結果DB222に格納する(S19)。例えば、行動認識部216は認識結果を示すデータを認識結果DB222に格納する。
 そして、サーバ200は一連の処理を終了する(S20)。
 ここで、図5(A)及び図5(B)に示す学習結果と図6(A)及び図6(B)に示す認識率の関係について説明する。
 図5(A)に示すように、“TV”の場合の音響の学習結果として、“TV音響”がある。また、図5(B)に示すように、“Office”の場合の加速度の学習結果として“TV加速度”がある。例えば、“TV加速度”に対して加速度の重みwa=「0.0」で重み付けし、“TV音響”に対して音響の重みws=「1.0」で重み付けする。すなわち、重み「0.0」で2つの学習結果“TV加速度”,“TV音響”に重み付けする。この場合、“TV”という生活行動の認識率は、図6(A)に示すように、「52%」となる。
 一方、例えば、“TV加速度”に対して加速度の重みwa=「1.0」を付与し、“TV音響”に対して音響の重みws=「1.0」を付与する。この場合、“TV”という生活行動の認識率は、図6(A)に示すように、重み「0.0」に対して「32%」となる。
 このように学習結果に対して、重みを各々変えて重み付した場合の認識率の一例が、図6(A)や図6(B)において示されている、と考えることができる。
 [第3の実施の形態]
 次に、第3の実施の形態について説明する。第3の実施の形態では、例えば、サーバ200において、端末100で取得した位置情報に基づいて、端末100が位置する場所に対応する学習結果214,215を選択して行動認識を行う例である。
 図7は第3の実施の形態における行動認識システム10の構成例、図9は本第3の実施の形態における動作例のフローチャートをそれぞれ表している。
 図7に示すように、端末100の測定部111は、更に、GPS(Global Positioning System)センサ1112を備える。GPSセンサ1112は、例えば、人工衛星との通信により、端末100の位置を示す位置データを取得する。GPSセンサ1112は、位置データを記憶部120へ出力する。
 端末100では、例えば、GPSセンサ1112で取得した位置データを音響ファイルデータに含めて、サーバ200へ送信する。
 サーバ200の受信部211は、音響ファイルデータから位置データと音圧レベルを抽出して、重み制御部213へ出力する。
 音響学習結果214-1~214-n(nは2以上の整数)は、場所(又は位置)ごとに音響の学習結果を含む。場所ごとの音響学習結果214-1~214-nの例として、自宅用の音響学習結果やオフィス用の音響学習結果などがある。
 また、加速度学習結果215-1~215-m(mは2以上の整数)も、場所ごとに加速度の学習結果を含む。場所ごとの加速度学習結果215-1~215-mの例として、自宅用の加速度学習結果やオフィス用の加速度学習結果などがある。
 図8は、ある場所Xにおける学習結果(A)を用いて行動認識を行った場合と、異なる場所Yにおける学習結果(B)を用いて行動認識を行った場合の認識率の例を表している。図8に示すように、端末100が位置する場所によって認識率が異なる場合がある。これは、例えば、場所によって端末100周囲の環境音は異なり、学習結果に影響を与えるためと考えられる。
 そこで、本第3の実施の形態では、サーバ200は、場所に応じた、複数の音響学習結果214-1~214-nと複数の加速度学習結果215-1~215-mを記憶部220に記憶しておく。そして、サーバ200は、端末100で取得した位置データに基づいて、複数の音響学習結果214-1~214-nの中からいずれを選択し、複数の加速度学習結果215-1~215-mの中からもいずれかを選択する。サーバ200は、選択した学習結果214,215を用いて、第2の実施の形態と同様に行動認識を行う。
 図9は本第3の実施の形態における動作例を表す図である。この場合、端末100は、GPSセンサ1112で位置データを取得し(S31,S32)、音響データや加速度データなどとともにセンサDB121に記憶する(S33)。そして、端末100は、これらのデータを含む音響ファイルデータをサーバ200へ送信する(S34)。
 サーバ200では、音響ファイルデータを受信して(S35)、音響ファイルデータから抽出した位置データに基づいて(S36)、端末100の場所を判別する(S38)。図9の例では、「自宅」、「オフィス」、「レストラン」、「その他」となっている。例えば、重み制御部213は、位置データと地図データに基づいて、「自宅」や「オフィス」などを選択してもよい。
 そして、サーバ200は、選択した学習結果214,215に対して、重みws,waで重み付けを行う。サーバ200は、重み付けを行った学習結果214,215を利用して行動認識を行う(S39~S42)。
 このように本第3の実施の形態では、端末100の位置に基づいて、行動認識に用いる最適な学習結果を選択し、選択した学習結果を用いて行動認識を行っている。従って、サーバ200は、例えば、端末100の位置に応じた、高精度の行動認識を行うことが可能となる。
 [その他の実施の形態]
 その他の実施の形態について説明する。
 第2の実施の形態においては、サーバ200は、音圧レベルデータに基づいて、音響の重みwsと加速度の重みwaを可変にする例について説明した。例えば、サーバ200は、加速度レベル(又は加速度データのレベル)に基づいて、音響の重みwsと加速度の重みwaを可変にしてもよい。
 図10は、加速度レベルに基づいて重みws,waを可変にする場合の行動認識システム10の構成例を表している。この場合、重み制御部213は、加速度レベルが加速度レベル閾値より小さいとき、音響の重みwsを音響重み閾値以上となるように(あるいは加速度の重みwaを加速度重み閾値より小さくなるように)振り分ける。
 また、重み制御部213は、加速度レベルが加速度レベル閾値以上のとき、音響の重みwsを音響重み閾値より小さくなるように(或いは加速度の重みwaを加速度重み閾値以上になるように)振り分ける。
 そして、重み制御部213は、このよう振り分けた重みws,waで各学習結果214,215に重み付けし、行動認識部216は重み付けされた各学習結果214,215に基づいて行動認識を行う。
 動作例についても、図3のS16において、音圧レベル情報に代えて、加速度レベル情報(又は加速度データ)を抽出すればよい。また、図4の行動認識処理においても、音圧レベルに代えて加速度レベルとし(S181~S185)、加速度レベルと加速度レベル閾値との比較により、上記各重みws,waとなるように設定すればよい。
 例えば、重み制御部213は、加速度レベルが加速度レベル閾値以上のとき、音響の重みwsを音響重み閾値より小さくなるように(或いは加速度の重みwaを加速度重み閾値以上になるように)振り分ける。この場合、音響の重みwsは加速度の重みwaよりも相対的に小さくなる。従って、サーバ200では、図6(B)に示すような認識率が右上がりのグラフとなる生活行動(音圧「小」)に対して、2つの重みws,waを固定(例えば「0.3」)にする場合と比較して、認識率を向上させることが可能となる。図6(B)の例では、重みを「0.3」(ws=1.0,wa=0.7)にした場合の認識率は「85%」、重みを「0.7」(ws=1.0,wa=0.7)にした場合の認識率が「90%」となり、“Office”の認識率を向上させることが可能となる。
 また、上述した実施の形態では、例えば、サーバ200において行動認識を行うものとして説明した。例えば、端末100において行動認識を行うようにしてもよい。この場合、端末100の制御部110は、特徴量算出部212、重み制御部213、及び行動認識部216を備え、記憶部120に、音響学習結果214と加速度学習結果215を記憶するようにしてもよい。端末100において行動認識処理が行われることで、例えば、上述した実施の形態の場合と比較して、サーバ200との通信がなくなる分、高速に行動認識を行うことが可能となる。例えば、行動認識装置は、サーバ200でもよいし端末100でもよい。
 又は、特徴量算出部212が端末100に備えられ、重み制御部213と行動認識部216はサーバ200に備えられてもよい。或いは、特徴量算出部212と重み制御部213が端末100に備えられ、行動認識部216がサーバ200に備えられてもよい。端末100からサーバ200へ、特徴量が送信されたり、特徴量と重み付けられた学習結果214,215が送信されたりしてもよい。
 さらに、上述した実施の形態において、音圧レベル読取部130が測定部111外に備えられている例を説明した。例えば、音圧レベル読取部130は測定部111内に含まれてもよい。例えば、音圧レベル読取部130はセンサの一部であってもよい。
10:行動認識システム        100:移動端末装置(端末)
110:制御部            111:測定部
1110:マイク           1111:加速度センサ
1112:GPSセンサ        120:記憶部
121:センサDB          130:音圧レベル読取部
140:通信部            200:クラウドサーバ装置(サーバ)
210:制御部            211:受信部
212:特徴量算出部         213:重み制御部
214(214-1~214-n):音響学習結果
215(215-1~215-m):加速度学習結果
216:行動認識部          220:記憶部
222:認識結果DB         300:ネットワーク

Claims (12)

  1.  音響についての第1の重みと加速度についての第2の重みを振り分けて、移動端末で測定された音響に関する第1の特徴量に対して前記第1の重みで重み付けし、前記移動端末で測定された加速度に関する第2の特徴量に対して前記第2の重みで重み付けする重み制御部と、
     前記第1の重みで重み付けられた前記第1の特徴量と前記第2の重みで重み付けられた前記第2の特徴量に基づいて、前記移動端末を使用するユーザの行動を認識する行動認識部と
     を備えることを特徴とする行動認識装置。
  2.  前記重み制御部は、前記移動端末で測定された音響の音圧レベルに基づいて、前記第1の重みと前記第2の重みを振り分けることを特徴とする請求項1記載の行動認識装置。
  3.  前記重み制御部は、前記移動端末で測定された加速度の加速度レベルに基づいて、前記第1の重みと前記第2の重みを振り分けることを特徴とする請求項1記載の行動認識装置。
  4.  前記重み制御部は、前記第1の特徴量に対する代表量に対して前記第1の重みで重み付けし、前記第2の特徴量に対する代表量に対して第2の重みで重み付けすることを特徴とする請求項1記載の行動認識装置。
  5.  前記第1の特徴量に対する代表量は前記第1の特徴量についての学習結果を表し、前記第2の特徴量に対する代表量は前記第2の特徴量についての学習結果を表すことを特徴とする請求項4記載の行動認識装置。
  6.  更に、前記移動端末で測定された音響に関するデータに対して前記第1の特徴量と前記移動端末で測定された加速度に関するデータに対して第2の特徴量を算出する特徴量算出部を備え、
     前記行動認識部は、前記第1の重みで重み付けられた前記第1の特徴量と、前記第2の重みで重み付けられた前記第2の特徴量と、前記特徴量算出部で算出された前記第1及び第2の特徴量に基づいて、ユーザの行動を認識することを特徴とする請求項1記載の行動認識装置。
  7.  前記重み制御部は、前記移動端末で測定された位置データに基づいて、前記移動端末が位置する場所に対応する前記第1及び第2の特徴量を選択し、選択した前記第1及び第2の特徴量に対して前記第1及び第2の重みでそれぞれ重み付けすることを特徴とする請求項1記載の行動認識装置。
  8.  前記第1の特徴量は前記第1の特徴量に対する代表量であり、前記第2の特徴量は前記第2の特徴量に対する代表量であることを特徴とする請求項7記載の行動認識装置。
  9.  前記第1の特徴量に対する代表量は前記場所に対応して複数の代表量を含み、前記第2の特徴量に対する代表量は前記場所に対応して複数の代表量を含むことを特徴とする請求項8記載の行動認識装置。
  10.  前記行動認識装置は、移動端末装置又はサーバ装置であることを特徴とする請求項1記載の行動認識装置。
  11.  行動認識装置におけるコンピュータに実行させる行動認識プログラムであって、
     音響についての第1の重みと加速度についての第2の重みを振り分けて、移動端末で測定された音響に関する第1の特徴量に対して前記第1の重みで重み付けし、前記移動端末で測定された加速度に関する第2の特徴量に対して前記第2の重みで重み付けする処理と、
     前記第1の重みで重み付けられた前記第1の特徴量と前記第2の重みで重み付けられた前記第2の特徴量に基づいて、前記移動端末を使用するユーザの行動を認識する処理と
     を前記コンピュータに実行させることを特徴とする行動認識プログラム。
  12.  重み制御部と行動認識部を有する行動認識装置における行動認識方法であって、
     前記重み制御部により、音響についての第1の重みと加速度についての第2の重みを振り分けて、移動端末で測定された音響に関する第1の特徴量に対して前記第1の重みで重み付けし、前記移動端末で測定された加速度に関する第2の特徴量に対して前記第2の重みで重み付けし、
     前記行動認識部により、前記第1の重みで重み付けられた前記第1の特徴量と前記第2の重みで重み付けられた前記第2の特徴量に基づいて、前記移動端末を使用するユーザの行動を認識する
     ことを特徴とする行動認識方法。
PCT/JP2016/063538 2016-05-02 2016-05-02 行動認識装置、行動認識方法、及び行動認識プログラム Ceased WO2017191669A1 (ja)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/JP2016/063538 WO2017191669A1 (ja) 2016-05-02 2016-05-02 行動認識装置、行動認識方法、及び行動認識プログラム

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2016/063538 WO2017191669A1 (ja) 2016-05-02 2016-05-02 行動認識装置、行動認識方法、及び行動認識プログラム

Publications (1)

Publication Number Publication Date
WO2017191669A1 true WO2017191669A1 (ja) 2017-11-09

Family

ID=60202872

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2016/063538 Ceased WO2017191669A1 (ja) 2016-05-02 2016-05-02 行動認識装置、行動認識方法、及び行動認識プログラム

Country Status (1)

Country Link
WO (1) WO2017191669A1 (ja)

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2009238116A (ja) * 2008-03-28 2009-10-15 Toshiba Corp 情報機器及び情報提示方法
JP2010016443A (ja) * 2008-07-01 2010-01-21 Toshiba Corp 状況認識装置、状況認識方法、及び無線端末装置
JP2016035760A (ja) * 2015-10-07 2016-03-17 ソニー株式会社 情報処理装置、情報処理方法およびコンピュータプログラム

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2009238116A (ja) * 2008-03-28 2009-10-15 Toshiba Corp 情報機器及び情報提示方法
JP2010016443A (ja) * 2008-07-01 2010-01-21 Toshiba Corp 状況認識装置、状況認識方法、及び無線端末装置
JP2016035760A (ja) * 2015-10-07 2016-03-17 ソニー株式会社 情報処理装置、情報処理方法およびコンピュータプログラム

Similar Documents

Publication Publication Date Title
US11290826B2 (en) Separating and recombining audio for intelligibility and comfort
US10453443B2 (en) Providing an indication of the suitability of speech recognition
US10452982B2 (en) Emotion estimating system
JP7285589B2 (ja) 対話型健康状態評価方法およびそのシステム
US20200312315A1 (en) Acoustic environment aware stream selection for multi-stream speech recognition
CN109460752B (zh) 一种情绪分析方法、装置、电子设备及存储介质
JP6093792B2 (ja) 測位装置、測位方法、測位プログラム、および、測位システム
WO2017187712A1 (ja) 情報処理装置
JP2018532151A (ja) 音声対応デバイス間の調停
US10430896B2 (en) Information processing apparatus and method that receives identification and interaction information via near-field communication link
US20160334437A1 (en) Mobile terminal, computer-readable recording medium, and activity recognition device
US20260018161A1 (en) Channel selection apparatus, channel selection method, and program
US9729982B2 (en) Hearing aid fitting device, hearing aid, and hearing aid fitting method
US11996093B2 (en) Information processing apparatus and information processing method
US20200372924A1 (en) Determining musical style using a variational autoencoder
WO2017191669A1 (ja) 行動認識装置、行動認識方法、及び行動認識プログラム
CN107203259B (zh) 用于使用单和/或多传感器数据融合确定移动设备使用者的概率性内容感知的方法和装置
JP7347994B2 (ja) 会議支援システム
JP6545950B2 (ja) 推定装置、推定方法、およびプログラム
EP2733659A1 (en) Apparatus for sensing socially-related parameters at spatial locations and associated method
KR102083466B1 (ko) 음악을 추천하기 위한 장치, 이를 위한 방법 및 이 방법을 수행하는 프로그램이 기록된 컴퓨터 판독 가능한 기록매체
JP2010238105A (ja) 平常/非平常判定システム、方法及びプログラム
JP2023115687A (ja) データ処理方法、プログラム及びデータ処理装置
WO2023209898A1 (ja) 音声分析装置、音声分析方法及び音声分析プログラム
AU2021238969B2 (en) Latent bio-signal estimation using bio-signal detectors

Legal Events

Date Code Title Description
NENP Non-entry into the national phase

Ref country code: DE

121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16901056

Country of ref document: EP

Kind code of ref document: A1

122 Ep: pct application non-entry in european phase

Ref document number: 16901056

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: JP