WO2025239246A1 - 出力装置及び音出力方法 - Google Patents

出力装置及び音出力方法

Info

Publication number
WO2025239246A1
WO2025239246A1 PCT/JP2025/016660 JP2025016660W WO2025239246A1 WO 2025239246 A1 WO2025239246 A1 WO 2025239246A1 JP 2025016660 W JP2025016660 W JP 2025016660W WO 2025239246 A1 WO2025239246 A1 WO 2025239246A1
Authority
WO
WIPO (PCT)
Prior art keywords
sound
data
acquired
output device
unit
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/JP2025/016660
Other languages
English (en)
French (fr)
Inventor
耕介 大角
利知 金岡
隆行 荒川
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Kyocera Corp
Original Assignee
Kyocera Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Kyocera Corp filed Critical Kyocera Corp
Publication of WO2025239246A1 publication Critical patent/WO2025239246A1/ja
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L15/10Speech classification or search using distance or distortion measures between unknown speech and reference templates
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R3/00Circuits for transducers

Definitions

  • This disclosure relates to an output device and a sound output method.
  • Patent Document 1 describes a portable music playback device equipped with a sound collection means for collecting external sounds.
  • An output device includes: an acquisition unit that acquires sound data; a control unit that determines whether or not the sound acquired by the acquisition unit is to be recorded, based on first data including data of the sound to be recorded; When the control unit receives input for sound playback, the control unit updates the first data based on sound data acquired by the acquisition unit within a predetermined time from the time when the input for sound playback was received.
  • a sound output method includes: Acquiring sound data; determining whether the acquired sound is to be recorded based on first data including data of the sound to be recorded; When an input for sound playback is received, the first data is updated based on data of the acquired sounds that are present within a predetermined time from the timing at which the input for sound playback is received.
  • FIG. 10 is a diagram illustrating an example of first data according to an embodiment. 10 is a timing chart showing an example of the operation of starting and ending recording. 10 is a flowchart illustrating an example of a procedure for sound output processing.
  • FIG. 10 is a diagram illustrating an example of first data according to another embodiment.
  • FIG. 10 is a diagram showing an example of first data according to yet another embodiment.
  • FIG. 10 is a diagram for explaining the processing of a determination unit according to yet another embodiment.
  • FIG. 10 is a block diagram of a sound output device according to yet another embodiment.
  • FIG. 10 is a diagram showing an example of first data and second data according to yet another embodiment.
  • the sound output device 1 (output device) shown in FIG. 1 is a hearable device.
  • the sound output device 1 is a bone conduction earphone.
  • the sound output device 1 is not limited to a bone conduction earphone as long as it is a hearable device.
  • the sound output device 1 may be an ear-hook earphone, a neck-hanging speaker, an inner-ear earphone, an in-ear earphone, a headphone, or a hearing aid.
  • the sound output device 1 is an inner-ear earphone or a headphone, it may have an external sound capture function.
  • the external sound capture function is a function that collects external sounds from the sound output device 1 and outputs them to the user.
  • External sounds are sounds emitted outside the sound output device 1.
  • external sounds include sounds emitted around the user. External sounds may also include sounds emitted by the user themselves.
  • the sound output device 1 includes a housing 1L, a housing 1R, and a fixing member 1F.
  • the housing 1L is placed against the user's left temple.
  • the housing 1R is placed against the user's right temple.
  • the fixing member 1F fixes the housing 1L and the housing 1R to the user's left and right temples, respectively.
  • the fixing member 1F includes a left ear hook that is placed over the user's left ear, a right ear hook that is placed over the user's right ear, and a band that connects these ear hooks.
  • the fixing member 1F may be fitted with a housing that can house each element of the sound output device 1.
  • the sound output device 1 is worn on the user's head. With the sound output device 1 worn on the user's head, the user can hear external sounds. However, if the user is paying attention to something else, they may miss external sounds containing necessary information. If the user feels that they have missed external sounds containing necessary information, they can have the sound output device 1 play the external sounds. By having the sound output device 1 play the external sounds, the user can confirm whether the external sounds contain the necessary information.
  • the sound output device 1 includes a microphone unit 10 that collects external sound data as an acquisition unit that acquires sound data.
  • the sound output device 1 also includes a speaker unit 11, an input unit 12, a memory unit 13, and a control unit 16.
  • the memory unit 13 and the control unit 16 may be housed in either the housing 1L or the housing 1R as shown in FIG. 1, or may be housed in a housing attached to the fixing member 1F.
  • the microphone unit 10 is capable of collecting external sounds around the sound output device 1.
  • the microphone unit 10 is composed of a left microphone and a right microphone.
  • the left microphone is housed in housing 1L.
  • the right microphone is housed in housing 1R.
  • the microphone unit 10 collects external sounds as stereo sounds using the left microphone and the right microphone.
  • the microphone unit 10 converts the collected external sounds into electrical signal data.
  • the microphone unit 10 outputs the external sound data after conversion into an electrical signal to the control unit 16.
  • the speaker unit 11 is capable of outputting sound.
  • the speaker unit 11 outputs sound by converting an electrical signal output from the control unit 16 into sound.
  • the speaker unit 11 is configured to include a left bone conduction speaker and a right bone conduction speaker.
  • the bone conduction speaker converts the electrical signal output from the control unit 16 into vibrations, thereby transmitting the vibrations to the user's skull.
  • the bone conduction speaker outputs sound to the user by transmitting the vibrations to the user's skull.
  • the left bone conduction speaker is housed in housing 1L.
  • the right bone conduction speaker is housed in housing 1R.
  • the input unit 12 is capable of accepting input from the user.
  • the input unit 12 is configured to include at least one input interface capable of accepting input from the user.
  • the input interface is, for example, a physical key or a capacitance key.
  • the input unit 12 is located on the surface of the housing 1L as shown in FIG. 1. However, the input unit 12 may also be located on the surface of the housing 1R.
  • the memory unit 13 is configured to include at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of these.
  • the semiconductor memory is, for example, RAM (Random Access Memory) or ROM (Read Only Memory).
  • the RAM is, for example, SRAM (Static Random Access Memory) or DRAM (Dynamic Random Access Memory).
  • the ROM is, for example, EEPROM (Electrically Erasable Programmable Read Only Memory).
  • the memory unit 13 may function as a main memory device, an auxiliary memory device, or a cache memory.
  • the memory unit 13 stores data used in the operation of the sound output device 1 and data obtained by the operation of the sound output device 1.
  • the memory unit 13 may also store programs executed by the control unit 16.
  • the memory unit 13 includes a buffer 14 and a category storage unit 15.
  • the buffer 14 stores data on external sounds collected by the microphone unit 10 that are determined to be recorded, as described below.
  • recording means storing sound data in the buffer 14.
  • the category storage unit 15 stores first data.
  • the first data includes data on sounds to be recorded.
  • the first data in this embodiment includes data on multiple sound categories (multiple categories), multiple score weights, and multiple feature quantities, as shown in Figure 3.
  • Sound categories may include any sound category. Sound categories may be set based on the user's usage patterns. In Figure 3, the sound categories are announcements, conversations between multiple people, speech by one person, sirens, screams, children playing, and loud voices. Hereinafter, the sound category of list number C will be referred to as "sound category C.”
  • the multiple score weights indicate the importance of sounds belonging to each of the multiple sound categories.
  • the score weight indicates the importance of recording sounds belonging to the corresponding sound category.
  • the sound category corresponding to a score weight means the sound category with the same list number C as the score weight in the listed first data.
  • the score weight corresponding to sound category C will be referred to as "score weight W C ". The higher the importance of recording sounds belonging to sound category C, the larger the score weight W C is set. The lower the importance of recording sounds belonging to sound category C, the smaller the score weight W C is set.
  • a sound feature indicates a sound feature belonging to a corresponding sound category.
  • the sound category corresponding to a sound feature means the sound category with the same list number C as the sound feature in the listed first data.
  • a sound feature corresponding to sound category C will be referred to as a "feature ⁇ C ".
  • a sound feature may be any sound feature that indicates a sound feature. Examples of sound features include Mel-Frequency Cepstrum Coefficients (MFCCs), formant frequencies, sound frequency spectra, or voiceprint values shown in a spectrogram. Examples of sound frequency spectra include amplitude spectra or phase spectra. When a sound feature is given as a string of numbers, it is also referred to as a "sound feature vector.”
  • a sound feature ⁇ C may be an average value, such as a time average value, of sound features belonging to sound category C.
  • the control unit 16 is configured to include at least one processor, at least one dedicated circuit, or a combination of these.
  • the processor is a general-purpose processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a dedicated processor specialized for specific processing.
  • the dedicated circuit is, for example, an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
  • the control unit 16 controls each part of the sound output device 1 and executes processing related to the operation of the sound output device 1.
  • the control unit 16 includes a determination unit 17, a selection unit 18, and an update unit 19.
  • the determination unit 17 acquires external sound data from the microphone unit 10. The determination unit 17 determines whether the acquired external sound is to be recorded. This process is explained below.
  • the determination unit 17 converts the stereo sound data acquired from the microphone unit 10 into monaural external sound data.
  • the monaural external sound data is used in the process of determining whether the external sound is to be recorded.
  • this process may use either one of the stereo external sound data instead of the monaural external sound data. Which of the stereo external sound data to use may be set in advance based on reliability, etc.
  • the determination unit 17 calculates the power of the external sound in the converted monaural sound.
  • the determination unit 17 calculates the power of the external sound, for example, by squaring the amplitude of the waveform data of the external sound. If the calculated power of the external sound is below a power threshold, the determination unit 17 determines that the external sound is not to be recorded.
  • the power threshold may be set based on the power of the external sound, such as an announcement sound, that the user wants to hear.
  • the determination unit 17 determines whether the external sound is to be recorded based on the first data in the category storage unit 15. First, the determination unit 17 calculates a similarity score S C.
  • the similarity score S C indicates the similarity between the external sound and a sound belonging to sound category C. The larger the similarity score S C , the higher the possibility that the external sound belongs to sound category C.
  • the determination unit 17 calculates the similarity score S C , for example, using equation (1) or equation (2). Equation (2) is an equation using cosine similarity.
  • Equation (1) 100-(x- ⁇ C ) 2 Equation (1)
  • the subscript C corresponds to the list number C in FIG.
  • the feature amount x in equation (1) and the feature amount x indicated with a superscript arrow in equation (2) are the feature amounts of the external sound.
  • the determination unit 17 calculates a total score S TC using a plurality of similarity scores S C and a plurality of score weights W C.
  • the determination unit 17 calculates the total score S TC , for example, by using equation (3).
  • S TC S C + W C formula (3)
  • the determination unit 17 After calculating the overall score S TC for each sound category C, the determination unit 17 acquires the overall score S T that is the maximum of the calculated overall scores S TC . The determination unit 17 determines whether the overall score S T is equal to or greater than an overall threshold. The overall threshold may be set taking into consideration user convenience. If the determination unit 17 determines that the overall score S T is equal to or greater than the overall threshold, it determines that the external sound is to be recorded. However, the determination unit 17 may also determine whether each overall score S TC is equal to or greater than the overall threshold. In this case, if the determination unit 17 determines that at least one of the multiple overall scores S TC is equal to or greater than the overall threshold, it determines that the external sound is to be recorded.
  • the determination unit 17 determines that the external sound is to be recorded, it starts recording the external sound. That is, the determination unit 17 stores the external sound data in the buffer 14. In this embodiment, the determination unit 17 stores the stereo external sound data in the buffer 14. However, the determination unit 17 may also store the monaural external sound data in the buffer 14.
  • the determination unit 17 After starting recording of the external sound, the determination unit 17 ends the recording of the external sound when the power of the external sound falls below a certain level or the external sound is no longer a target special sound. Through such processing by the determination unit 17, recording starts or ends at the timing shown in Figure 4, for example. In Figure 4, the determination unit 17 starts recording the external sound at times t0, t2, t4, and t6. The determination unit 17 ends recording the external sound at times t1, t3, t5, and t7. In Figure 4, external sound data D0, D1, D2, and D3 are stored in the buffer 14.
  • the selection unit 18 can accept recording and playback input (sound playback input) via the input unit 12. If the user feels that they have missed an external sound containing necessary information, they can input recording and playback input via the input unit 12.
  • the selection unit 18 When the selection unit 18 receives a recording/playback input from the input unit 12, it selects external sound data to be played back from the external sound data stored in the buffer 14. For example, the selection unit 18 selects external sound data from the external sound data stored in the buffer 14 up to a predetermined time prior to the time the recording/playback input was received.
  • the predetermined time may be set taking into consideration user convenience. For example, in FIG. 4, the recording/playback input is received at time t7. In FIG. 4, time t2 is a predetermined time prior to time t7. In FIG. 4, the selection unit 18 selects external sound data D1, D2, and D3 as the external sound data to be played back.
  • the selection unit 18 plays the selected external sound data by the speaker unit 11. For example, the selection unit 18 outputs the external sound collected by the left microphone of the microphone unit 10 from the left speaker of the speaker unit 11, out of the stereo sound data of the selected external sound. The selection unit 18 also outputs the external sound collected by the right microphone of the microphone unit 10 from the right speaker of the speaker unit 11, out of the stereo sound data of the selected external sound. By having the selection unit 18 play the external sound, the user can hear any external sound that they missed.
  • the update unit 19 When the update unit 19 receives a recording/playback input via the input unit 12, it updates the first data in the category storage unit 15 based on the external sound data stored in the buffer 14. Here, if the user feels that they have missed an external sound containing necessary information, they input a recording/playback input via the input unit 12. In other words, the external sound immediately prior to the time the recording/playback input was received is likely to contain a sound that provides the user with necessary information, i.e., a sound that should be recorded. Therefore, the update unit 19 updates the first data based on the data of the external sound immediately prior to the time the recording/playback input was received. "Most recent" means within a predetermined time from the time the input was received. The predetermined time may be set in advance taking into consideration user convenience. Alternatively, the control unit 16 may receive the predetermined time from the user via the input unit 12. The predetermined time may be, for example, within 5 seconds or within 5 phrases.
  • the update unit 19 may use external sound data from the first time to the second time as the external sound data immediately prior to the time when the recording/playback input was accepted.
  • the first time is the time when the recording/playback input was accepted.
  • the second time is a time that goes back a set amount of time from the first time.
  • the set amount of time may be set taking into account the time from when the user feels that they have missed an external sound containing necessary information until the recording/playback input is accepted.
  • the set amount of time may be 30 seconds, for example.
  • the recording/playback input is accepted at time t7. Therefore, in FIG. 4, the first time is time t7. Also, in FIG.
  • the update unit 19 uses external sound data D2 and D3 from time t4 to time t7 as the external sound data immediately prior to the time when the recording/playback input was accepted.
  • the update unit 19 may use, as the external sound data immediately preceding the time when the input for recording and playback was accepted, the external sound data recorded at the time when the input for recording and playback was accepted up to a predetermined number of times before.
  • the predetermined number may be set taking into consideration the time from when the user feels that they have missed an external sound containing necessary information until the input for recording and playback is accepted. For example, in FIG. 4, the predetermined number is 3.
  • the external sound data recorded at the time when the input for recording and playback was accepted is external sound data D3. Therefore, the update unit 19 uses the external sound data D1, D2, and D3 up to three times before external sound data D3.
  • the updating unit 19 When the updating unit 19 acquires data on the external sound immediately prior to the timing of receiving the recording/playback input, it identifies the sound category to which the acquired external sound belongs.
  • the sound category to which the external sound belongs will be referred to as the "external sound category.”
  • the updating unit 19 calculates a similarity score S C using, for example, Equation (1) or Equation (2), and if the calculated similarity score S C is equal to or greater than the similarity threshold, it identifies category C as the external sound category.
  • the updating unit 19 may use the data on the similarity score S C calculated by the determination unit 17 to identify the external sound category.
  • the update unit 19 When the update unit 19 identifies the category of the external sound, it updates the first data in the category storage unit 15 based on the identified category of the external sound. As an example, the update unit 19 increases the score weight W C of a sound category C that is the same as the category of the external sound in the first data in the category storage unit 15. The degree to which the score weight W C is increased may be set in advance taking into consideration the convenience of the user. As another example, if the category of the external sound does not exist in the first data in the category storage unit 15, the update unit 19 registers the category of the external sound as a new sound category C in the first data. In this case, the update unit 19 may calculate the sound feature amount ⁇ C of the new sound category C from the data of the external sound. Furthermore, the update unit 19 may register an initial value for the score weight W C of the new sound category C. This initial value may be set in advance according to the user's usage mode.
  • (Operation of sound output device) 5 is a flowchart showing an example of the procedure for the sound output process.
  • This sound output process corresponds to the sound output method according to the present embodiment. For example, when the power of the sound output device 1 is turned on, the sound output device 1 starts the process of step S1.
  • the microphone unit 10 collects external sounds around the sound output device 1 (step S1). The microphone unit 10 converts the collected external sounds into electrical signal data. The microphone unit 10 outputs the external sound data after the conversion into an electrical signal to the control unit 16.
  • the determination unit 17 acquires external sound data from the microphone unit 10 and determines whether the power of the acquired external sound is equal to or greater than the power threshold (step S2). If the determination unit 17 determines that the power of the external sound is equal to or greater than the power threshold (step S2: YES), the process proceeds to step S3. On the other hand, if the determination unit 17 determines that the power of the external sound is below the power threshold (step S2: NO), the process proceeds to step S4.
  • the determination unit 17 determines whether or not the external sound is to be recorded, based on the first data in the category storage unit 15. In the present embodiment, the determination unit 17 calculates a total score S T using formula (1) or (2) and formula (3). If the determination unit 17 determines that the total score S T is equal to or greater than the total threshold, it determines that the external sound is to be recorded.
  • step S3 determines that the external sound is not to be recorded (step S3: NO)
  • step S4 determines that the external sound is to be recorded
  • step S5 determines that the external sound is to be recorded
  • step S4 the determination unit 17 ends recording of external sound. However, if recording has not started before the processing of step S4, the determination unit 17 maintains a state in which external sound is not being recorded.
  • step S5 the determination unit 17 starts recording external sounds.
  • step S6 the control unit 16 determines whether or not a recording/playback input has been received by the input unit 12. If the control unit 16 does not determine that a recording/playback input has been received by the input unit 12 (step S6: NO), the sound output device 1 returns to the processing of step S1. On the other hand, if the control unit 16 determines that a recording/playback input has been received by the input unit 12 (step S6: YES), the control unit 16 proceeds to the processing of step S7.
  • step S7 the selection unit 18 selects external sound data to be played from the external sound data stored in the buffer 14.
  • the selection unit 18 plays the selected external sound data to be played back by the speaker unit 11.
  • step S8 when the update unit 19 receives input for recording and playback, it updates the first data in the category storage unit 15 based on the external sound data stored in the buffer 14.
  • the determination unit 17 may determine whether to automatically play the acquired external sound based on the multiple similarity scores and the multiple score weights. As an example, when the determination unit 17 determines that the total score ST is equal to or greater than the total threshold, the determination unit 17 may further determine whether the total score ST is equal to or greater than a first playback threshold. When the determination unit 17 determines that the total score ST is equal to or greater than the first playback threshold, the determination unit 17 may determine to automatically play the acquired external sound. When the determination unit 17 determines that the total score ST is below the first playback threshold, the determination unit 17 may determine not to automatically play the acquired external sound. When the determination unit 17 determines to automatically play the acquired external sound, the determination unit 17 automatically plays back data of the acquired external sound by the speaker unit 11. The first playback threshold is greater than the total threshold. The first playback threshold may be set based on user convenience.
  • the determination unit 17 determines whether or not an external sound is to be recorded, based on the first data including data on the sound to be recorded. By determining whether or not an external sound is to be recorded, it is possible to record external sound that includes information necessary for the user. In other words, only external sound data that includes information necessary for the user can be stored in the buffer 14. With this configuration, when external sound is played back, only external sound data that includes information necessary for the user is played back. Therefore, the user can hear only external sound that includes the necessary information.
  • the update unit 19 updates the first data in the category storage unit 15 based on the data of the external sound immediately prior to the time when the recording/playback input is received.
  • the external sound immediately prior to the time when the recording/playback input is received is likely to contain a sound that is information necessary for the user, i.e., a sound that should be recorded.
  • the updated first data reflects the data of the sound that is information necessary for the user. With this configuration, it is possible to accurately determine whether an external sound is to be recorded based on the updated first data.
  • this embodiment can improve sound recording technology.
  • a sound output device 1 calculates a first probability as to whether an external sound is to be recorded, and determines whether the external sound is to be recorded based on the calculated first probability.
  • a Gaussian Mixture Model (GMM) is used to calculate the first probability according to the other embodiment.
  • the category storage unit 15 stores first data as shown in FIG. 6.
  • the first data in another embodiment includes multiple sound categories, multiple second probabilities, multiple first feature vectors, and multiple second feature vectors.
  • the sound category data is listed and assigned a list number C (C is an integer satisfying 1 ⁇ C), in the same or similar manner as in FIG. 3.
  • Sound categories are set to be the same as or similar to the sound categories shown in Figure 3. As in Figure 3, the sound category of list number C is written as "Sound Category C.”
  • the second probability is the probability that the external sound belongs to the corresponding sound category.
  • the sound category corresponding to the second probability means the sound category with the same list number C as the second probability in the listed first data.
  • the second probability corresponding to sound category C will be referred to as the "second probability ⁇ C .”
  • the first feature vector is an average value, such as a time average, of the feature vectors of sounds belonging to the corresponding sound category.
  • the sound category corresponding to the first feature vector means the sound category with the same list number C as the first feature in the listed first data.
  • the first feature vector corresponding to sound category C will be referred to as the "first feature vector M C ".
  • the second feature vector is a covariance matrix of feature vectors of sounds belonging to a corresponding sound category.
  • the sound category corresponding to the second feature vector means the sound category with the same list number C as the second feature vector in the listed first data.
  • the second feature vector corresponding to sound category C will be referred to as the "second feature vector ⁇ C .”
  • the probability p(x) that the external sound is acquired is expressed by equation (4).
  • the second probability p(C) is the second probability that the external sound belongs to sound category C.
  • the subscript C corresponds to the list number C in FIG. 6.
  • the second probability p(C) corresponds to the second probability ⁇ C in FIG. 6, as shown on the right side of equation (4).
  • C) is the conditional probability that an external sound belonging to sound category C is acquired.
  • C) is expressed as N(x
  • the determination unit 17 calculates a first probability p(C
  • the selection unit 18 When the selection unit 18 receives a recording/playback input from the input unit 12, it selects the external sound data to be played back from the external sound data stored in the buffer 14, in the same or similar manner as in the above-described embodiment. The selection unit 18 plays back the external sound data selected for playback through the speaker unit 11, in the same or similar manner as in the above-described embodiment.
  • the updating unit 19 acquires data on the external sound immediately prior to the timing of receiving the recording/playback input, in the same manner as or similar to the embodiment described above.
  • the updating unit 19 updates the first data stored in the category storage unit 15 based on the acquired external sound data.
  • the updating unit 19 identifies the category of the external sound, in the same manner as or similar to the embodiment described above.
  • the updating unit 19 updates the second probability ⁇ C of the category C that is the same as the identified category of the external sound.
  • the updating unit 19 registers the category of the external sound in the first data as a new sound category C.
  • the updating unit 19 may calculate the second probability ⁇ C of the sound of the new sound category C, the first feature vector M C , and the second feature vector ⁇ C from the external sound data.
  • a sound output device 1 obtains a first probability p(C
  • D X ) is a first probability that the external sound D X belongs to a sound category C.
  • the machine learning model is machine-learned to output the first probability p(C
  • the data of the external sound D X may be data in any format depending on the machine learning model used.
  • the data of the external sound D X may be a feature vector of the external sound or waveform data of the external sound.
  • the category storage unit 15 stores neural network parameters.
  • the first data includes neural network parameters. These parameters are, for example, weighting and bias terms between nodes in each layer of the neural network.
  • a determination unit 17 acquires a first probability p(C
  • steps S4 to S7 is the same as or similar to that in the embodiment described above.
  • step S8 the update unit 19 updates the neural network parameters stored in the category storage unit 15 based on the external sound data stored in the buffer 14.
  • the determination unit 17 may determine whether or not to automatically play the external sound based on the first probability p(C
  • the determination unit 17 may determine not to automatically play the acquired external sound. If the determination unit 17 determines to automatically play the acquired external sound, the determination unit 17 automatically plays back data of the acquired external sound by the speaker unit 11.
  • the words to be recorded are proper nouns such as place names or personal names.
  • the words to be recorded are Yokohama and Tanaka.
  • the multiple scores indicate the importance of recording each of the sounds containing the multiple words to be recorded.
  • the scores indicate the importance of recording sounds containing the corresponding words to be recorded.
  • the words to be recorded corresponding to the scores are the words to be recorded associated with the scores in the first data. The higher the importance of recording sounds containing the corresponding words to be recorded, the higher the score is set. The lower the importance of recording sounds containing the corresponding words to be recorded, the higher the score is set.
  • the determination unit 17 determines whether the word of the acquired external sound matches any of the words to be recorded in the category storage unit 15. If the determination unit 17 determines that the word of the acquired external sound matches any of the words to be recorded in the category storage unit 15, it obtains a score for the recorded word that matches the word of the external sound. For example, in Figure 8, the determination unit 17 obtains the word "Yokohama" from the external sound data. The determination unit 17 references the first data in the category storage unit 15, as shown in Figure 7, and obtains a score of 80 associated with the word "Yokohama".
  • the determination unit 17 determines whether the acquired score is equal to or greater than the score threshold.
  • the score threshold may be set taking into consideration user convenience. If the determination unit 17 determines that the acquired score is equal to or greater than the score threshold, it determines that the external sound is to be recorded. If the determination unit 17 determines that the external sound is to be recorded, it stores external sound data corresponding to a sentence containing words that match the words to be recorded in the buffer 14. For example, if the determination unit 17 determines that the score of 80 associated with the word "Yokohama" in Figure 7 is equal to or greater than the score threshold, it stores external sound data corresponding to the sentence "I'm going to Yokohama today" as shown in Figure 8 in the buffer 14. If the determination unit 17 determines that the acquired score is below the score threshold, it determines that the external sound is not to be recorded.
  • the selection unit 18 When the selection unit 18 receives a recording/playback input from the input unit 12, it selects the external sound data to be played back from the external sound data stored in the buffer 14, in the same or similar manner as in the above-described embodiment. The selection unit 18 plays back the external sound data selected for playback through the speaker unit 11, in the same or similar manner as in the above-described embodiment.
  • the update unit 19 updates the first data stored in the category storage unit 15 based on the external sound data stored in the buffer 14.
  • the update unit 19 acquires external sound data immediately prior to the time the recording/playback input was received, in the same or similar manner as in the above-described embodiment.
  • the update unit 19 acquires word data from the external sound data immediately prior to the time the recording/playback input was received, in the same or similar manner as the processing of the determination unit 17. If the acquired word data is not stored in the category storage unit 15, the update unit 19 stores the acquired word data in the category storage unit 15 as a new word to be recorded. In this case, the update unit 19 may set an initial value for the score of the new word to be registered.
  • This initial value may be set in advance according to the user's usage pattern. If the acquired word data is stored in the category storage unit 15, the update unit 19 increases the score of the same word to be recorded as the acquired word. The degree to which the word score is increased may be set in advance in consideration of user convenience.
  • steps S1 and S2 are the same as or similar to that in the above-described embodiment.
  • the determination unit 17 acquires words from the external sound data.
  • the determination unit 17 determines whether the acquired word of the external sound matches any of the words to be recorded in the category storage unit 15. If the determination unit 17 determines that the acquired word of the external sound does not match any of the words to be recorded in the category storage unit 15, it determines that the external sound is not to be recorded (step S3: NO). On the other hand, if the determination unit 17 determines that the acquired word of the external sound matches any of the words to be recorded in the category storage unit 15, it acquires a score associated with the word to be recorded that matches the word of the external sound.
  • step S3 determines that the acquired score is equal to or greater than the score threshold. If the determination unit 17 determines that the external sound is to be recorded (step S3: YES). On the other hand, if the determination unit 17 determines that the acquired score is below the score threshold, it determines that the external sound is not to be recorded (step S3: NO).
  • steps S4 to S8 are the same as or similar to the embodiment described above.
  • the determination unit 17 may determine whether or not to automatically play the acquired external sound based on the score. As an example of this, in this case, if the determination unit 17 determines that the acquired score is equal to or greater than the score threshold, it may further determine whether or not the acquired score is equal to or greater than a third playback threshold. If the determination unit 17 determines that the acquired score is equal to or greater than the third playback threshold, it may determine that the acquired external sound will be automatically played. If the determination unit 17 determines that the acquired score is below the third playback threshold, it may determine that the acquired external sound will not be automatically played. If the determination unit 17 determines that the acquired external sound will be automatically played, it automatically plays the data of the acquired external sound by the speaker unit 11. The third playback threshold is greater than the score threshold. The third playback threshold may be set based on user convenience.
  • a sound output device 101 (output device) 9 , a sound output device 101 (output device) according to yet another embodiment is connectable to a network 2.
  • the network 2 may be any network including a mobile communication network, the Internet, etc.
  • the sound output device 101 determines whether or not the acquired external sound is to be recorded, based on the position information of the sound output device 101 and the first data.
  • Server 3 may be, for example, a dedicated computer configured to function as a server, a general-purpose personal computer, or a cloud computing system.
  • the server 3 stores multiple pieces of second data. Each piece of second data is associated with location information.
  • the second data includes data on sounds to be recorded at the associated location.
  • the location associated with the second data may be set with consideration for user convenience.
  • the location associated with the second data may be, for example, a train station or an airport.
  • the second data includes multiple sound categories, multiple score weights, and multiple feature data, similar to or identical to the first data.
  • Sound output device 101 is equipped with a microphone unit 10, a speaker unit 11, an input unit 12, a memory unit 13, and a control unit 16, which are the same as or similar to sound output device 1 shown in FIG. 2.
  • sound output device 101 is equipped with a positioning unit 20 and a communication unit 21.
  • Positioning unit 20 and communication unit 21 may be housed in either housing 1L or housing 1R as shown in FIG. 1, or may be housed in a housing attached to fixing member 1F.
  • the positioning unit 20 is capable of acquiring position information of the sound output device 101.
  • the positioning unit 20 is configured to include at least one receiving module compatible with a satellite positioning system.
  • the receiving module is, for example, a receiving module compatible with GPS (Global Positioning System).
  • the communication unit 21 is configured to include at least one communication module that can be connected to the network 2.
  • the communication module is, for example, a communication module that supports mobile communication standards such as LTE (Long Term Evolution), 4G (4th Generation), or 5G (5th Generation).
  • the control unit 16 acquires the position information of the sound output device 101 at predetermined time intervals using the positioning unit 20.
  • the predetermined time interval may be set taking into consideration the user's average moving speed, etc.
  • the control unit 16 transmits the position information of the sound output device 101 via the network 2 using the communication unit 21.
  • the server 3 receives location information of the sound output device 101 from the sound output device 101 via the network 2. Upon receiving the location information of the sound output device 101, the server 3 transmits, via the network 2, second data associated with the location information of the sound output device 101 from among the multiple second data stored on the server 3 to the sound output device 101. For example, it is assumed that the location of the sound output device 101 is within Station A. In this case, the server 3 transmits the second data associated with Station A to the sound output device 101.
  • the control unit 16 receives second data associated with the location information of the sound output device 101 from the server 3 via the network 2 using the communication unit 21.
  • the control unit 16 stores the received second data in the category storage unit 15.
  • the category storage unit 15 stores first data and second data as shown in Figure 10. This first data is the same as the first data shown in Figure 3.
  • the second data is data associated with Station A. In Figure 10, the first data and second data are collectively listed and assigned list number C.
  • the determination unit 17 calculates the similarity score SC using equation (1) or equation (2) in the same or similar manner as in the above-described embodiment.
  • the determination unit 17 calculates the overall score STC using equation (3) in the same or similar manner as in the above-described embodiment.
  • the determination unit 17 acquires the maximum overall score ST from the multiple overall scores STC in the same or similar manner as in the above-described embodiment.
  • the determination unit 17 determines whether the overall score ST is equal to or greater than an overall threshold.
  • the overall threshold may be the same as or different from the overall score in the above-described embodiment. If the determination unit 17 determines that the overall score ST is equal to or greater than the overall threshold, it determines that the external sound is to be recorded.
  • the determination unit 17 determines that the overall score ST is below the overall threshold, it determines that the external sound is not to be recorded. In the same or similar manner as in the above-described embodiment, the determination unit 17 may determine whether each overall score STC is equal to or greater than an overall threshold. In this case, when the determination unit 17 determines that at least one of the plurality of total scores S TC is equal to or greater than the total threshold value, it determines that the external sound is to be recorded.
  • the selection unit 18 When the selection unit 18 receives a recording/playback input from the input unit 12, it selects the external sound data to be played back from the external sound data stored in the buffer 14, in the same or similar manner as in the above-described embodiment. The selection unit 18 plays back the external sound data selected for playback through the speaker unit 11, in the same or similar manner as in the above-described embodiment.
  • the update unit 19 acquires data on the external sound immediately prior to the timing at which the recording/playback input was received, in the same or similar manner as in the above-described embodiment.
  • the update unit 19 updates the first data stored in the category storage unit 15 based on the acquired external sound data.
  • the update unit 19 transmits the data on the external sound immediately prior to the timing at which the recording/playback input was received, along with location information of the sound output device 101, to the server 3 via the network 2 via the communication unit 21.
  • the server 3 can update the second data associated with the location information.
  • the control unit 16 acquires position information of the sound output device 101 at predetermined time intervals using the positioning unit 20. As described above, the control unit 16 transmits the user's position information to the server 3, thereby receiving second data associated with the user's position information from the server 3. The control unit 16 stores the received second data in the category storage unit 15.
  • steps S1 to S7 is the same as or similar to that in the embodiment described above.
  • step S8 when the update unit 19 receives a recording/playback input via the input unit 12, it updates the first data in the category storage unit 15 based on the external sound data stored in the buffer 14, in the same or similar manner as in the above-described embodiment.
  • the update unit 19 transmits the external sound data immediately prior to the time the recording/playback input was received, along with location information of the sound output device 101, to the server 3 via the network 2 via the communication unit 21.
  • the determination unit 17 may determine whether to automatically play the acquired external sound based on the multiple similarity scores and the multiple score weights, in the same or similar manner as in the above-described embodiment. As an example, when the determination unit 17 determines that the total score ST is equal to or greater than the total threshold, the determination unit 17 may determine whether the total score ST is equal to or greater than a first playback threshold, in the same or similar manner as in the above-described embodiment. When the determination unit 17 determines that the total score ST is equal to or greater than the first playback threshold, the determination unit 17 may determine to automatically play the acquired external sound. When the determination unit 17 determines that the total score ST is below the first playback threshold, the determination unit 17 may determine not to automatically play the acquired external sound. When the determination unit 17 determines to automatically play the acquired external sound, the determination unit 17 automatically plays back data of the acquired external sound via the speaker unit 11.
  • each functional unit, means, or step can be added to other embodiments so as not to cause logical inconsistencies, or can be replaced with each functional unit, means, or step of other embodiments.
  • multiple functional units, means, or steps can be combined into one or separated.
  • the above-described embodiments of the present disclosure are not limited to faithful implementation of each of the described embodiments, but can also be implemented by combining features or omitting some features as appropriate.
  • the sound output device 1, 101 has been described as a hearable device such as a bone conduction earphone as shown in FIG. 1.
  • the sound output device 1, 101 does not have to be a hearable device.
  • the sound output device 1, 101 may be a device such as a smartphone.
  • the sound output device 1, 101 may further include a communication unit that acquires sound data from an earphone or the like, instead of the microphone unit 10 and speaker unit 11.
  • the sound output device 1, 101 may transmit external sound data stored in the buffer 14 to the earphone or the like via the communication unit.
  • a general-purpose computer functions as the sound output device 1 or sound output device 101 according to the above-described embodiments.
  • a program describing the processing content for realizing each function of the sound output device 1 or sound output device 101 according to the above-described embodiments is stored in the memory of the general-purpose computer, and the program is read and executed by a processor. Therefore, the present disclosure can also be realized as a program executable by a processor, or a non-transitory computer-readable medium that stores the program.
  • an output device is an acquisition unit that acquires sound data; a control unit that determines whether or not the sound acquired by the acquisition unit is to be recorded, based on first data including data of the sound to be recorded; When the control unit receives input for sound playback, the control unit updates the first data based on sound data acquired by the acquisition unit within a predetermined time from the time when the input for sound playback was received.
  • the first data is Multiple categories and a plurality of score weights indicating the importance of sounds belonging to each of the plurality of categories;
  • the control unit may determine whether the acquired sound is to be recorded based on the similarity between the acquired sound and sounds belonging to each of the multiple categories and the multiple score weights.
  • the control unit may determine whether to automatically play back the acquired sound based on the plurality of similarities and the plurality of score weights.
  • the control unit may identify a category to which a sound within a predetermined time from the timing of receiving the input for sound playback belongs, and update the first data based on the identified category.
  • the control unit calculating a total score based on the plurality of similarities and the plurality of score weights; If the total score is equal to or greater than a total threshold, the acquired sound may be determined to be a recording target, and if the total score is equal to or greater than a first playback threshold, the acquired sound may be automatically played back.
  • the control unit calculates a first probability that the sound acquired by the acquisition unit is the recording target based on the first data; When it is determined that the calculated first probability is equal to or greater than a probability threshold, it may be determined that the sound acquired by the acquisition unit is to be recorded.
  • the first data includes a plurality of sound categories, a plurality of second probabilities, a plurality of first feature vectors, and a plurality of second feature vectors;
  • the second probability is a second probability as to whether the sound acquired by the acquisition unit belongs to the corresponding sound category;
  • the first feature vector is an average value of feature vectors of sounds belonging to the corresponding sound category,
  • the second feature vector is a covariance matrix of feature vectors of sounds belonging to the corresponding sound category,
  • the control unit may calculate the first probability using a feature vector of the sound acquired by the acquisition unit and the first data.
  • the first data includes parameters of a neural network used in a machine learning model; the machine learning model is machine-learned to output the first probability when sound data acquired by the acquisition unit is input,
  • the control unit may calculate the first probability using the machine learning model.
  • the control unit may update parameters of the neural network based on sound data within a predetermined time period from the timing at which the input for sound playback is received.
  • the control unit may automatically play back the sound acquired by the acquisition unit when the first probability is equal to or greater than a second playback threshold.
  • the first data is A plurality of words to be recorded; a plurality of scores indicating the importance of recording each of the sounds comprising the plurality of target words;
  • the control unit may determine that the acquired sound is to be recorded when a score of a word included in the sound acquired by the acquisition unit is equal to or greater than a threshold value.
  • the control unit In the output device described in (11), The control unit acquiring word data from sound data within a predetermined time period from the timing of receiving the sound playback input; If the acquired word matches the word to be recorded, increase the score of the word to be recorded; If the acquired word does not match the word to be recorded, the acquired word may be registered in the first data as a new word to be recorded.
  • the control unit may automatically play back the sound acquired by the acquisition unit when the score is equal to or greater than a third playback threshold.
  • a positioning unit capable of acquiring position information of the output device
  • the control unit may determine whether or not the sound acquired by the acquisition unit is to be recorded, based on the location information acquired from the positioning unit and the first data.
  • a sound output method includes: Acquiring sound data; determining whether the acquired sound is to be recorded based on first data including data of the sound to be recorded; When an input for sound playback is received, the first data is updated based on data of the acquired sounds that are present within a predetermined time from the timing at which the input for sound playback is received.
  • references such as “first” and “second” are identifiers used to distinguish the configuration.
  • Configurations distinguished by descriptions such as “first” and “second” in this disclosure may have their numbers swapped.
  • first data may swap the identifiers “first” and “second” with second data.
  • the identifier swapping occurs simultaneously.
  • the configurations remain distinguished even after the identifier swapping.
  • Identifiers may be deleted. Configurations from which identifiers have been deleted are distinguished by their symbols. Identifier descriptions such as “first” and “second” in this disclosure should not be used solely to interpret the order of the configurations or to justify the existence of identifiers with smaller numbers.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Multimedia (AREA)
  • Reverberation, Karaoke And Other Acoustics (AREA)

Abstract

出力装置は、音のデータを取得する取得部と、制御部とを備える。制御部は、記録対象となる音のデータを含む第1データに基づいて、取得部によって取得された音が記録対象であるか否かを判定する。制御部は、音再生の入力を受け付けると、取得部によって取得された音のうちの音再生の入力を受け付けたタイミングから所定時間内の音のデータに基づいて、第1データを更新する。

Description

出力装置及び音出力方法 関連出願へのクロスリファレンス
 本出願は、2024年5月17日に日本国に特許出願された特願2024-080942の優先権を主張するものであり、この先の出願の開示全体をここに参照のために取り込む。
 本開示は、出力装置及び音出力方法に関する。
 従来、ユーザの周囲の音を録音する技術が知られている。例えば、特許文献1には、外部音を集音する集音手段を具備する携帯音楽再生装置が記載されている。
特開2001-256771号公報
 本開示の一実施形態に係る出力装置は、
 音のデータを取得する取得部と、
 記録対象となる音のデータを含む第1データに基づいて、前記取得部によって取得された音が記録対象であるか否かを判定する制御部と、を備え、
 前記制御部は、音再生の入力を受け付けると、前記取得部によって取得された音のうちの前記音再生の入力を受け付けたタイミングから所定時間内の音のデータに基づいて、前記第1データを更新する。
 本開示の一実施形態に係る音出力方法は、
 音のデータを取得することと、
 記録対象となる音のデータを含む第1データに基づいて、前記取得された音が記録対象であるか否かを判定することと、
 音再生の入力を受け付けると、前記取得された音のうちの前記音再生の入力を受け付けたタイミングから所定時間内の音のデータに基づいて、前記第1データを更新することと、を含む。
一実施形態に係る音出力装置の概略構成を示す図である。 図2に示す音出力装置のブロック図である。 一実施形態に係る第1データの一例を示す図である。 録音開始及び録音終了の動作の一例を示すタイミングチャートである。 音出力処理の手順例を示すフローチャートである。 他の実施形態に係る第1データの一例を示す図である。 さらに他の実施形態に係る第1データの一例を示す図である。 さらに他の実施形態に係る判定部の処理を説明するための図である。 さらに他の実施形態に係る音出力装置のブロック図である。 さらに他の実施形態に係る第1データ及び第2データの一例を示す図である。
 音を録音する従来の技術には、改善の余地がある。本開示の一実施形態によれば、音を録音する技術を改善することができる。
 以下、本開示に係る実施形態について、図面を参照して説明する。
 図1に示すような音出力装置1(出力装置)は、ヒアラブルデバイスである。本実施形態では、音出力装置1は、骨伝導イヤホンである。ただし、音出力装置1は、ヒアラブルデバイスであれば、骨伝導イヤホンに限定されない。他の例として、音出力装置1は、耳掛け型イヤホン、首掛け型スピーカ、インナーイヤー型イヤホン、カナル型イヤホン、ヘッドホン又は補聴器であってもよい。音出力装置1は、インナーイヤー型イヤホン又はヘッドホンである場合、外部音の取り込み機能を有してよい。外部音の取り込み機能は、音出力装置1の外部音を集音してユーザに出力する機能である。外部音とは、音出力装置1の外部で発せられる音である。一例として、外部音には、ユーザの周囲で発せられる音が含まれる。外部音には、ユーザ自身が発する音が含まれてよい。
 音出力装置1は、筐体1Lと、筐体1Rと、固定部材1Fとを含む。筐体1Lは、ユーザの左側のこめかみ部分に当てられる。筐体1Rは、ユーザの右側のこめかみ部分に当てられる。固定部材1Fは、筐体1L及び筐体1Rをそれぞれユーザの左側及び右側のこめかみ部分に固定する。固定部材1Fは、ユーザの左耳に掛けられる左用のイヤーフックと、ユーザの右耳に掛けられる右用のイヤーフックと、これらのイヤーフックを接続するバンドとを含む。固定部材1Fは、音出力装置1の各要素を収容可能な筐体が取り付けられてもよい。
 音出力装置1は、ユーザの頭部に装着される。ユーザは、音出力装置1を頭部に装着した状態で、外部音を聞くことができる。しかしながら、ユーザは、他の物事に注意を向けていると、必要な情報を含む外部音を聞き逃してしまうことがある。ユーザは、必要な情報を含む外部音を聞き逃したと感じた場合、音出力装置1に外部音を再生させることができる。音出力装置1に外部音を再生させることにより、ユーザは、必要な情報が外部音に含まれるか否かを確認することができる。
 図2に示すように、本実施形態に係る音出力装置1は、音のデータを取得する取得部として、外部音のデータを集音するマイク部10を備える。音出力装置1は、マイク部10に加えて、スピーカ部11と、入力部12と、記憶部13と、制御部16とを備える。記憶部13及び制御部16は、図1に示すような、筐体1L及び筐体1Rの何れかに収容されてもよいし、固定部材1Fに取り付けられる筐体に収容されてもよい。
 マイク部10は、音出力装置1の周囲の外部音を集音可能である。マイク部10は、左用のマイク及び右用のマイクを含んで構成される。左用のマイクは、筐体1Lに収容される。右用のマイクは、筐体1Rに収容される。マイク部10は、左用のマイク及び右用のマイクによって、外部音をステレオ音として集音する。マイク部10は、集音した外部音を電気信号のデータに変換する。マイク部10は、電気信号に変換後の外部音のデータを制御部16に出力する。
 スピーカ部11は、音を出力可能である。スピーカ部11は、制御部16から出力される電気信号を音に変換することにより、音を出力する。本実施形態では、スピーカ部11は、左用の骨伝導スピーカ及び右用の骨伝導スピーカを含んで構成される。骨伝導スピーカは、制御部16から出力される電気信号を振動に変換することにより、ユーザの頭蓋骨に振動を伝達させる。骨伝導スピーカは、ユーザの頭蓋骨に振動を伝達させることにより、音をユーザに対して出力する。左用の骨伝導スピーカは、筐体1Lに収容される。右用の骨伝導スピーカは、筐体1Rに収容される。
 入力部12は、ユーザからの入力を受け付け可能である。入力部12は、ユーザからの入力を受け付け可能な少なくとも1つの入力用インタフェースを含んで構成される。入力用インタフェースは、例えば、物理キー又は静電容量キー等である。入力部12は、図1に示すように筐体1Lの表面に位置する。ただし、入力部12は、筐体1Rの表面に位置してもよい。
 記憶部13は、少なくとも1つの半導体メモリ、少なくとも1つの磁気メモリ、少なくとも1つの光メモリ又はこれらのうちの少なくとも2種類の組み合わせを含んで構成される。半導体メモリは、例えば、RAM(Random Access Memory)又はROM(Read Only Memory)等である。RAMは、例えば、SRAM(Static Random Access Memory)又はDRAM(Dynamic Random Access Memory)等である。ROMは、例えば、EEPROM(Electrically Erasable Programmable Read Only Memory)等である。記憶部13は、主記憶装置、補助記憶装置又はキャッシュメモリとして機能してよい。記憶部13は、音出力装置1の動作に用いられるデータと、音出力装置1の動作によって得られたデータとを記憶する。記憶部13には、制御部16によって実行されるプログラムが記憶されてよい。
 記憶部13は、バッファ14と、カテゴリ格納部15とを含む。
 バッファ14には、マイク部10によって集音された外部音のデータのうち、後述するように記録対象であると判定された外部音のデータが記憶される。本実施形態において、録音とは、バッファ14に音のデータを記憶させることを意味する。
 カテゴリ格納部15には、第1データが格納される。第1データは、記録対象となる音のデータを含む。本実施形態に係る第1データは、図3に示すような、複数の音のカテゴリ(複数のカテゴリ)と、複数のスコア重みと、複数の特徴量のデータとを含む。図3では、第1データは、リスト化されてリスト番号C(Cは、1≦Cを満たす整数)が付される。
 音のカテゴリは、任意の音のカテゴリを含んでよい。音のカテゴリは、ユーザの使用態様に基づいて設定されてよい。図3では、音のカテゴリは、アナウンス、複数人での会話、一人のスピーチ、サイレン、悲鳴、子供の遊ぶ声及び大声である。以下、リスト番号Cの音のカテゴリは、「音のカテゴリC」と記載される。
 複数のスコア重みは、複数の音のカテゴリのそれぞれに属する音の重要度を示す。本実施形態では、スコア重みは、対応する音のカテゴリに属する音を録音する重要度を示す。スコア重みに対応する音のカテゴリとは、リスト化された第1データにおいてスコア重みと同じリスト番号Cの音のカテゴリを意味する。以下、音のカテゴリCに対応するスコア重みは、「スコア重みW」と記載される。音のカテゴリCに属する音を録音する重要度が高いほど、スコア重みWは、大きく設定される。音のカテゴリCに属する音を録音する重要度が低いほど、スコア重みWは、小さく設定される。
 音の特徴量は、対応する音のカテゴリに属する音の特徴量を示す。音の特徴量に対応する音のカテゴリとは、リスト化された第1データにおいて音の特徴量と同じリスト番号Cの音のカテゴリを意味する。以下、音のカテゴリCに対応する音の特徴量は、「特徴量μ」と記載される。音の特徴量は、音の特徴を示すものであれば、任意のものであってよい。音の特徴量は、例えば、メル周波数ケプストラム係数(MFCC:Mel-Frequency Cepstrum Coefficient)、フォルマント周波数、音の周波数スペクトル又はスペクトログラムに示された声紋の値等である。音の周波数スペクトルは、例えば、振幅スペクトル又は位相スペクトル等である。音の特徴量は、数字の文字列で与えられる場合、「音の特徴量ベクトル」とも称される。音の特徴量μは、音のカテゴリCに属する音の特徴量の時間平均値等の平均値であってよい。
 制御部16は、少なくとも1つのプロセッサ、少なくとも1つの専用回路又はこれらの組み合わせを含んで構成される。プロセッサは、CPU(Central Processing Unit)若しくはGPU(Graphics Processing Unit)等の汎用プロセッサ又は特定の処理に特化した専用プロセッサである。専用回路は、例えば、FPGA(Field-Programmable Gate Array)又はASIC(Application Specific Integrated Circuit)等である。制御部16は、音出力装置1の各部を制御しながら、音出力装置1の動作に関わる処理を実行する。
 制御部16は、判定部17と、選択部18と、更新部19とを含む。
 判定部17は、マイク部10から外部音のデータを取得する。判定部17は、取得した外部音が記録対象であるか否かを判定する。以下、この処理について説明する。
 まず、判定部17は、マイク部10から取得したステレオ音のデータをモノラル音の外部音のデータに変換する。本実施形態では、外部音が記録対象であるか否かを判定する処理には、モノラル音の外部音のデータが用いられる。ただし、この処理には、モノラル音の外部音のデータの代わりに、ステレオ音の外部音のデータの何れか一方が用いられてもよい。ステレオ音の外部音のデータの何れを用いるかは、信頼度等に基づいて予め設定されてよい。
 判定部17は、変換後のモノラル音の外部音のパワーを算出する。判定部17は、例えば、外部音の波形データの振幅を二乗することにより、外部音のパワーを算出する。判定部17は、算出した外部音のパワーがパワー閾値を下回る場合、外部音が記録対象ではないと判定する。パワー閾値は、ユーザが聞きたいアナウンス音等の外部音のパワーに基づいて設定されてよい。
 判定部17は、算出した外部音のパワーがパワー閾値以上である場合、カテゴリ格納部15の第1データに基づいて、外部音が記録対象であるか否かを判定する。まず、判定部17は、類似度スコアSを算出する。類似度スコアSは、外部音と音のカテゴリCに属する音との間の類似度を示す。類似度スコアSが大きいほど、外部音が音のカテゴリCに属する可能性が高くなる。判定部17は、例えば、式(1)又は式(2)によって、類似度スコアSを算出する。式(2)は、コサイン類似度(cosine similarity)を用いた式である。
 S=100-(x-μ        式(1)
 式(1)及び式(2)において、添え字Cは、図3におけるリスト番号Cに対応する。
 式(1)の特徴量x及び式(2)の上付き矢印とともに示される特徴量xは、外部音の特徴量である。
 式(1)又は式(2)から、音のカテゴリCに属する音と外部音とが完全に一致する場合、類似度スコアSが100になることが分かる。一方、音のカテゴリCに属する音と外部音との間の類似度が低いほど、類似度スコアSが小さくなることが分かる。
 判定部17は、複数の類似度スコアSと複数のスコア重みWとによって、総合スコアSTCを算出する。判定部17は、例えば、式(3)によって、総合スコアSTCを算出する。
 STC=S+W             式(3)
 判定部17は、各音のカテゴリCについての総合スコアSTCを算出すると、算出した複数の総合スコアSTCのうちから最大値となる総合スコアSを取得する。判定部17は、総合スコアSが総合閾値以上であるか否かを判定する。総合閾値は、ユーザの利便性を考慮して設定されてよい。判定部17は、総合スコアSが総合閾値以上であると判定した場合、外部音が記録対象であると判定する。ただし、判定部17は、各総合スコアSTCが総合閾値以上であるか否かを判定してもよい。この場合、判定部17は、複数の総合スコアSTCの少なくとも1つが総合閾値以上であると判定した場合、外部音が記録対象であると判定する。
 判定部17は、外部音が記録対象であると判定した場合、外部音の録音を開始する。つまり、判定部17は、外部音のデータをバッファ14に記憶させる。本実施形態では、判定部17は、ステレオ音の外部音のデータをバッファ14に記憶させる。ただし、判定部17は、モノラル音の外部音のデータをバッファ14に記憶させてもよい。
 判定部17は、外部音の録音を開始した後、外部音のパワーが下回るか又は外部音が特音対象ではなくなると、外部音の録音を終了する。このような判定部17の処理によって、例えば、図4に示すようなタイミングで録音が開始又は終了する。図4では、判定部17は、時刻t0,t2,t4,t6で、外部音の録音を開始する。判定部17は、時刻t1,t3,t5,t7で、外部音の録音を終了する。図4では、外部音のデータD0,D1,D2,D3がバッファ14に記憶される。
 選択部18は、録音再生の入力(音再生の入力)を入力部12によって受け付け得る。ユーザは、必要な情報を含む外部音を聞き逃したと感じた場合、録音再生の入力を入力部12から入力する。
 選択部18は、録音再生の入力を入力部12によって受け付けると、バッファ14に記憶させた外部音のデータのうちから、再生する外部音のデータを選択する。選択部18は、例えば、バッファ14に記憶させた外部音のデータのうち、録音再生の入力を受け付けた時刻から所定時間過去に遡った時刻までの外部音のデータを選択する。所定時間は、ユーザの利便性を考慮して設定されてよい。例えば、図4では、時刻t7で、録音再生の入力が受け付けられる。図4では、時刻t2は、時刻t7から所定時間過去に遡った時刻となる。図4では、選択部18は、外部音のデータD1,D2,D3を再生する外部音のデータとして選択する。
 選択部18は、再生すると選択した外部音のデータをスピーカ部11によって再生する。例えば、選択部18は、選択した外部音であるステレオ音のデータのうち、マイク部10の左用のマイクによって集音した外部音をスピーカ部11の左用のスピーカに出力させる。また、選択部18は、選択した外部音であるステレオ音のデータのうち、マイク部10の右用のマイクによって集音した外部音をスピーカ部11の右用のスピーカに出力させる。選択部18によって外部音が再生されることにより、ユーザは、聞き逃した外部音を聞くことができる。
 更新部19は、録音再生の入力を入力部12によって受け付けると、バッファ14に記憶された外部音のデータに基づいて、カテゴリ格納部15の第1データを更新する。ここで、ユーザは、必要な情報を含む外部音を聞き逃したと感じると、録音再生の入力を入力部12から入力する。つまり、録音再生の入力を受け付けたタイミングの直近の外部音には、ユーザにとって必要な情報となる音すなわち記録対象とすべき音が含まれる可能性が高い。そこで、更新部19は、外部音のうちで録音再生の入力を受け付けた直近の外部音のデータに基づいて、第1データを更新する。直近とは、入力を受け付けたタイミングから所定時間以内のことである。所定時間は、ユーザの利便性を考慮して予め設定されてよい。又は、所定時間は、制御部16がユーザから入力部12を介して受け付けてもよい。所定時間以内は、例えば、5秒以内又は5文節以内であってよい。
 更新部19は、録音再生の入力を受け付けたタイミングの直近の外部音のデータとして、第1時刻から第2時刻までの外部音のデータを用いてよい。第1時刻は、録音再生の入力を受け付けた時刻である。第2時刻は、第1時刻から過去に設定時間遡った時刻である。設定時間は、必要な情報を含む外部音をユーザが聞き逃したと感じてから録音再生の入力を受け付けるまでの時間を考慮して設定されてよい。設定時間は、例えば、30秒である。例えば、図4では、上述したように、時刻t7で、録音再生の入力が受け付けられる。そのため、図4では、第1時刻は、時刻t7となる。また、図4では、時刻t7から過去に設定時間遡った時刻は、時刻t4となる。そのため、第2時刻は、時刻t4となる。更新部19は、録音再生の入力を受け付けたタイミングの直近の外部音のデータとして、時刻t4から時刻t7までの外部音のデータD2,D3を用いる。
 更新部19は、録音再生の入力を受け付けたタイミングの直近の外部音のデータとして、録音再生の入力を受け付けたタイミングで録音された外部音のデータから所定数前までの外部音のデータを用いてもよい。所定数は、必要な情報を含む外部音をユーザが聞き逃したと感じてから録音再生の入力を受け付けるまでの時間を考慮して設定されてよい。例えば、図4では、所定数は、3であるものとする。図4では、録音再生の入力を受け付けたタイミングで録音された外部音のデータは、外部音のデータD3となる。そのため、更新部19は、外部音のデータD3から3つ前までの外部音のデータD1,D2,D3を用いる。
 更新部19は、録音再生の入力を受け付けたタイミングの直近の外部音のデータを取得すると、取得した外部音が属する音のカテゴリを特定する。以下、外部音が属する音のカテゴリは、「外部音のカテゴリ」と記載される。更新部19は、例えば、式(1)又は式(2)によって類似度スコアSを算出し、算出した類似度スコアSが類似度閾値以上である場合、カテゴリCを外部音のカテゴリとして特定する。又は、更新部19は、判定部17が算出した類似度スコアSのデータを用いて、外部音のカテゴリを特定してよい。
 更新部19は、外部音のカテゴリを特定すると、特定した外部音のカテゴリに基づいて、カテゴリ格納部15の第1データを更新する。一例として、更新部19は、カテゴリ格納部15の第1データのうち、外部音のカテゴリと同じ音のカテゴリCのスコア重みWを増加させる。スコア重みWを増加させる度合いは、ユーザの利便性を考慮して予め設定されてよい。他の例として、更新部19は、外部音のカテゴリがカテゴリ格納部15の第1データに存在しない場合、外部音のカテゴリを新たな音のカテゴリCとして第1データに登録する。この場合、更新部19は、新たな音のカテゴリCの音の特徴量μを外部音のデータから算出してよい。また、更新部19は、新たな音のカテゴリCのスコア重みWには、初期値を登録してよい。この初期値は、ユーザの使用態様に応じて予め設定されてよい。
 (音出力装置の動作)
 図5は、音出力処理の手順例を示すフローチャートである。この音出力処理は、本実施形態に係る音出力方法に相当する。音出力装置1は、例えば、音出力装置1の電源がオンになると、ステップS1の処理を開始する。
 マイク部10は、音出力装置1の周囲の外部音を集音する(ステップS1)。マイク部10は、集音した外部音を電気信号のデータに変換する。マイク部10は、電気信号に変換後の外部音のデータを制御部16に出力する。
 判定部17は、マイク部10から外部音のデータを取得し、取得した外部音のパワーがパワー閾値以上であるか否かを判定する(ステップS2)。判定部17は、外部音のパワーがパワー閾値以上であると判定した場合(ステップS2:YES)、ステップS3の処理に進む。一方、判定部17は、外部音のパワーがパワー閾値を下回ると判定した場合(ステップS2:NO)、ステップS4の処理に進む。
 ステップS3の処理では、判定部17は、カテゴリ格納部15の第1データに基づいて、外部音が記録対象であるか否かを判定する。本実施形態では、判定部17は、式(1)又は式(2)と、式(3)とによって、総合スコアSを算出する。判定部17は、総合スコアSが総合閾値以上であると判定した場合、外部音が記録対象であると判定する。
 判定部17は、外部音が記録対象ではないと判定した場合(ステップS3:NO)、ステップS4の処理に進む。一方、判定部17は、外部音が記録対象であると判定した場合(ステップS3:YES)、ステップS5の処理に進む。
 ステップS4の処理では、判定部17は、外部音の録音を終了する。ただし、判定部17は、ステップS4の処理前に録音が開始されていない場合、外部音を録音しない状態を維持する。
 ステップS5の処理では、判定部17は、外部音の録音を開始する。
 ステップS6の処理では、制御部16は、録音再生の入力を入力部12によって受け付けたか否かを判定する。制御部16が録音再生の入力を入力部12によって受け付けたと判定しない場合(ステップS6:NO)、音出力装置1は、ステップS1の処理に戻る。一方、制御部16は、録音再生の入力を入力部12によって受け付けたと判定した場合(ステップS6:YES)、ステップS7の処理に進む。
 ステップS7の処理では、選択部18は、バッファ14に記憶させた外部音のデータのうちから、再生する外部音のデータを選択する。選択部18は、再生すると選択した外部音のデータをスピーカ部11によって再生する。
 ステップS8の処理では、更新部19は、録音再生の入力を受け付けると、バッファ14に記憶された外部音のデータに基づいて、カテゴリ格納部15の第1データを更新する。
 ここで、ステップS3の処理において、判定部17は、複数の類似度スコアと複数のスコア重みとに基づいて、取得した外部音を自動的に再生するか否かを判定してもよい。この一例として、判定部17は、総合スコアSが総合閾値以上であると判定した場合、さらに総合スコアSが第1再生閾値以上であるか否かを判定してもよい。判定部17は、総合スコアSが第1再生閾値以上であると判定した場合、取得した外部音を自動的に再生すると判定してよい。判定部17は、総合スコアSが第1再生閾値を下回ると判定した場合、取得した外部音を自動的に再生しないと判定してよい。判定部17は、取得した外部音を自動的に再生すると判定した場合、取得した外部音のデータをスピーカ部11によって自動的に再生する。第1再生閾値は、総合閾値よりも大きい。第1再生閾値は、ユーザの利便性に基づいて設定されてもよい。
 このように本実施形態に係る音出力装置1では、判定部17は、記録対象となる音のデータを含む第1データに基づいて、外部音が記録対象であるか否かを判定する。外部音が記録対象であるか否かを判定することにより、ユーザにとって必要な情報を含む外部音を録音することができる。つまり、バッファ14には、ユーザにとって必要な情報を含む外部音のデータのみを記憶させることができる。このような構成により、外部音が再生される際、ユーザにとって必要な情報を含む外部音のデータのみが再生される。したがって、ユーザは、必要な情報を含む外部音のみを聞くことができる。
 さらに、本実施形態では、更新部19は、録音再生の入力を受け付けたタイミングの直近の外部音のデータに基づいて、カテゴリ格納部15の第1データを更新する。上述したように、ユーザは、必要な情報を含む外部音を聞き逃したと感じると、録音再生の入力を入力部12から入力する。つまり、録音再生の入力を受け付けたタイミングの直近の外部音には、ユーザにとって必要な情報となる音すなわち記録対象とすべき音が含まれる可能性が高い。録音再生の入力を受け付けたタイミングの直近の外部音のデータに基づいてカテゴリ格納部15の第1データを更新することにより、更新後の第1データは、ユーザにとって必要な情報となる音のデータを反映したものとなる。このような構成により、更新後の第1データに基づいて、外部音が記録対象であるか否かを精度よく判定することができる。
 よって、本実施形態によれば、音を録音する技術を改善することができる。
 (他の実施形態に係る音出力装置)
 他の実施形態に係る音出力装置1は、外部音が記録対象であるかについての第1確率を算出し、算出した第1確率に基づいて外部音が記録対象であるか否かを判定する。他の実施形態に係る第1確率の算出には、混合ガウス分布(GMM:Gaussian Mixture Model)が用いられる。
 他の実施形態に係るカテゴリ格納部15には、図6に示すような、第1データが格納される。他の実施形態に係る第1データは、複数の音のカテゴリと、複数の第2確率と、複数の第1特徴量ベクトルと、複数の第2特徴量ベクトルとを含む。図6では、図3と同じ又は類似に、音のカテゴリのデータは、リスト化されてリスト番号C(Cは、1≦Cを満たす整数)が付される。
 音のカテゴリは、図3に示すような音のカテゴリと同じ又は類似に設定される。図3と同じく、リスト番号Cの音のカテゴリは、「音のカテゴリC」と記載される。
 第2確率は、対応する音のカテゴリに外部音が属するかについての確率である。第2確率に対応する音のカテゴリとは、リスト化された第1データにおいて第2確率と同じリスト番号Cの音のカテゴリを意味する。以下、音のカテゴリCに対応する第2確率は、「第2確率ω」と記載される。
 第1特徴量ベクトルは、対応する音のカテゴリに属する音の特徴ベクトルの時間平均等の平均値である。第1特徴量ベクトルに対応する音のカテゴリとは、リスト化された第1データにおいて第1特徴量と同じリスト番号Cの音のカテゴリを意味する。以下、音のカテゴリCに対応する第1特徴量ベクトルは、「第1特徴量ベクトルΜ」と記載される。
 第2特徴量ベクトルは、対応する音のカテゴリに属する音の特徴量ベクトルの共分散行列である。第2特徴量ベクトルに対応する音のカテゴリとは、リスト化された第1データにおいて第2特徴量ベクトルと同じリスト番号Cの音のカテゴリを意味する。以下、音のカテゴリCに対応する第2特徴量ベクトルは、「第2特徴量ベクトルΣ」と記載される。
 ここで、混合ガウス分布を用いると、外部音の特徴量ベクトルを「特徴量ベクトルx」と記載する場合、その外部音が取得される確率p(x)は、式(4)によって表される。
 式(4)において、第2確率p(C)は、音のカテゴリCに外部音が属するかについての第2確率である。添え字Cは、図6におけるリスト番号Cに対応する。第2確率p(C)は、式(4)の右辺に示されるように、図6における第2確率ωに対応する。
 式(4)において、確率p(x|C)は、音のカテゴリCに属する外部音が取得される条件付き確率である。確率p(x|C)は、式(4)の右辺に示されるように、図6に示すような第1特徴量ベクトルμ及び第2特徴量ベクトルΣを用いると、N(x|μ,Σ)と表される。
 式(4)から式(5)が導かれる。式(5)は、第1確率p(C|x)を算出する式である。第1確率p(C|x)は、特徴量ベクトルxを有する外部音が音のカテゴリCに属する第1確率である。
 他の実施形態に係る判定部17は、外部音の特徴量ベクトルxと、カテゴリ格納部15の第1データと、式(5)とによって、第1確率p(C|x)を算出する。判定部17は、算出した第1確率p(C|x)が確率閾値以上であるか否かを判定する。判定部17は、算出した第1確率p(C|x)が確率閾値以上であると判定した場合、外部音が記録対象であると判定する。判定部17は、算出した第1確率p(C|x)が確率閾値を下回ると判定した場合、外部音が記録対象ではないと判定する。
 選択部18は、録音再生の入力を入力部12によって受け付けると、上述した実施形態と同じ又は類似に、バッファ14に記憶させた外部音のデータのうちから、再生する外部音のデータを選択する。上述した実施形態と同じ又は類似に、選択部18は、再生すると選択した外部音のデータをスピーカ部11によって再生する。
 更新部19は、録音再生の入力を入力部12によって受け付けると、上述した実施形態と同じ又は類似に、録音再生の入力を受け付けたタイミングの直近の外部音のデータを取得する。更新部19は、取得した外部音のデータに基づいて、カテゴリ格納部15に格納された第1データを更新する。一例として、更新部19は、上述した実施形態と同じ又は類似に、外部音のカテゴリを特定する。更新部19は、特定した外部音のカテゴリと同じカテゴリCの第2確率ωを更新する。他の例として、更新部19は、外部音のカテゴリがカテゴリ格納部15の第1データに存在しない場合、外部音のカテゴリを新たな音のカテゴリCとして第1データに登録する。更新部19は、新たな音のカテゴリCの音の第2確率ω、第1特徴量ベクトルM及び第2特徴量ベクトルΣを外部音のデータから算出してよい。
 (他の実施形態に係る音出力装置の動作)
 他の実施形態に係る音出力装置1の動作は、図5に示すフローチャートによって説明される。
 ステップS1,S2の処理は、上述した実施形態と同じ又は類似である。
 ステップS3の処理では、判定部17は、外部音の特徴量ベクトルxと、カテゴリ格納部15の第1データと、式(5)とによって、第1確率p(C|x)を算出する。判定部17は、算出した第1確率p(C|x)が確率閾値以上であると判定した場合、外部音が記録対象であると判定する(ステップS3:YES)。一方、判定部17は、算出した第1確率p(C|x)が確率閾値を下回ると判定した場合、外部音が記録対象ではないと判定する(ステップS3:NO)。
 ステップS4~S8の処理は、上述した実施形態と同じ又は類似である。
 ここで、ステップS3の処理において、判定部17は、第1確率p(C|x)に基づいて、取得した外部音を自動的に再生するか否かを判定してもよい。この一例として、判定部17は、第1確率p(C|x)が確率閾値以上であると判定した場合、第1確率p(C|x)が第2再生閾値以上であるか否かを判定してもよい。判定部17は、第1確率p(C|x)が第2再生閾値以上であると判定した場合、取得した外部音を自動的に再生すると判定してよい。判定部17は、第1確率p(C|x)が第2再生閾値を下回ると判定した場合、取得した外部音を自動的に再生しないと判定してよい。判定部17は、取得した外部音を自動的に再生すると判定した場合、取得した外部音のデータをスピーカ部11によって自動的に再生する。第2再生閾値は、確率閾値よりも大きい。第2再生閾値は、ユーザの利便性に基づいて設定されてもよい。
 他の実施形態に係る音出力装置1のその他の構成及び効果は、上述した実施形態と同じ又は類似である。
 (さらに他の実施形態に係る音出力装置)
 さらに他の実施形態に係る音出力装置1は、機械学習モデルによって、第1確率p(C|D)を取得する。第1確率p(C|D)は、外部音Dが音のカテゴリCに属するかについての第1確率である。機械学習モデルは、外部音Dのデータが入力されると、第1確率p(C|D)を出力するように機械学習したものである。外部音Dのデータは、用いられる機械学習モデルに応じた任意の形式のデータであってよい。一例として、外部音Dのデータは、外部音の特徴量ベクトルであってもよいし、外部音の波形データであってもよい。機械学習モデルの学習データは、例えば、外部音の波形又は特徴量と、その外部音が属するカテゴリCをラベル付けしたものである。機械学習モデルは、ニューラルネットワークを用いたものである。ニューラルネットワークは、入力層と、中間層と、出力層とを含む。ニューラルネットワークの各層は、少なくとも1つのノードを含む。
 さらに他の実施形態に係るカテゴリ格納部15には、ニューラルネットワークのパラメータが格納される。つまり、さらに他の実施形態に係る第1データは、ニューラルネットワークのパラメータを含む。このパラメータは、例えば、ニューラルネットワークの各層のノード間の重み付け及びバイアス項等である。
 さらに他の実施形態に係る判定部17は、外部音Dのデータを機械学習モデルに入力することにより、第1確率p(C|D)を取得する。判定部17は、取得した第1確率p(C|D)が確率閾値以上であるか否かを判定する。判定部17は、取得した第1確率p(C|D)が確率閾値以上であると判定した場合、外部音が記録対象であると判定する。判定部17は、取得した第1確率p(C|D)が確率閾値を下回ると判定した場合、外部音が記録対象ではないと判定する。
 選択部18は、録音再生の入力を入力部12によって受け付けると、上述した実施形態と同じ又は類似に、バッファ14に記憶させた外部音のデータのうちから、再生する外部音のデータを選択する。上述した実施形態と同じ又は類似に、選択部18は、再生すると選択した外部音のデータをスピーカ部11によって再生する。
 更新部19は、録音再生の入力を入力部12によって受け付けると、上述した実施形態と同じ又は類似に、録音再生の入力を受け付けたタイミングの直近の外部音のデータを取得する。更新部19は、取得した外部音のデータに基づいて、カテゴリ格納部15に格納されたニューラルネットワークのパラメータを更新する。
 (さらに他の実施形態に係る音出力装置の動作)
 さらに他の実施形態に係る音出力装置1の動作は、図5に示すフローチャートによって説明される。
 ステップS1,S2の処理は、上述した実施形態と同じ又は類似である。
 ステップS3の処理では、判定部17は、外部音Dのデータを機械学習モデルに入力することにより、第1確率p(C|D)を取得する。判定部17は、取得した第1確率p(C|D)が確率閾値以上であると判定した場合、外部音が記録対象であると判定する(ステップS3:YES)。判定部17は、取得した第1確率p(C|D)が確率閾値を下回ると判定した場合、外部音が記録対象ではないと判定する(ステップS3:NO)。
 ステップS4~S7の処理は、上述した実施形態と同じ又は類似である。
 ステップS8の処理では、更新部19は、バッファ14に記憶された外部音のデータに基づいて、カテゴリ格納部15に格納されたニューラルネットワークのパラメータを更新する。
 ここで、ステップS3の処理において、判定部17は、上述した実施形態と同じ又は類似に、第1確率p(C|D)に基づいて、外部音を自動的に再生するか否かを判定してもよい。この一例として、判定部17は、第1確率p(C|D)が確率閾値以上であると判定した場合、上述した実施形態と同じ又は類似に、第1確率p(C|D)が第2再生閾値以上であるか否かを判定してもよい。上述した実施形態と同じ又は類似に、判定部17は、第1確率p(C|D)が第2再生閾値以上であると判定した場合、取得した外部音を自動的に再生すると判定してよい。判定部17は、第1確率p(C|D)が第2再生閾値を下回ると判定した場合、取得した外部音を自動的に再生しないと判定してよい。判定部17は、取得した外部音を自動的に再生すると判定した場合、取得した外部音のデータをスピーカ部11によって自動的に再生する。
 (さらに他の実施形態に係る音出力装置)
 さらに他の実施形態に係る音出力装置1は、外部音に含まれる単語によって、外部音が記録対象であるか否かを判定する。
 さらに他の実施形態に係るカテゴリ格納部15には、図7に示すような、第1データが格納される。第1データは、複数の記録対象の単語と、複数のスコアとを含む。第1データにおいて、複数の記録対象の単語のそれぞれと、複数のスコアのそれぞれとは、対応付けられる。
 記録対象の単語は、地名又は人名等の固有名詞である。図7では、記録対象の単語は、横浜及び田中である。
 複数のスコアは、複数の記録対象の単語を含む音のそれぞれを録音する重要度を示す。本実施形態では、スコアは、対応する記録対象の単語を含む音を録音する重要度を示す。スコアに対応する記録対象の単語は、第1データにおいてスコアに対応付けられた記録対象の単語である。スコアは、対応する記録対象の単語を含む音を録音する重要度が高いほど、大きく設定される。スコアは、対応する記録対象の単語を含む音を録音する重要度が低いほど、大きく設定される。
 さらに他の実施形態に係る判定部17は、外部音のデータに対して音声認識処理を実行する。判定部17は、外部音のデータに対して音声認識処理を実行することにより、外部音に含まれる文字列のデータを取得する。判定部17は、取得した外部音の文字列のデータに対して形態素解析を実行することにより、外部音の文字列から形態素のデータを取得する。判定部17は、外部音の形態素のデータから単語のデータを取得する。例えば、判定部17は、外部音のデータに対して音声認識処理を実行することにより、図8に示すような「今日は横浜に行きます」との文字列のデータを取得したものとする。この場合、判定部17は、取得した文字列のデータに対して形態素解析を実行することにより、「今日」、「は」、「横浜」、「に」、「行き」及び「ます」との形態素のデータを取得する。判定部17は、形態素のデータから「横浜」との単語のデータを取得する。
 判定部17は、取得した外部音の単語がカテゴリ格納部15の記録対象の単語の何れかと一致するか否かを判定する。判定部17は、取得した外部音の単語がカテゴリ格納部15の記録対象の単語の何れかと一致すると判定した場合、外部音の単語と一致する記録対象の単語のスコアを取得する。例えば、図8では、判定部17は、外部音のデータから「横浜」との単語を取得する。判定部17は、カテゴリ格納部15の図7に示すような第1データを参照し、「横浜」との単語に対応付けられたスコアの80を取得する。
 判定部17は、取得したスコアがスコア閾値以上であるか否かを判定する。スコア閾値は、ユーザの利便性を考慮して設定されてよい。判定部17は、取得したスコアがスコア閾値以上であると判定した場合、外部音が記録対象であると判定する。判定部17は、外部音が記録対象であると判定した場合、記録対象の単語と一致する単語を含む外部音について、一文に対応する外部音のデータをバッファ14に記憶させる。例えば、判定部17は、図7における、「横浜」との単語に対応付けられたスコアの80がスコア閾値以上であると判定した場合、図8に示すような「今日は横浜に行きます」との一文に対応する外部音のデータをバッファ14に記憶させる。判定部17は、取得したスコアがスコア閾値を下回ると判定した場合、外部音が記録対象ではないと判定する。
 選択部18は、録音再生の入力を入力部12によって受け付けると、上述した実施形態と同じ又は類似に、バッファ14に記憶させた外部音のデータのうちから、再生する外部音のデータを選択する。上述した実施形態と同じ又は類似に、選択部18は、再生すると選択した外部音のデータをスピーカ部11によって再生する。
 更新部19は、録音再生の入力を入力部12によって受け付けると、バッファ14に記憶された外部音のデータに基づいて、カテゴリ格納部15に格納された第1データを更新する。更新部19は、上述した実施形態と同じ又は類似に、録音再生の入力を受け付けたタイミングの直近の外部音のデータを取得する。更新部19は、判定部17の処理と同じ又は類似にして、録音再生の入力を受け付けたタイミングの直近の外部音のデータから単語のデータを取得する。更新部19は、取得した単語のデータがカテゴリ格納部15に格納されていない場合、取得した単語のデータを新たな記録対象の単語としてカテゴリ格納部15に格納する。この場合、更新部19は、新たな登録対象の単語のスコアには初期値を設定してよい。この初期値は、ユーザの使用態様に応じて予め設定されてよい。更新部19は、取得した単語のデータがカテゴリ格納部15に格納されている場合、取得した単語と同じ記録対象の単語のスコアを増加させる。単語のスコアを増加させる度合いは、ユーザの利便性を考慮して予め設定されてよい。
 (さらに他の実施形態に係る音出力装置の動作)
 さらに他の実施形態に係る音出力装置1の動作は、図5に示すフローチャートによって説明される。
 ステップS1,S2の処理は、上述した実施形態と同じ又は類似である。
 ステップS3の処理では、判定部17は、外部音のデータから単語を取得する。判定部17は、取得した外部音の単語がカテゴリ格納部15の記録対象の単語の何れかと一致するか否かを判定する。判定部17は、取得した外部音の単語がカテゴリ格納部15の記録対象の単語の何れとも一致しないと判定した場合、外部音が記録対象ではないと判定する(ステップS3:NO)。一方、判定部17は、取得した外部音の単語がカテゴリ格納部15の記録対象の単語の何れかと一致すると判定した場合、外部音の単語と一致する記録対象の単語に対応付けられたスコアを取得する。判定部17は、取得したスコアがスコア閾値以上であると判定した場合、外部音が記録対象であると判定する(ステップS3:YES)。一方、判定部17は、取得したスコアがスコア閾値を下回ると判定した場合、外部音が記録対象ではないと判定する(ステップS3:NO)。
 ステップS4~S8の処理は、上述した実施形態と同じ又は類似である。
 ここで、ステップS3の処理において、判定部17は、スコアに基づいて、取得した外部音を自動的に再生するか否かを判定してもよい。この一例として、この場合、判定部17は、取得したスコアがスコア閾値以上であると判定した場合、さらに取得したスコアが第3再生閾値以上であるか否かを判定してもよい。判定部17は、取得したスコアが第3再生閾値以上であると判定した場合、取得した外部音を自動的に再生すると判定してよい。判定部17は、取得したスコアが第3再生閾値を下回ると判定した場合、取得した外部音を自動的に再生しないと判定してよい。判定部17は、取得した外部音を自動的に再生すると判定した場合、取得した外部音のデータをスピーカ部11によって自動的に再生する。第3再生閾値は、スコア閾値よりも大きい。第3再生閾値は、ユーザの利便性に基づいて設定されてもよい。
 (さらに他の実施形態に係る音出力装置)
 図9に示すように、さらに他の実施形態に係る音出力装置101(出力装置)は、ネットワーク2に接続可能である。ネットワーク2は、移動体通信網及びインターネット等を含む任意のネットワークであってよい。音出力装置101は、音出力装置101の位置情報と第1データとに基づいて、取得された外部音が記録対象であるか否かを判定する。
 サーバ3は、例えば、サーバとして機能するように構成された専用のコンピュータ、汎用のパーソナルコンピュータ又はクラウドコンピューティングシステム等である。
 サーバ3には、複数の第2データが格納される。複数の第2データには、それぞれ、位置情報が対応付けられる。第2データは、対応付けられた位置における記録対象となる音のデータを含む。第2データに対応付けられる位置は、ユーザの利便性を考慮して設定されてよい。第2データに対応付けられる位置は、例えば、駅構内又は空港等である。第2データは、第1データと同じ又は類似に、複数の音のカテゴリと、複数のスコア重みと、複数の特徴量のデータとを含む。
 音出力装置101は、図2に示すような音出力装置1と同じ又は類似に、マイク部10と、スピーカ部11と、入力部12と、記憶部13と、制御部16とを備える。これらに加えて、音出力装置101は、測位部20と、通信部21とを備える。測位部20及び通信部21は、図1に示すような、筐体1L及び筐体1Rの何れかに収容されてもよいし、固定部材1Fに取り付けられる筐体に収容されてもよい。
 測位部20は、音出力装置101の位置情報を取得可能である。測位部20は、衛星測位システムに対応する少なくとも1つの受信モジュールを含んで構成される。受信モジュールは、例えば、GPS(Global Positioning System)に対応した受信モジュールである。
 通信部21は、ネットワーク2に接続可能な少なくとも1つの通信モジュールを含んで構成される。通信モジュールは、例えば、LTE(Long Term Evolution)、4G(4th Generation)又は5G(5th Generation)等の移動体通信規格に対応した通信モジュールである。
 制御部16は、所定時間毎に、音出力装置101の位置情報を測位部20によって取得する。所定時間は、ユーザの平均移動速度等を考慮して設定されてよい。音出力装置101がユーザの頭部に装着されることにより、音出力装置101の位置情報は、ユーザの位置情報となる。制御部16は、ネットワーク2を介して音出力装置101の位置情報を通信部21によって送信する。
 サーバ3は、ネットワーク2を介して音出力装置101から、音出力装置101の位置情報を受信する。サーバ3は、音出力装置101の位置情報を受信すると、サーバ3に格納された複数の第2データのうちから、音出力装置101の位置情報が対応付けられた第2データをネットワーク2を介して音出力装置101に送信する。例えば、音出力装置101の位置は、A駅構内であるものとする。この場合、サーバ3は、A駅構内に対応付けられた第2データを音出力装置101に送信する。
 制御部16は、ネットワーク2を介してサーバ3から、音出力装置101の位置情報が対応付けられた第2データを通信部21によって受信する。制御部16は、受信した第2データをカテゴリ格納部15に格納させる。
 カテゴリ格納部15には、図10に示すような第1データ及び第2データが格納される。この第1データは、図3に示す第1データと同じである。第2データは、A駅構内に対応付けられたデータである。図10では、第1データ及び第2データは、まとめてリスト化されてリスト番号Cが付される。
 判定部17は、上述した実施形態と同じ又は類似に、式(1)又は式(2)によって類似度スコアSを算出する。判定部17は、上述した実施形態と同じ又は類似に、式(3)によって、総合スコアSTCを算出する。判定部17は、上述した実施形態と同じ又は類似に、複数の総合スコアSTCのうちから最大値となる総合スコアSを取得する。判定部17は、総合スコアSが総合閾値以上であるか否かを判定する。総合閾値は、上述した実施形態における総合スコアと同じであってもよいし、異なってもよい。判定部17は、総合スコアSが総合閾値以上であると判定した場合、外部音が記録対象であると判定する。判定部17は、総合スコアSが総合閾値を下回ると判定した場合、外部音が記録対象ではないと判定する。上述した実施形態と同じ又は類似に、判定部17は、各総合スコアSTCが総合閾値以上であるか否かを判定してもよい。この場合、判定部17は、複数の総合スコアSTCの少なくとも1つが総合閾値以上であると判定した場合、外部音が記録対象であると判定する。
 選択部18は、録音再生の入力を入力部12によって受け付けると、上述した実施形態と同じ又は類似に、バッファ14に記憶させた外部音のデータのうちから、再生する外部音のデータを選択する。上述した実施形態と同じ又は類似に、選択部18は、再生すると選択した外部音のデータをスピーカ部11によって再生する。
 更新部19は、録音再生の入力を入力部12によって受け付けると、上述した実施形態と同じ又は類似に、録音再生の入力を受け付けたタイミングの直近の外部音のデータを取得する。更新部19は、取得した外部音のデータに基づいて、カテゴリ格納部15に格納された第1データを更新する。さらに他の実施形態では、更新部19は、音出力装置101の位置情報とともに、録音再生の入力を受け付けたタイミングの直近の外部音のデータを、ネットワーク2を介してサーバ3に通信部21によって送信する。このような構成により、サーバ3は、位置情報に対応付けられた第2データを更新することができる。
 (さらに他の実施形態に係る音出力装置の動作)
 さらに他の実施形態に係る音出力装置1の動作は、図5に示すフローチャートによって説明される。ただし、制御部16は、図5に示す処理と並行して、所定時間毎に、音出力装置101の位置情報を測位部20によって取得する。上述したように、制御部16は、ユーザの位置情報をサーバ3に送信することにより、サーバ3からユーザの位置情報に対応付けられた第2データを受信する。制御部16は、受信した第2データをカテゴリ格納部15に格納させる。
 ステップS1~S7の処理は、上述した実施形態と同じ又は類似である。
 ステップS8の処理では、更新部19は、録音再生の入力を入力部12によって受け付けると、上述した実施形態と同じ又は類似に、バッファ14に記憶された外部音のデータに基づいて、カテゴリ格納部15の第1データを更新する。さらに他の実施形態では、更新部19は、音出力装置101の位置情報とともに、録音再生の入力を受け付けたタイミングの直近の外部音のデータを、ネットワーク2を介してサーバ3に通信部21によって送信する。
 ここで、ステップS3の処理において、判定部17は、上述した実施形態と同じ又は類似に、複数の類似度スコアと複数のスコア重みとに基づいて、取得した外部音を自動的に再生するか否かを判定してもよい。この一例として、判定部17は、総合スコアSが総合閾値以上であると判定した場合、上述した実施形態と同じ又は類似に、総合スコアSが第1再生閾値以上であるか否かを判定してもよい。判定部17は、総合スコアSが第1再生閾値以上であると判定した場合、取得した外部音を自動的に再生すると判定してよい。判定部17は、総合スコアSが第1再生閾値を下回ると判定した場合、取得した外部音を自動的に再生しないと判定してよい。判定部17は、取得した外部音を自動的に再生すると判定した場合、取得した外部音のデータをスピーカ部11によって自動的に再生する。
 本開示を諸図面及び実施例に基づき説明してきたが、当業者であれば本開示に基づき種々の変形又は修正を行うことが容易であることに注意されたい。したがって、これらの変形又は修正は本開示の範囲に含まれることに留意されたい。例えば、各機能部に含まれる機能等は論理的に矛盾しないように再配置可能である。複数の機能部等は、1つに組み合わせられたり、分割されたりしてよい。上述した本開示に係る各実施形態は、それぞれ説明した各実施形態に忠実に実施することに限定されるものではなく、適宜、各特徴を組み合わせたり、一部を省略したりして実施され得る。つまり、本開示の内容は、当業者であれば本開示に基づき種々の変形及び修正を行うことができる。したがって、これらの変形及び修正は本開示の範囲に含まれる。例えば、各実施形態において、各機能部、各手段又は各ステップ等は論理的に矛盾しないように他の実施形態に追加し、若しくは、他の実施形態の各機能部、各手段又は各ステップ等と置き換えることが可能である。また、各実施形態において、複数の各機能部、各手段又は各ステップ等を1つに組み合わせたり、或いは分割したりすることが可能である。また、上述した本開示の各実施形態は、それぞれ説明した各実施形態に忠実に実施することに限定されるものではなく、適宜、各特徴を組み合わせたり、一部を省略したりして実施することもできる。
 例えば、上述した実施形態では、音出力装置1,101は、図1に示すような骨伝導イヤホン等のヒアラブルデバイスであるものとして説明した。ただし、音出力装置1,101は、ヒアラブルデバイスではなくてもよい。他の例として、音出力装置1,101は、スマートフォン等の装置であってもよい。この場合、音出力装置1,101は、マイク部10及びスピーカ部11の代わりに、イヤホン等から音のデータを取得する通信部をさらに備えてもよい。また、音出力装置1,101は、当該通信部によって、バッファ14に記憶された外部音のデータをイヤホン等に送信してもよい。
 例えば、汎用のコンピュータを、上述した実施形態に係る音出力装置1又は音出力装置101として機能させる実施形態も可能である。具体的には、上述した実施形態に係る音出力装置1又は音出力装置101の各機能を実現する処理内容を記述したプログラムを、汎用のコンピュータのメモリに格納し、プロセッサによって当該プログラムを読み出して実行させる。したがって、本開示は、プロセッサが実行可能なプログラム、又は、当該プログラムを記憶する非一時的なコンピュータ可読媒体としても実現可能である。
 一実施形態において、(1)出力装置は、
 音のデータを取得する取得部と、
 記録対象となる音のデータを含む第1データに基づいて、前記取得部によって取得された音が記録対象であるか否かを判定する制御部と、を備え、
 前記制御部は、音再生の入力を受け付けると、前記取得部によって取得された音のうちの前記音再生の入力を受け付けたタイミングから所定時間内の音のデータに基づいて、前記第1データを更新する。
 (2)上記(1)に記載の出力装置において、
 前記第1データは、
  複数のカテゴリと、
  前記複数のカテゴリのそれぞれに属する音の重要度を示す複数のスコア重みと、を含み、
 前記制御部は、前記複数のカテゴリのそれぞれに属する音と前記取得された音との類似度と、前記複数のスコア重みとに基づいて、前記取得された音が記録対象であるか否かを判定してもよい。
 (3)上記(2)に記載の出力装置において、
 前記制御部は、前記複数の類似度と前記複数のスコア重みとに基づいて、前記取得された音を自動的に再生するか否かを判定してもよい。
 (4)上記(2)又は(3)に記載の出力装置において、
 前記制御部は、前記音再生の入力を受け付けたタイミングから所定時間内の音が属するカテゴリを特定し、特定した当該カテゴリに基づいて、前記第1データを更新してもよい。
 (5)上記(2)から(4)までの何れか1つに記載の出力装置において、
 前記制御部は、
  前記複数の類似度と前記複数のスコア重みとによって総合スコアを算出し、
  前記総合スコアが総合閾値以上である場合、前記取得された音を記録対象であると判定し、前記総合スコアが第1再生閾値以上である場合、前記取得された音を自動的に再生してもよい。
 (6)上記(1)から(5)までの何れか1つに記載の出力装置において、
 前記制御部は、前記第1データに基づいて、前記取得部によって取得された音が前記記録対象であるかについての第1確率を算出し、
 算出した前記第1確率が確率閾値以上であると判定した場合、前記取得部によって取得された音が記録対象であると判定してもよい。
 (7)上記(6)に記載の出力装置において、
 前記第1データは、複数の音のカテゴリと、複数の第2確率と、複数の第1特徴量ベクトルと、複数の第2特徴量ベクトルとを含み、
 前記第2確率は、対応する前記音のカテゴリに前記取得部によって取得された音が属するかについての第2確率であり、
 前記第1特徴量ベクトルは、対応する前記音のカテゴリに属する音の特徴量ベクトルの平均値であり、
 前記第2特徴量ベクトルは、対応する前記音のカテゴリに属する音の特徴量ベクトルの共分散行列であり、
 前記制御部は、前記取得部によって取得された音の特徴量ベクトルと前記第1データとによって、前記第1確率を算出してもよい。
 (8)上記(6)に記載の出力装置において、
 前記第1データは、機械学習モデルに用いられるニューラルネットワークのパラメータを含み、
 前記機械学習モデルは、前記取得部によって取得された音のデータが入力されると、前記第1確率を出力するように機械学習したものであり、
 前記制御部は、前記機械学習モデルによって前記第1確率を算出してもよい。
 (9)上記(8)に記載の出力装置において、
 前記制御部は、前記音再生の入力を受け付けたタイミングから所定時間内の音のデータに基づいて、前記ニューラルネットワークのパラメータを更新してもよい。
 (10)上記(6)から(9)までの何れか1つに記載の出力装置において、
 前記制御部は、前記第1確率が第2再生閾値以上である場合、前記取得部によって取得された音を自動的に再生してもよい。
 (11)上記(1)に記載の出力装置において、
 前記第1データは、
  複数の記録対象の単語と、
  前記複数の記録対象の単語を含む音のそれぞれを録音する重要度を示す複数のスコアと、を含み、
 前記制御部は、前記取得部によって取得された音に含まれる単語のスコアが閾値以上である場合、前記取得された音が記録対象であると判定してもよい。
 (12)上記(11)に記載の出力装置において、
 前記制御部は、
 前記音再生の入力を受け付けたタイミングから所定時間内の音のデータから単語のデータを取得し、
 取得した前記単語が前記記録対象の単語と一致する場合、前記記録対象の単語のスコアを増加させ、
 取得した前記単語が前記記録対象の単語と一致しない場合、取得した前記単語を新たな記録対象の単語として前記第1データに登録してもよい。
 (13)上記(11)又は(12)に記載の出力装置において、
 前記制御部は、前記スコアが第3再生閾値以上である場合、前記取得部によって取得された音を自動的に再生してもよい。
 
 (14)上記(1)に記載の出力装置において、
 前記出力装置の位置情報を取得可能な測位部をさらに備え、
 前記制御部は、前記測位部から取得した前記位置情報と前記第1データとに基づいて、前記取得部によって取得された音が記録対象であるか否かを判定してもよい。
 一実施形態において、(15)音出力方法は、
 音のデータを取得することと、
 記録対象となる音のデータを含む第1データに基づいて、前記取得された音が記録対象であるか否かを判定することと、
 音再生の入力を受け付けると、前記取得された音のうちの前記音再生の入力を受け付けたタイミングから所定時間内の音のデータに基づいて、前記第1データを更新することと、を含む。
 本開示において「第1」及び「第2」等の記載は、当該構成を区別するための識別子である。本開示における「第1」及び「第2」等の記載で区別された構成は、当該構成における番号を交換することができる。例えば、第1データは、第2データと識別子である「第1」と「第2」とを交換することができる。識別子の交換は同時に行われる。識別子の交換後も当該構成は区別される。識別子は削除してよい。識別子を削除した構成は、符号で区別される。本開示における「第1」及び「第2」等の識別子の記載のみに基づいて、当該構成の順序の解釈、小さい番号の識別子が存在することの根拠に利用してはならない。
 1,101 音出力装置(1F:固定部材、1L:筐体、1R:筐体、10:マイク部、11:スピーカ部、12:入力部、13:記憶部、14:バッファ、15:カテゴリ格納部、16:制御部、17:判定部、18:選択部、19:更新部、20:測位部、21:通信部)
 2 ネットワーク
 3 サーバ

Claims (15)

  1.  音のデータを取得する取得部と、
     記録対象となる音のデータを含む第1データに基づいて、前記取得部によって取得された音が記録対象であるか否かを判定する制御部と、を備え、
     前記制御部は、音再生の入力を受け付けると、前記取得部によって取得された音のうちの前記音再生の入力を受け付けたタイミングから所定時間内の音のデータに基づいて、前記第1データを更新する、出力装置。
  2.  前記第1データは、
      複数のカテゴリと、
      前記複数のカテゴリのそれぞれに属する音の重要度を示す複数のスコア重みと、を含み、
     前記制御部は、前記複数のカテゴリのそれぞれに属する音と前記取得された音との類似度と、前記複数のスコア重みとに基づいて、前記取得された音が記録対象であるか否かを判定する、請求項1に記載の出力装置。
  3.  前記制御部は、前記複数の類似度と前記複数のスコア重みとに基づいて、前記取得された音を自動的に再生するか否かを判定する、請求項2に記載の出力装置。
  4.  前記制御部は、前記音再生の入力を受け付けたタイミングから所定時間内の音が属するカテゴリを特定し、特定した当該カテゴリに基づいて、前記第1データを更新する、請求項2又は3に記載の出力装置。
  5.  前記制御部は、
      前記複数の類似度と前記複数のスコア重みとによって総合スコアを算出し、
      前記総合スコアが総合閾値以上である場合、前記取得された音を記録対象であると判定し、前記総合スコアが第1再生閾値以上である場合、前記取得された音を自動的に再生する、請求項2から4までの何れか一項に記載の出力装置。
  6.  前記制御部は、前記第1データに基づいて、前記取得部によって取得された音が前記記録対象であるかについての第1確率を算出し、
     算出した前記第1確率が確率閾値以上であると判定した場合、前記取得部によって取得された音が記録対象であると判定する、請求項1から5までの何れか一項に記載の出力装置。
  7.  前記第1データは、複数の音のカテゴリと、複数の第2確率と、複数の第1特徴量ベクトルと、複数の第2特徴量ベクトルとを含み、
     前記第2確率は、対応する前記音のカテゴリに前記取得部によって取得された音が属するかについての第2確率であり、
     前記第1特徴量ベクトルは、対応する前記音のカテゴリに属する音の特徴量ベクトルの平均値であり、
     前記第2特徴量ベクトルは、対応する前記音のカテゴリに属する音の特徴量ベクトルの共分散行列であり、
     前記制御部は、前記取得部によって取得された音の特徴量ベクトルと前記第1データとによって、前記第1確率を算出する、請求項6に記載の出力装置。
  8.  前記第1データは、機械学習モデルに用いられるニューラルネットワークのパラメータを含み、
     前記機械学習モデルは、前記取得部によって取得された音のデータが入力されると、前記第1確率を出力するように機械学習したものであり、
     前記制御部は、前記機械学習モデルによって前記第1確率を算出する、請求項6に記載の出力装置。
  9.  前記制御部は、前記音再生の入力を受け付けたタイミングから所定時間内の音のデータに基づいて、前記ニューラルネットワークのパラメータを更新する、請求項8に記載の出力装置。
  10.  前記制御部は、前記第1確率が第2再生閾値以上である場合、前記取得部によって取得された音を自動的に再生する、請求項6から9までの何れか一項に記載の出力装置。
  11.  前記第1データは、
      複数の記録対象の単語と、
      前記複数の記録対象の単語を含む音のそれぞれを録音する重要度を示す複数のスコアと、を含み、
     前記制御部は、前記取得部によって取得された音に含まれる単語のスコアが閾値以上である場合、前記取得された音が記録対象であると判定する、請求項1に記載の出力装置。
  12.  前記制御部は、
     前記音再生の入力を受け付けたタイミングから所定時間内の音のデータから単語のデータを取得し、
     取得した前記単語が前記記録対象の単語と一致する場合、前記記録対象の単語のスコアを増加させ、
     取得した前記単語が前記記録対象の単語と一致しない場合、取得した前記単語を新たな記録対象の単語として前記第1データに登録する、請求項11に記載の出力装置。
  13.  前記制御部は、前記スコアが第3再生閾値以上である場合、前記取得部によって取得された音を自動的に再生する、請求項11又は12に記載の出力装置。
  14.  前記出力装置の位置情報を取得可能な測位部をさらに備え、
     前記制御部は、前記測位部から取得した前記位置情報と前記第1データとに基づいて、前記取得部によって取得された音が記録対象であるか否かを判定する、請求項1に記載の出力装置。
  15.  音のデータを取得することと、
     記録対象となる音のデータを含む第1データに基づいて、前記取得された音が記録対象であるか否かを判定することと、
     音再生の入力を受け付けると、前記取得された音のうちの前記音再生の入力を受け付けたタイミングから所定時間内の音のデータに基づいて、前記第1データを更新することと、を含む、音出力方法。
PCT/JP2025/016660 2024-05-17 2025-05-02 出力装置及び音出力方法 Pending WO2025239246A1 (ja)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2024-080942 2024-05-17
JP2024080942 2024-05-17

Publications (1)

Publication Number Publication Date
WO2025239246A1 true WO2025239246A1 (ja) 2025-11-20

Family

ID=97720210

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2025/016660 Pending WO2025239246A1 (ja) 2024-05-17 2025-05-02 出力装置及び音出力方法

Country Status (1)

Country Link
WO (1) WO2025239246A1 (ja)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2005064744A (ja) * 2003-08-08 2005-03-10 Yamaha Corp 聴覚補助装置
JP2013114723A (ja) * 2011-11-30 2013-06-10 Sony Corp 情報処理装置、プログラムおよび情報処理方法
JP2014123887A (ja) * 2012-12-21 2014-07-03 Sharp Corp 音声再生装置
WO2020178961A1 (ja) * 2019-03-04 2020-09-10 マクセル株式会社 ヘッドマウント情報処理装置

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2005064744A (ja) * 2003-08-08 2005-03-10 Yamaha Corp 聴覚補助装置
JP2013114723A (ja) * 2011-11-30 2013-06-10 Sony Corp 情報処理装置、プログラムおよび情報処理方法
JP2014123887A (ja) * 2012-12-21 2014-07-03 Sharp Corp 音声再生装置
WO2020178961A1 (ja) * 2019-03-04 2020-09-10 マクセル株式会社 ヘッドマウント情報処理装置

Similar Documents

Publication Publication Date Title
US12567435B1 (en) Context driven device arbitration
US7603276B2 (en) Standard-model generation for speech recognition using a reference model
KR102450853B1 (ko) 음성 인식 장치 및 방법
JP6129134B2 (ja) 音声対話装置、音声対話システム、端末、音声対話方法およびコンピュータを音声対話装置として機能させるためのプログラム
CN112420026A (zh) 优化关键词检索系统
JP2016045487A (ja) 情報処理装置、情報処理システム、情報処理方法、及び情報処理プログラム
US20240134908A1 (en) Sound search
CN114067782B (zh) 音频识别方法及其装置、介质和芯片系统
EP3667660A1 (en) Information processing device and information processing method
US12190877B1 (en) Device arbitration for speech processing
CN112289300B (zh) 音频处理方法、装置及电子设备和计算机可读存储介质
JP2024504435A (ja) オーディオ信号生成システム及び方法
KR20230146898A (ko) 대화 처리 방법 및 대화 시스템
US20240005937A1 (en) Audio signal processing method and system for enhancing a bone-conducted audio signal using a machine learning model
CN112735381B (zh) 一种模型更新方法及装置
CN119943063B (zh) 基于不可学习语音样本的语音防克隆方法及装置
JPWO2011122522A1 (ja) 感性表現語選択システム、感性表現語選択方法及びプログラム
JP2005227794A (ja) 標準モデル作成装置及び標準モデル作成方法
JP6233103B2 (ja) 音声合成装置、音声合成方法及び音声合成プログラム
CN115862586A (zh) 音色特征提取模型的训练和音频合成的方法及装置
JP4877112B2 (ja) 音声処理装置およびプログラム
JP2019144524A (ja) ワード検出システム、ワード検出方法及びワード検出プログラム
JP2020197629A (ja) 音声テキスト変換システムおよび音声テキスト変換装置
WO2025229916A1 (ja) 音処理システム及び音処理方法
JP4297433B2 (ja) 音声合成方法及びその装置

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25803495

Country of ref document: EP

Kind code of ref document: A1