WO2025257952A1 - 情報処理装置及び方法 - Google Patents
情報処理装置及び方法Info
- Publication number
- WO2025257952A1 WO2025257952A1 PCT/JP2024/021262 JP2024021262W WO2025257952A1 WO 2025257952 A1 WO2025257952 A1 WO 2025257952A1 JP 2024021262 W JP2024021262 W JP 2024021262W WO 2025257952 A1 WO2025257952 A1 WO 2025257952A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- user
- utterance
- speech
- content data
- time
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
Definitions
- the present invention relates to technology for registering voices spoken by a user in association with that user.
- Patent Document 1 discloses a method for identifying which participant is speaking in a video conference call by comparing the direction of a sound source calculated from audio data with the coordinate information of each participant in an image.
- Patent Document 1 requires a device to photograph conference participants, making the system configuration complicated.
- the present invention aims to register the content of a user's speech in association with that user using a relatively simple configuration.
- the present invention provides an information processing device comprising: an acquisition unit that acquires speech content data indicating the content of speech made by users; an identification unit that identifies the time of speech made by each of the users and the user who made the speech; and a registration unit that compares the time of speech identified based on the speech content data with the time of speech identified for each user, and associates and registers the speech content data with the user who made the speech indicated by the speech content data.
- FIG. 1 is a block diagram showing an example of a configuration of an information processing system 1 according to an embodiment of the present invention.
- FIG. 2 is a block diagram showing an example of a hardware configuration of an information processing device 30 according to the embodiment.
- FIG. 2 is a block diagram showing an example of a functional configuration of an information processing device 30.
- 3 is a diagram illustrating an example of data stored in a storage unit 33 of an information processing device 30.
- FIG. 3 is a diagram illustrating an example of data stored in a storage unit 33 of an information processing device 30.
- FIG. 10 is a diagram illustrating an example of a correspondence relationship between speech content and speech timing.
- FIG. 10 is a diagram illustrating an example of a correspondence relationship between speech content and speech timing.
- 3 is a diagram illustrating an example of information stored in a storage unit 33 of an information processing device 30.
- FIG. 10 is a flowchart showing an example of an operation of the information processing device 30.
- FIG. 10 is a diagram illustrating offsets
- FIG. 1 is a diagram illustrating an example of the configuration of an information processing system 1 according to an embodiment of the present invention.
- the information processing system 1 includes multiple user terminals 10 used by multiple users to participate in a web conference, a web conference system 20 that provides web conference services to the multiple users, an information processing device 30 corresponding to the information processing device of the present invention, and a communication network 2, including the Internet, that communicatively connects these devices.
- the user terminals 10 are computers such as smartphones, wearable devices, tablets, or personal computers, and are equipped with at least functions for wired or wireless communication and audio input/output.
- the web conference system 20 is a system composed of one or more computers, enabling each user's audio or video to be shared among these users via the communication network 2, thereby realizing a web conference with users located at remote locations.
- the information processing device 30 is a computer that performs processing to record what each user said among multiple users participating in a web conference realized by the web conference system.
- the information processing device 30 may be composed of a single computer or multiple computers.
- Figure 2 is a diagram showing the hardware configuration of information processing device 30.
- Information processing device 30 is physically configured as a computer including a processor 3001, memory 3002, storage 3003, communication device 3004, input device 3005, output device 3006, and a bus connecting these. Each of these devices operates on power supplied from a battery (not shown).
- the term "device" can be interpreted as a circuit, device, unit, etc.
- the hardware configuration of information processing device 30 may be configured to include one or more of the devices shown in Figure 2, or may be configured without including some of the devices.
- information processing device 30 may be configured by multiple devices in different housings connected for communication.
- Each function of the information processing device 30 is realized by loading specific software (programs) onto hardware such as the processor 3001 and memory 3002, causing the processor 3001 to perform calculations, control communications via the communication device 3004, and control at least one of reading and writing data from and to the memory 3002 and storage 3003.
- the processor 3001 for example, runs an operating system to control the entire computer.
- the processor 3001 may be configured as a central processing unit (CPU) that includes an interface with peripheral devices, a control unit, an arithmetic unit, registers, etc.
- CPU central processing unit
- the processor 3001 reads programs (program code), software modules, data, etc. from at least one of the storage 3003 and the communication device 3004 into the memory 3002, and executes various processes in accordance with these.
- the programs used are those that cause a computer to execute at least some of the operations described below.
- the functional blocks of the information processing device 30 may be implemented by a control program stored in the memory 3002 and running on the processor 3001.
- the various processes may be executed by a single processor 3001, or may be executed simultaneously or sequentially by two or more processors 3001.
- the processor 3001 may be implemented on one or more chips.
- the programs may also be transmitted to the information processing device 30 via a telecommunications line.
- Memory 3002 is a computer-readable recording medium and may be composed of, for example, at least one of ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. Memory 3002 may also be called a register, cache, main memory (primary storage device), etc. Memory 3002 can store executable programs (program code), software modules, etc. for implementing the method of this embodiment.
- ROM Read Only Memory
- EPROM Erasable Programmable ROM
- EEPROM Electrical Erasable Programmable ROM
- RAM Random Access Memory
- Memory 3002 may also be called a register, cache, main memory (primary storage device), etc.
- Memory 3002 can store executable programs (program code), software modules, etc. for implementing the method of this embodiment.
- Storage 3003 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc.
- Storage 3003 may also be referred to as an auxiliary storage device.
- the communication device 3004 is hardware (transmission/reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as a network device, network controller, network card, communication module, etc.
- Each device such as the processor 3001 and memory 3002, is connected by a bus for communicating information.
- the bus may be configured using a single bus, or different buses may be used between each device.
- the information processing device 30 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a field-programmable gate array (FPGA), and some or all of the functional blocks may be realized by this hardware.
- the processor 3001 may be implemented using at least one of these pieces of hardware.
- the user terminal 10 is physically configured as a computer device including a processor, memory, storage, communication device, input device, output device, and buses connecting these.
- the processor, memory, and storage of the user terminal 10 are hardware similar to the processor 3001, memory 3002, and storage 3003 of the information processing device 30.
- the communication device of the user terminal 10 may be configured to include high-frequency switches, duplexers, filters, frequency synthesizers, etc. to realize at least one of frequency division duplex (FDD) and time division duplex (TDD).
- FDD frequency division duplex
- TDD time division duplex
- the transmitting and receiving antennas, amplifier units, transmitting and receiving units, and transmission path interfaces may be realized by communication devices.
- the transmitting and receiving units may be implemented as physically or logically separated transmitting and receiving units.
- the input device of the user terminal 10 is an input device that accepts input from outside (for example, a key, microphone, switch, button, various sensors, etc.).
- the output device of the user terminal 10 is an output device (e.g., a display, speaker, LED lamp, etc.) that outputs to the outside.
- the input device and output device may be integrated (e.g., a touch screen).
- a configuration for realizing audio input/output functions e.g., a microphone and speaker
- a configuration for realizing display functions e.g., a display
- Figure 3 is a block diagram showing the functional configuration of information processing device 30.
- processor 2001 reads programs and the like from storage 2003 to memory 2002 and executes them, thereby realizing the functions of speech content acquisition unit 31, speech timing acquisition unit 32, memory unit 33, voice analysis unit 34, identification unit 35, registration unit 36, and processing unit 37.
- the speech content acquisition unit 31 acquires speech content data indicating the content spoken by each user in a web conference from the web conference system 20, for example, after the end of each web conference, along with a conference ID (ID: Identification) for identifying that web conference, via the communication network 2.
- the speech content data is audio data indicating the time series of speech made in the web conference. Therefore, if multiple users speak simultaneously during a certain period of time, the speech content data will contain a mixture of speech from these multiple users for at least a portion of the time.
- the conference ID and speech content data are associated with each other and stored in the storage unit 23.
- the speech time acquisition unit 32 acquires speech time data indicating the time of speech by each user using each user terminal 10, a user ID for identifying that user, and the conference ID of that web conference via the communication network 2 from the user terminals 10 of users who participated in the web conference, i.e., user terminals 10 that were communicatively connected to the web conference system 20 during the web conference. More specifically, each user terminal 10 records the date and time when the volume of the voice picked up by the microphone during the web conference exceeds a threshold as the speech start date and time, and records the date and time when the volume of the voice picked up by the microphone during the web conference exceeds a threshold and then falls below that threshold as the speech end date and time.
- the speech start date and time and speech end date and time set are transmitted to the information processing device 30 via the communication network 2 together with the user ID of the user logged in to the web conference at that user terminal 10 and the conference ID of the web conference.
- the pair of speech start date and time and speech end date and time corresponds to speech time data indicating when the speech was made.
- the conference ID, user ID, and speech time data are associated with each other and stored in the storage unit 33.
- the speech analysis unit 34 uses voice activity detection (VAD) to distinguish between speech segments and non-speech segments in the speech content data stored in the memory unit 33, and identifies the start date and time and end date and time of each speech segment. It then performs speech recognition processing on the speech content data corresponding to the speech segments, and outputs the speech recognition results and the start date and time and end date and time of the speech (i.e., the time of the speech). This converts the speech content data in audio format into speech content data in text format with the time of the speech identified.
- VAD voice activity detection
- the registration unit 36 associates each piece of speech content data corresponding to that speech section with the user ID of that user and registers them in the memory unit 33.
- utterance content 3 will be registered in association with user ID "U001" as the speech content of user A with the longest speech time.
- the registration unit 36 associates the speech content data for that period with the user IDs of those multiple users and registers them.
- Fig. 9 is a flowchart showing an example of the operation of the information processing device 30.
- the utterance content acquisition unit 31 and the utterance time acquisition unit 32 each acquire each piece of data (step S11). That is, the utterance content acquisition unit 31 acquires utterance content data together with the conference ID of the Web conference from the Web conference system 20 via the communication network 2. Furthermore, the utterance time acquisition unit 32 acquires utterance time data, a user ID, and a conference ID from each user terminal 10 via the communication network 2. The acquired data are stored in the storage unit 33.
- the registration unit 36 compares the time of the utterance identified by the voice analysis unit 34 with the time of the utterance identified for each user by the identification unit 35, and associates each piece of utterance content data with the user ID of the user who spoke the content indicated by the utterance content data, and registers this in the storage unit 33 (step S14).
- the processing unit 37 performs a predetermined process using the content registered by the registration unit 36 (step S15).
- the identification unit 35 identified the time when each user spoke based on the volume of the speech of that user.
- the identification unit 35 may identify the start date and time and end date and time of each speech section based on the results of speech section detection performed by the speech analysis unit 34 on the user's speech, and identify the period between the start date and time and the end date and time as the time when the user spoke.
- the registration unit 36 compares the time of an utterance identified by the voice analysis unit 34 based on the utterance content data with the time of an utterance identified by the identification unit 35 from the utterance time data for each user, and registers each piece of utterance content data in association with the user who uttered the content indicated by the utterance content data.
- the time axis used by the voice analysis unit 34 to identify the time of an utterance based on the utterance content data is misaligned with the time axis used by the identification unit 35 to identify the time of an utterance for each user.
- the time axis used by the voice analysis unit 34 to identify the time of an utterance and the time axis used by the identification unit 35 to identify the time of an utterance are both based on system time measured by the system, but it is possible that these system times do not completely match.
- the registration unit 36 therefore extracts speech content data from only one user during one web conference, and determines as an offset value the difference between the speech time identified based on the speech content data (first speech time) and the speech time closest to the first speech time (referred to as the second speech time) among the speech times identified for each user by the identification unit 35. Then, when comparing the speech time identified by the voice analysis unit 34 with the speech time identified by the identification unit 35, the registration unit 36 takes this offset value into consideration when making the comparison.
- the registration unit 36 associates and registers the utterance content data with the user based on the difference between the time of the utterance by one user included in the utterance content data and the time of the utterance identified by the identification unit 35. This makes it possible to perform registration without being affected by the difference, even if there is a discrepancy between the time axis used by the voice analysis unit 34 to identify the time of the utterance based on the utterance content data and the time axis used by the identification unit 35 to identify each user from the utterance time data.
- the processing of the processing unit 37 is not limited to the example of the embodiment and may be such.
- the processing unit 37 may create a prompt to be input to an algorithm that responds in an interactive format based on a speech record in which the results of speech recognition of speech content data are recorded for each user.
- the algorithm that responds in an interactive format is, for example, a generation AI (artificial intelligence) that generates an outline of meeting minutes
- the prompt for the generation AI includes, for example, instructions, context, input, output, etc.
- the processing unit 37 generates a prompt with instructions to generate an outline of meeting minutes based on a speech record in which the results of speech recognition of speech content data are recorded for each user, and further specifies the context, input, and output.
- the voice analysis is performed by the voice analysis unit 34 of the information processing device 30, but this may also be performed by the web conference system 20, and the information processing device 30 may acquire the voice analysis results.
- the present invention is an information processing device characterized by comprising an acquisition unit that acquires utterance content data indicating the content uttered by users, an identification unit that identifies the time of utterance by each of the users and the user who spoke, and a registration unit that compares the time of utterance identified based on the utterance content data with the time of utterance identified for each user, and associates the utterance content data with the user who spoke the content indicated by the utterance content data and registers them.
- the utterance content data acquired by the acquisition unit may be voice data indicating the content uttered by the users, or may be text data indicating the content uttered by the users.
- the identification unit 35 identifies the time of speech by each user based on the volume of the speech of that user, but the method of identifying the speech time is not limited to this, and any well-known method may be used.
- the identification unit 35 may use a method called voice activity detection (VAD) to distinguish between speech intervals and non-speech intervals, and identify the start and end times of each speech interval as the speech time.
- VAD voice activity detection
- the present invention is applicable to cloud-based web conferencing systems and on-premise web conferencing systems. Furthermore, the present invention is not limited to web conferencing systems, but can be applied to any conferencing system that shares speech content from multiple users, regardless of specifications or protocols.
- each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are directly or indirectly connected (e.g., wired, wireless, etc.) and these multiple devices.
- the functional block may also be realized by combining software with the single device or multiple devices.
- Functions include, but are not limited to, judgment, determination, assessment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, election, establishment, comparison, assumption, expectation, regard, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment.
- a functional block (component) that performs transmission functions is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these are implemented.
- the information processing device 30 in one embodiment of the present disclosure may function as a computer that performs the processing of the present disclosure.
- LTE Long Term Evolution
- LTE-A Long Term Evolution-Advanced
- SUPER 3G IMT-Advanced
- 4G 4th generation mobile communication system
- 5G 5th generation mobile communication system
- FRA Full Radio Access
- NR new Ra
- the present invention may be applied to at least one of systems using IEEE 802.11 (Wi-Fi (registered trademark)), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, UWB (ULtra-WIDE Band), Bluetooth (registered trademark), or other appropriate systems, as well as next-generation systems that are based on and extend these. It may also be applied to a combination of multiple systems (for example, a combination of at least one of LTE and LTE-A with 5G).
- the present invention may also be an information processing method characterized by comprising an acquisition step of acquiring speech content data indicating the content of speech made by users; an identification step of identifying the time of speech made by each of the users and the user who made the speech; and a registration step of comparing the time of speech identified based on the speech content data with the time of speech identified for each user, and registering the speech content data in association with the user who made the speech indicated by the speech content data.
- the processing procedures, sequences, flowcharts, etc. of each aspect/embodiment described in this disclosure may be rearranged as long as there is no contradiction.
- the methods described in this disclosure present various step elements in an exemplary order, and are not limited to the specific order presented.
- Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.
- the determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (for example, comparison with a predetermined value).
- Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
- Software, instructions, information, etc. may also be transmitted or received via a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, or Digital Subscriber Line (DSL)) and/or wireless technologies (such as infrared or microwave), then such wired and/or wireless technologies are included within the definition of a transmission medium.
- wired technologies such as coaxial cable, fiber optic cable, twisted pair, or Digital Subscriber Line (DSL)
- wireless technologies such as infrared or microwave
- information, parameters, etc. described in this disclosure may be expressed using absolute values, relative values from a predetermined value, or other corresponding information.
- the phrase “based on” does not mean “based only on,” unless expressly stated otherwise. In other words, the phrase “based on” means both “based only on” and “based at least on.”
- any reference to an element using a designation such as "first,” “second,” etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.
- a and B are different may mean “A and B are different from each other.” Note that this term may also mean “A and B are each different from C.” Terms such as “separate” and “combined” may also be interpreted in the same way as “different.”
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Telephonic Communication Services (AREA)
Abstract
発話内容取得部31は、Web会議システム20から、Web会議において各々のユーザによって発話された内容を示す発話内容データをそのWeb会議を識別するための会議IDとともに通信網2経由で取得する。発話時期取得部32は、Web会議に参加したユーザのユーザ端末10から、各ユーザ端末10を利用する各ユーザによる発話の時期を示す発話時期データと、そのユーザを識別するためのユーザIDと、そのWeb会議の会議IDとを通信網2経由で取得する。音声分析部34は、発話内容データにおいて音声区間検出により音声区間と非音声区間とを判別して各音声区間の開始日時及び終了日時を特定し、さらに、音声区間に相当する発話内容データに対して音声認識処理を実行して、その音声認識結果とその音声の開始日時及び終了日時(つまり発話の時期)を出力する。特定部35は、各々のユーザによる発話の時期と、そのユーザのユーザIDとを特定する。登録部36は、記憶部33に記憶されている発話内容データに基づいて音声分析部34により特定される発話の時期と、特定部35により各々のユーザについて特定された発話の時期とを比較して、各発話内容データと、その発話内容データが示す内容を発話したユーザとを対応付けて記憶部33に登録する。
Description
本発明は、ユーザが発話した音声とそのユーザとを対応付けて登録するための技術に関する。
近年、Web会議システムが広範に利用されている。この種のシステムにおいて、複数の会議参加者のうち誰がどのような発言をしたかを記録することが求められている。例えば特許文献1には、テレビ電話会議において、音声データから計算した音源の方向と、画像内における各参加者の座標情報とを比較することによって、どの参加者が話者であるかを特定することが開示されている。
特許文献1に記載の仕組みでは、会議参加者を撮影する装置が必要になるなど、システム構成が複雑となる。
そこで、本発明は、比較的簡易な構成により、ユーザが発話した内容とそのユーザとを対応付けて登録することを目的とする。
上記課題を解決するため、本発明は、ユーザによって発話された内容を示す発話内容データを取得する取得部と、各々の前記ユーザによる発話の時期と、発話した当該ユーザとを特定する特定部と、前記発話内容データに基づいて特定される発話の時期と各々のユーザについて特定された前記発話の時期とを比較して、前記発話内容データと、当該発話内容データが示す内容を発話したユーザとを対応付けて登録する登録部とを備えることを特徴とする情報処理装置を提供する。
本発明によれば、比較的簡易な構成により、ユーザが発話した内容とそのユーザとを対応付けて登録することができる。
[実施形態]
[構成]
図1は、本発明の実施形態に係る情報処理システム1の構成の一例を示す図である。情報処理システム1は、複数のユーザがそれぞれWeb会議に参加するために利用する複数のユーザ端末10と、複数のユーザにWeb会議サービスを提供するWeb会議システム20と、本発明の情報処理装置に相当する情報処理装置30と、これらを通信可能に接続する、インターネットを含む通信網2とを備えている。ユーザ端末10は、例えばスマートホン、ウェアラブル端末、タブレット又はパーソナルコンピュータなどのコンピュータであり、有線又は無線によって通信を行う機能と音声の入出力を行う機能とを少なくとも備えている。Web会議システム20は、1以上のコンピュータによって構成されたシステムであり、各ユーザの音声又は映像を通信網2経由でこれらユーザ間において共有可能として、互いに遠隔拠点にいるユーザとのWeb会議を実現する。情報処理装置30は、Web会議システムによって実現されるWeb会議に参加している複数のユーザのうちどのユーザがどのような発言をしたかを記録するための処理を行うコンピュータである。この情報処理装置30は、単体のコンピュータで構成されていてもよいし、複数のコンピュータによって構成されていてもよい。
[構成]
図1は、本発明の実施形態に係る情報処理システム1の構成の一例を示す図である。情報処理システム1は、複数のユーザがそれぞれWeb会議に参加するために利用する複数のユーザ端末10と、複数のユーザにWeb会議サービスを提供するWeb会議システム20と、本発明の情報処理装置に相当する情報処理装置30と、これらを通信可能に接続する、インターネットを含む通信網2とを備えている。ユーザ端末10は、例えばスマートホン、ウェアラブル端末、タブレット又はパーソナルコンピュータなどのコンピュータであり、有線又は無線によって通信を行う機能と音声の入出力を行う機能とを少なくとも備えている。Web会議システム20は、1以上のコンピュータによって構成されたシステムであり、各ユーザの音声又は映像を通信網2経由でこれらユーザ間において共有可能として、互いに遠隔拠点にいるユーザとのWeb会議を実現する。情報処理装置30は、Web会議システムによって実現されるWeb会議に参加している複数のユーザのうちどのユーザがどのような発言をしたかを記録するための処理を行うコンピュータである。この情報処理装置30は、単体のコンピュータで構成されていてもよいし、複数のコンピュータによって構成されていてもよい。
図2は、情報処理装置30のハードウェア構成を示す図である。情報処理装置30は、物理的には、プロセッサ3001、メモリ3002、ストレージ3003、通信装置3004、入力装置3005、出力装置3006、及びこれらを接続するバスなどを含むコンピュータとして構成されている。これらの各装置は図示せぬ電池から供給される電力によって動作する。なお、以下の説明では、「装置」という文言は、回路、デバイス、ユニットなどに読み替えることができる。情報処理装置30のハードウェア構成は、図2に示した各装置を1つ又は複数含むように構成されてもよいし、一部の装置を含まずに構成されてもよい。また、それぞれ筐体が異なる複数の装置が通信接続されて、情報処理装置30を構成してもよい。
情報処理装置30における各機能は、プロセッサ3001、メモリ3002などのハードウェア上に所定のソフトウェア(プログラム)を読み込ませることによって、プロセッサ3001が演算を行い、通信装置3004による通信を制御したり、メモリ3002及びストレージ3003におけるデータの読み出し及び書き込みの少なくとも一方を制御したりすることによって実現される。
プロセッサ3001は、例えば、オペレーティングシステムを動作させてコンピュータ全体を制御する。プロセッサ3001は、周辺装置とのインターフェース、制御装置、演算装置、レジスタなどを含む中央処理装置(CPU:Central Processing Unit)によって構成されてもよい。
プロセッサ3001は、プログラム(プログラムコード)、ソフトウェアモジュール、データなどを、ストレージ3003及び通信装置3004の少なくとも一方からメモリ3002に読み出し、これらに従って各種の処理を実行する。プログラムとしては、後述する動作の少なくとも一部をコンピュータに実行させるプログラムが用いられる。情報処理装置30の機能ブロックは、メモリ3002に格納され、プロセッサ3001において動作する制御プログラムによって実現されてもよい。各種の処理は、1つのプロセッサ3001によって実行されてもよいが、2以上のプロセッサ3001により同時又は逐次に実行されてもよい。プロセッサ3001は、1以上のチップによって実装されてもよい。なお、プログラムは、電気通信回線を介して情報処理装置30に送信されてもよい。
メモリ3002は、コンピュータ読み取り可能な記録媒体であり、例えば、ROM(Read Only Memory)、EPROM(Erasable Programmable ROM)、EEPROM(Electrically Erasable Programmable ROM)、RAM(Random Access Memory)などの少なくとも1つによって構成されてもよい。メモリ3002は、レジスタ、キャッシュ、メインメモリ(主記憶装置)などと呼ばれてもよい。メモリ3002は、本実施形態に係る方法を実施するために実行可能なプログラム(プログラムコード)、ソフトウェアモジュールなどを保存することができる。
ストレージ3003は、コンピュータ読み取り可能な記録媒体であり、例えば、CD-ROM(Compact Disc ROM)などの光ディスク、ハードディスクドライブ、フレキシブルディスク、光磁気ディスク(例えば、コンパクトディスク、デジタル多用途ディスク、Blu-ray(登録商標)ディスク)、スマートカード、フラッシュメモリ(例えば、カード、スティック、キードライブ)、フロッピー(登録商標)ディスク、磁気ストリップなどの少なくとも1つによって構成されてもよい。ストレージ3003は、補助記憶装置と呼ばれてもよい。
通信装置3004は、有線ネットワーク及び無線ネットワークの少なくとも一方を介してコンピュータ間の通信を行うためのハードウェア(送受信デバイス)であり、例えばネットワークデバイス、ネットワークコントローラ、ネットワークカード、通信モジュールなどともいう。
プロセッサ3001、メモリ3002などの各装置は、情報を通信するためのバスによって接続される。バスは、単一のバスを用いて構成されてもよいし、装置間ごとに異なるバスを用いて構成されてもよい。
情報処理装置30は、マイクロプロセッサ、デジタル信号プロセッサ(DSP:Digital Signal Processor)、ASIC(Application Specific Integrated Circuit)、PLD(Programmable Logic Device)、FPGA(Field Programmable Gate Array)などのハードウェアを含んで構成されてもよく、そのハードウェアにより、各機能ブロックの一部又は全てが実現されてもよい。例えば、プロセッサ3001は、これらのハードウェアの少なくとも1つを用いて実装されてもよい。
なお、ユーザ端末10は、物理的には、プロセッサ、メモリ、ストレージ、通信装置、入力装置、出力装置及びこれらを接続するバスなどを含むコンピュータ装置として構成されている。ユーザ端末10のプロセッサ、メモリ、ストレージは、情報処理装置30のプロセッサ3001、メモリ3002、ストレージ3003と同様のハードウェアである。ユーザ端末10の通信装置は、例えば周波数分割複信(FDD:Frequency Division Duplex)及び時分割複信(TDD:Time Division Duplex)の少なくとも一方を実現するために、高周波スイッチ、デュプレクサ、フィルタ、周波数シンセサイザなどを含んで構成されてもよい。例えば、送受信アンテナ、アンプ部、送受信部、伝送路インターフェースなどは、通信装置によって実現されてもよい。送受信部は、送信部と受信部とで、物理的に、または論理的に分離された実装がなされてもよい。ユーザ端末10の入力装置は、外部からの入力を受け付ける入力デバイス(例えば、キー、マイクロホン、スイッチ、ボタン、各センサ等)である。ユーザ端末10の出力装置は、外部への出力を実施する出力デバイス(例えば、ディスプレイ、スピーカー、LEDランプなど)である。なお、入力装置及び出力装置は、一体となった構成(例えば、タッチスクリーン)であってもよい。入力装置及び出力装置において、音声入出力機能を実現するための構成(例えばマイクロホン及びスピーカー)は必須であり、表示機能を実現するための構成(例えばディスプレイ)は必ずしも必須ではない。
図3は、情報処理装置30の機能構成を示すブロック図である。情報処理装置30において、プロセッサ2001がプログラムなどをストレージ2003からメモリ2002に読み出して実行することで、発話内容取得部31と、発話時期取得部32と、記憶部33と、音声分析部34と、特定部35と、登録部36と、処理部37という機能を実現する。
発話内容取得部31は、Web会議システム20から、例えば各々のWeb会議の終了後に、そのWeb会議において各々のユーザによって発話された内容を示す発話内容データをそのWeb会議を識別するための会議ID(ID:Identification、識別子)とともに通信網2経由で取得する。本実施形態における発話内容データは、Web会議において発話がなされたときの時系列の音声を示す音声データである。従って、或る時間帯で複数のユーザが同時に発話した場合には、その発話内容データにおいて、これら複数のユーザによって発話された音声が少なくとも一部の期間で混在している。図4に例示するように、これら会議ID及び発話内容データは互いに対応付けられて記憶部23に記憶される。
図3において、発話時期取得部32は、Web会議に参加したユーザのユーザ端末10、つまりWeb会議の期間中にWeb会議システム20に通信接続したユーザ端末10から、各ユーザ端末10を利用する各ユーザによる発話の時期を示す発話時期データと、そのユーザを識別するためのユーザIDと、そのWeb会議の会議IDとを通信網2経由で取得する。より具体的には、各ユーザ端末10は、Web会議の期間中において、マイクロホンにより収音した音声の音量が閾値を超える場合にそのときの日時を発話開始日時として記録し、Web会議の期間中においてマイクロホンにより収音した音声の音量が閾値を超えてからその閾値を下回った場合にそのときの日時を発話終了日時として記録しておき、例えばWeb会議終了後にこれら発話開始日時及び発話終了日時の組を、そのユーザ端末10においてWeb会議にログインしているユーザのユーザID及びWeb会議の会議IDとともに通信網2経由で情報処理装置30に送信する。これら発話開始日時及び発話終了日時の組は、発話した時期を示す発話時期データに相当する。図5に例示するように、これら会議ID、ユーザID及び発話時期データは互いに対応付けられて記憶部33に記憶される。
図3において、音声分析部34は、記憶部33に記憶されている発話内容データにおいて、音声区間検出(VAD:Voice Activity Detection)により音声区間と非音声区間とを判別して各音声区間の開始日時及び終了日時を特定し、さらに、音声区間に相当する発話内容データに対して音声認識処理を実行して、その音声認識結果とその音声の開始日時及び終了日時(つまり発話の時期)を出力する。これにより、音声形式の発話内容データが、発話時期が特定されたテキスト形式の発話内容データに変換される。
特定部35は、各々のユーザによる発話の時期と、そのユーザのユーザIDとを特定する。具体的には、特定部35は、前述した各ユーザの発話時期データ、つまり、各ユーザによる発話音声の音量に基づいて、そのユーザによる発話の時期を特定する。
登録部36は、記憶部33に記憶されている発話内容データに基づいて音声分析部34により特定される発話の時期と、特定部35により各々のユーザについて発話時期データから特定された発話の時期とを比較して、各発話内容データと、その発話内容データが示す内容を発話したユーザとを対応付けて記憶部33に登録する。より具体的には、登録部36は、或る音声区間について音声分析部34により特定された発話開始日時と、或るユーザについて特定部35により特定された発話開始日時との差が閾値以内で、且つ、その音声区間について音声分析部34により特定された発話終了日時と、そのユーザについて特定部35により特定された発話終了日時との差が閾値以内である場合に、その音声区間に対応する各発話内容データと、そのユーザのユーザIDとを対応付けて記憶部33に登録する。
例えば図6に例示するように、或る発話内容1の音声区間の発話時期が「00:10」~「00:45」である場合に、ユーザA(図5のユーザID「U001」のユーザとする)の発話時期が「00:10」~「00:45」であった場合、発話内容1はユーザAの発話内容としてユーザID「U001」に対応付けて登録されることになる。また、別の発話内容2の音声区間の発話時期が「00:47」~「01:23」である場合に、ユーザB(図5のユーザID「U002」のユーザとする)の発話時期が「00:47」~「01:23」であった場合、発話内容2はユーザBの発話内容としてユーザID「U002」に対応付けて登録されることになる。
ここで、前述したように、或る時間帯で複数のユーザが同時に発話した場合には、その発話内容データにおいて、これら複数のユーザによって発話された音声が少なくとも一部の期間で混在している。この場合、登録部36は、その発話音声が混在している期間におけるその発話内容データと、その期間において最も長い発話時間のユーザのユーザIDとを対応付けて登録する。
例えば図7に例示するように、或る発話内容3の音声区間の発話時期が「03:10」~「04:35」である場合に、ユーザA(図5のユーザID「U001」のユーザ)の発話時期が「03:10」~「03:50」であり、ユーザB(図5のユーザID「U002」のユーザ)の発話時期が「03:40」~「04:10」であり、ユーザC(図5のユーザID「U003」のユーザ)の発話時期が「03:57」~「04:35」であった場合、発話内容3は最も長い発話時間のユーザAの発話内容としてユーザID「U001」に対応付けて登録されることになる。なお、発話内容データにおいて複数のユーザによる音声が混在している期間において最も長い発話時間のユーザが複数いる場合には、登録部36は、その期間におけるその発話内容データと、その複数のユーザのユーザIDとを対応付けて登録する。
以上のような登録部36の登録により、図8に例示するように、これら会議ID、ユーザID、発話時期データ及び発話内容データは互いに対応付けられて記憶部33に記憶される。
図3において、処理部37は、登録部36により登録された内容を用いて所定の処理を行う。例えば処理部37は、音声認識した結果であるテキスト形式の発話内容データをユーザIDごとに記録した発話記録を例えば議事録として作成して出力する。ここでいう出力とは、送信、表示、記憶媒体への書き込み、プリントアウト等のあらゆる出力形態を含む。
[動作]
次に、本実施形態の動作を説明する。図9は、情報処理装置30の動作の一例を示すフローチャートである。図9において、発話内容取得部31及び発話時期取得部32はそれぞれ各データを取得する(ステップS11)。つまり、発話内容取得部31は、Web会議システム20から発話内容データをWeb会議の会議IDとともに通信網2経由で取得する。また、発話時期取得部32は、各ユーザ端末10から、発話時期データと、ユーザIDと、会議IDとを通信網2経由で取得する。これら取得された各データは記憶部33に記憶される。
次に、本実施形態の動作を説明する。図9は、情報処理装置30の動作の一例を示すフローチャートである。図9において、発話内容取得部31及び発話時期取得部32はそれぞれ各データを取得する(ステップS11)。つまり、発話内容取得部31は、Web会議システム20から発話内容データをWeb会議の会議IDとともに通信網2経由で取得する。また、発話時期取得部32は、各ユーザ端末10から、発話時期データと、ユーザIDと、会議IDとを通信網2経由で取得する。これら取得された各データは記憶部33に記憶される。
次に、音声分析部34は、記憶部33に記憶されている発話内容データにおいて、音声区間検出により音声区間と非音声区間とを判別して各音声区間の開始日時及び終了日時を特定し、さらに、音声区間に相当する発話内容データに対して音声認識処理を実行して、その音声認識結果とその音声の開始日時及び終了日時を特定する(ステップS12)。
次に、特定部35は、各ユーザによる発話音声の音量に基づいて、そのユーザによる発話の時期を特定する(ステップS13)。
次に、登録部36は、音声分析部34により特定された発話の時期と、特定部35により各々のユーザについて特定された発話の時期とを比較して、各発話内容データと、その発話内容データが示す内容を発話したユーザのユーザIDとを対応付けて記憶部33に登録する(ステップS14)。
処理部37は、登録部36により登録された内容を用いて所定の処理を行う(ステップS15)。
本実施形態によれば、会議参加者を撮影する装置を必要とせず、比較的簡易な構成により、ユーザが発話した内容とそのユーザとを対応付けて登録することが可能となる。
[変形例]
本発明は、上述した実施形態に限定されない。上述した実施形態を以下のように変形してもよい。また、以下の2つ以上の変形例を組み合わせて実施してもよい。
本発明は、上述した実施形態に限定されない。上述した実施形態を以下のように変形してもよい。また、以下の2つ以上の変形例を組み合わせて実施してもよい。
[変形例1]
上記実施形態において、特定部35は、各ユーザによる発話音声の音量に基づいて、そのユーザによる発話の時期を特定していたが、これに代えて、ユーザによる発話音声に対して音声分析部34が音声区間検出を行った結果に基づいて、各音声区間の開始日時及び終了日時を特定し、その開始日時及び終了日時の期間をユーザによる発話の時期として特定するようにしてもよい。
上記実施形態において、特定部35は、各ユーザによる発話音声の音量に基づいて、そのユーザによる発話の時期を特定していたが、これに代えて、ユーザによる発話音声に対して音声分析部34が音声区間検出を行った結果に基づいて、各音声区間の開始日時及び終了日時を特定し、その開始日時及び終了日時の期間をユーザによる発話の時期として特定するようにしてもよい。
[変形例2]
上記実施形態において、登録部36は、音声分析部34により発話内容データに基づいて特定される発話の時期と、特定部35により各々のユーザについて発話時期データから特定された発話の時期とを比較して、各発話内容データと、その発話内容データが示す内容を発話したユーザとを対応付けて登録していた。ここで、音声分析部34により発話内容データに基づいて特定される発話の時期を特定するときの時間軸と、特定部35により各々のユーザについて発話時期データから特定するときの時間軸とがずれる場合も考えられる。つまり、音声分析部34が発話の時期を特定するときの時間軸と、特定部35が発話の時期を特定するときの時間軸は、それぞれシステムによって計測されるシステム時間に基づいているが、このシステム時間は双方で完全一致しない場合が考えられる。
上記実施形態において、登録部36は、音声分析部34により発話内容データに基づいて特定される発話の時期と、特定部35により各々のユーザについて発話時期データから特定された発話の時期とを比較して、各発話内容データと、その発話内容データが示す内容を発話したユーザとを対応付けて登録していた。ここで、音声分析部34により発話内容データに基づいて特定される発話の時期を特定するときの時間軸と、特定部35により各々のユーザについて発話時期データから特定するときの時間軸とがずれる場合も考えられる。つまり、音声分析部34が発話の時期を特定するときの時間軸と、特定部35が発話の時期を特定するときの時間軸は、それぞれシステムによって計測されるシステム時間に基づいているが、このシステム時間は双方で完全一致しない場合が考えられる。
そこで、登録部36は、1つのWeb会議の期間中に1人のユーザのみの発話内容データを抽出し、その発話内容データに基づいて特定される発話の時期(第1の発話時期)と、特定部35により各々のユーザについて特定された発話の時期のうち、第1の発話時期に最も近い発話の時期(第2の発話時期という)との差分をオフセット値とする。そして、登録部36は、音声分析部34により特定される発話の時期と、特定部35により特定される発話の時期とを比較するときに、このオフセット値を考慮してその比較を行う。
例えば図10に例示するように、或る発話内容1の音声区間の発話時期(第1の発話時期)が「00:11」~「00:46」である場合に、ユーザA(図5のユーザID「U001」のユーザ)の発話時期(第2の発話時期)が「00:10」~「00:45」であり、第1の発話時期に最も近い発話時期が第2の発話時期であった場合、これら第1の発話時期及び第2の発話時期の差分1秒がオフセット値Tとなる。このような場合、登録部36は、音声分析部34により特定される発話の時期よりもオフセット値T=1秒だけ前の時期を実際の発話の時期として、特定部35により特定される発話の時期との比較を行う。
このように、登録部36は、発話内容データに含まれる1人のユーザによる発話の時期と、特定部35により特定された発話の時期との差分に基づいて、発話内容データとユーザとを対応付けて登録する。これにより、音声分析部34により発話内容データに基づいて特定される発話の時期を特定するときの時間軸と、特定部35により各々のユーザについて発話時期データから特定するときの時間軸とがずれている場合であっても、そのずれの影響を受けずに登録を行うことが可能となる。
[変形例3]
処理部37の処理は実施形態の例に限定されず、そのようなものであってもよい。例えば処理部37は、発話内容データを音声認識した結果をユーザごとに記録した発話記録に基づいて、対話形式で応答を行うアルゴリズムに入力するプロンプトを作成するようにしてもよい。対話形式で応答を行うアルゴリズムとは、例えば会議の議事録の概要を生成する生成AI(Artificial Intelligence)等であり、生成AIに対するプロンプトは例えばインストイラクション、コンテキスト、インプット、アウトプット等からなる。処理部37は、発話内容データを音声認識した結果をユーザごとに記録した発話記録に基づいて会議の議事録の概要を生成することをインストラクションとし、さらに、コンテキスト、インプット、アウトプットを指定したプロンプトを作成する。
処理部37の処理は実施形態の例に限定されず、そのようなものであってもよい。例えば処理部37は、発話内容データを音声認識した結果をユーザごとに記録した発話記録に基づいて、対話形式で応答を行うアルゴリズムに入力するプロンプトを作成するようにしてもよい。対話形式で応答を行うアルゴリズムとは、例えば会議の議事録の概要を生成する生成AI(Artificial Intelligence)等であり、生成AIに対するプロンプトは例えばインストイラクション、コンテキスト、インプット、アウトプット等からなる。処理部37は、発話内容データを音声認識した結果をユーザごとに記録した発話記録に基づいて会議の議事録の概要を生成することをインストラクションとし、さらに、コンテキスト、インプット、アウトプットを指定したプロンプトを作成する。
[変形例4]
上記実施形態では音声分析を情報処理装置30の音声分析部34が行っていたが、これをWeb会議システム20が行って、情報処理装置30はその音声分析結果を取得するようにしてもよい。つまり、本発明は、ユーザによって発話された内容を示す発話内容データを取得する取得部と、各々の前記ユーザによる発話の時期と、発話した当該ユーザとを特定する特定部と、前記発話内容データに基づいて特定される発話の時期と各々のユーザについて特定された前記発話の時期とを比較して、前記発話内容データと、当該発話内容データが示す内容を発話したユーザとを対応付けて登録する登録部とを備えることを特徴とする情報処理装置であるが、取得部によって取得される発話内容データは、ユーザによって発話された内容を示す音声データであってもよいし、ユーザによって発話された内容を示すテキストデータであってもよい。
上記実施形態では音声分析を情報処理装置30の音声分析部34が行っていたが、これをWeb会議システム20が行って、情報処理装置30はその音声分析結果を取得するようにしてもよい。つまり、本発明は、ユーザによって発話された内容を示す発話内容データを取得する取得部と、各々の前記ユーザによる発話の時期と、発話した当該ユーザとを特定する特定部と、前記発話内容データに基づいて特定される発話の時期と各々のユーザについて特定された前記発話の時期とを比較して、前記発話内容データと、当該発話内容データが示す内容を発話したユーザとを対応付けて登録する登録部とを備えることを特徴とする情報処理装置であるが、取得部によって取得される発話内容データは、ユーザによって発話された内容を示す音声データであってもよいし、ユーザによって発話された内容を示すテキストデータであってもよい。
[変形例5]
上記実施形態において、特定部35は各ユーザによる発話音声の音量に基づいてそのユーザによる発話の時期を特定していたが、発話時期の特定方法はこれに限らず、周知の手法を用いてもよい。例えば、特定部35は、音声区間検出(VAD:Voice Activity Detection)と呼ばれる手法を用いて音声区間と非音声区間とを識別し、各音声区間が開始される時期及び終了する時期を発話時期として特定するようにしてもよい。
上記実施形態において、特定部35は各ユーザによる発話音声の音量に基づいてそのユーザによる発話の時期を特定していたが、発話時期の特定方法はこれに限らず、周知の手法を用いてもよい。例えば、特定部35は、音声区間検出(VAD:Voice Activity Detection)と呼ばれる手法を用いて音声区間と非音声区間とを識別し、各音声区間が開始される時期及び終了する時期を発話時期として特定するようにしてもよい。
[変形例6]
本発明は、クラウド型のWeb会議システムやオンプレミス型のWeb会議システムに対して適用可能である。さらに、本発明は、Web会議システムに対してのみ適用されるものではなく、複数のユーザによって発話された内容を共有する会議システムであれば、どのような仕様やプロトロルのものであっても適用可能である。
本発明は、クラウド型のWeb会議システムやオンプレミス型のWeb会議システムに対して適用可能である。さらに、本発明は、Web会議システムに対してのみ適用されるものではなく、複数のユーザによって発話された内容を共有する会議システムであれば、どのような仕様やプロトロルのものであっても適用可能である。
[その他の変形例]
なお、上記実施形態の説明に用いたブロック図は、機能単位のブロックを示している。これらの機能ブロック(構成部)は、ハードウェア及びソフトウェアの少なくとも一方の任意の組み合わせによって実現される。また、各機能ブロックの実現方法は特に限定されない。すなわち、各機能ブロックは、物理的又は論理的に結合した1つの装置を用いて実現されてもよいし、物理的又は論理的に分離した2つ以上の装置を直接的又は間接的に(例えば、有線、無線などを用いて)接続し、これら複数の装置を用いて実現されてもよい。機能ブロックは、上記1つの装置又は上記複数の装置にソフトウェアを組み合わせて実現されてもよい。
なお、上記実施形態の説明に用いたブロック図は、機能単位のブロックを示している。これらの機能ブロック(構成部)は、ハードウェア及びソフトウェアの少なくとも一方の任意の組み合わせによって実現される。また、各機能ブロックの実現方法は特に限定されない。すなわち、各機能ブロックは、物理的又は論理的に結合した1つの装置を用いて実現されてもよいし、物理的又は論理的に分離した2つ以上の装置を直接的又は間接的に(例えば、有線、無線などを用いて)接続し、これら複数の装置を用いて実現されてもよい。機能ブロックは、上記1つの装置又は上記複数の装置にソフトウェアを組み合わせて実現されてもよい。
機能には、判断、決定、判定、計算、算出、処理、導出、調査、探索、確認、受信、送信、出力、アクセス、解決、選択、選定、確立、比較、想定、期待、見做し、報知(broadcasting)、通知(notifying)、通信(communicating)、転送(forwarding)、構成(configuring)、再構成(reconfiguring)、割り当て(allocating、mapping)、割り振り(assigning)などがあるが、これらに限られない。たとえば、送信を機能させる機能ブロック(構成部)は、送信制御部(transmitting unit)や送信機(transmitter)と呼称される。いずれも、上述したとおり、実現方法は特に限定されない。
例えば、本開示の一実施の形態における情報処理装置30などは、本開示の処理を行うコンピュータとして機能してもよい。
本開示において説明した各態様/実施形態は、LTE(Long Term Evolution)、LTE-A(LTE-Advanced)、SUPER 3G、IMT-Advanced、4G(4th generation mobile communication system)、5G(5th generation mobile communication system)、FRA(Future Radio Access)、NR(new Radio)、W-CDMA(登録商標)、GSM(登録商標)、CDMA2000、UMB(ULtra Mobile Broadband)、IEEE 802.11(Wi-Fi(登録商標))、IEEE 802.16(WiMAX(登録商標))、IEEE 802.20、UWB(ULtra-WIDeBand)、Bluetooth(登録商標)、その他の適切なシステムを利用するシステム及びこれらに基づいて拡張された次世代システムの少なくとも一つに適用されてもよい。また、複数のシステムが組み合わされて(例えば、LTE及びLTE-Aの少なくとも一方と5Gとの組み合わせ等)適用されてもよい。
本発明は、ユーザによって発話された内容を示す発話内容データを取得する取得ステップと、各々の前記ユーザによる発話の時期と、発話した当該ユーザとを特定する特定ステップと、前記発話内容データに基づいて特定される発話の時期と各々のユーザについて特定された前記発話の時期とを比較して、前記発話内容データと、当該発話内容データが示す内容を発話したユーザとを対応付けて登録する登録ステップとを備えることを特徴とする情報処理方法であってもよい。本開示において説明した各態様/実施形態の処理手順、シーケンス、フローチャートなどは、矛盾の無い限り、順序を入れ替えてもよい。例えば、本開示において説明した方法については、例示的な順序を用いて様々なステップの要素を提示しており、提示した特定の順序に限定されない。
入出力された情報等は特定の場所(例えば、メモリ)に保存されてもよいし、管理テーブルを用いて管理してもよい。入出力される情報等は、上書き、更新、又は追記され得る。出力された情報等は削除されてもよい。入力された情報等は他の装置へ送信されてもよい。
判定は、1ビットで表される値(0か1か)によって行われてもよいし、真偽値(Boolean:true又はfalse)によって行われてもよいし、数値の比較(例えば、所定の値との比較)によって行われてもよい。
以上、本開示について詳細に説明したが、当業者にとっては、本開示が本開示中に説明した実施形態に限定されるものではないということは明らかである。本開示は、請求の範囲の記載により定まる本開示の趣旨及び範囲を逸脱することなく修正及び変更態様として実施することができる。したがって、本開示の記載は、例示説明を目的とするものであり、本開示に対して何ら制限的な意味を有するものではない。
ソフトウェアは、ソフトウェア、ファームウェア、ミドルウェア、マイクロコード、ハードウェア記述言語と呼ばれるか、他の名称で呼ばれるかを問わず、命令、命令セット、コード、コードセグメント、プログラムコード、プログラム、サブプログラム、ソフトウェアモジュール、アプリケーション、ソフトウェアアプリケーション、ソフトウェアパッケージ、ルーチン、サブルーチン、オブジェクト、実行可能ファイル、実行スレッド、手順、機能などを意味するよう広く解釈されるべきである。また、ソフトウェア、命令、情報などは、伝送媒体を介して送受信されてもよい。例えば、ソフトウェアが、有線技術(同軸ケーブル、光ファイバケーブル、ツイストペア、デジタル加入者回線(DSL:Digital Subscriber Line)など)及び無線技術(赤外線、マイクロ波など)の少なくとも一方を使用してウェブサイト、サーバ、又は他のリモートソースから送信される場合、これらの有線技術及び無線技術の少なくとも一方は、伝送媒体の定義内に含まれる。
本開示において説明した情報、信号などは、様々な異なる技術のいずれかを使用して表されてもよい。例えば、上記の説明全体に渡って言及され得るデータ、命令、コマンド、情報、信号、ビット、シンボル、チップなどは、電圧、電流、電磁波、磁界若しくは磁性粒子、光場若しくは光子、又はこれらの任意の組み合わせによって表されてもよい。
なお、本開示において説明した用語及び本開示の理解に必要な用語については、同一の又は類似する意味を有する用語と置き換えてもよい。
なお、本開示において説明した用語及び本開示の理解に必要な用語については、同一の又は類似する意味を有する用語と置き換えてもよい。
また、本開示において説明した情報、パラメータなどは、絶対値を用いて表されてもよいし、所定の値からの相対値を用いて表されてもよいし、対応する別の情報を用いて表されてもよい。
本開示において使用する「に基づいて」という記載は、別段に明記されていない限り、「のみに基づいて」を意味しない。言い換えれば、「に基づいて」という記載は、「のみに基づいて」と「に少なくとも基づいて」の両方を意味する。
本開示において使用する「第1」、「第2」などの呼称を使用した要素へのいかなる参照も、それらの要素の量又は順序を全般的に限定しない。これらの呼称は、2つ以上の要素間を区別する便利な方法として本開示において使用され得る。したがって、第1及び第2の要素への参照は、2つの要素のみが採用され得ること、又は何らかの形で第1の要素が第2の要素に先行しなければならないことを意味しない。
上記の各装置の構成における「部」を、「手段」、「回路」、「デバイス」等に置き換えてもよい。
本開示において、「含む(include)」、「含んでいる(including)」及びそれらの変形が使用されている場合、これらの用語は、用語「備える(comprising)」と同様に、包括的であることが意図される。さらに、本開示において使用されている用語「又は(or)」は、排他的論理和ではないことが意図される。
本開示において、例えば、英語でのa,an及びtheのように、翻訳により冠詞が追加された場合、本開示は、これらの冠詞の後に続く名詞が複数形であることを含んでもよい。
本開示において、「AとBが異なる」という用語は、「AとBが互いに異なる」ことを意味してもよい。なお、当該用語は、「AとBがそれぞれCと異なる」ことを意味してもよい。「離れる」、「結合される」などの用語も、「異なる」と同様に解釈されてもよい。
1:情報処理システム、2:通信網、20:Web会議システム、30:情報処理装置、31:発話内容取得部、32:発話時期取得部、33:記憶部、34:音声分析部、35:特定部、36:登録部、37:処理部、3001:プロセッサ、3002:メモリ、3003:ストレージ、3004:通信装置。
Claims (10)
- ユーザによって発話された内容を示す発話内容データを取得する取得部と、
各々の前記ユーザによる発話の時期と、発話した当該ユーザとを特定する特定部と、
前記発話内容データに基づいて特定される発話の時期と各々のユーザについて特定された前記発話の時期とを比較して、前記発話内容データと、当該発話内容データが示す内容を発話したユーザとを対応付けて登録する登録部と
を備えることを特徴とする情報処理装置。 - 前記特定部は、各々の前記ユーザによる発話音声の音量に基づいて、当該ユーザによる発話の時期を特定する
ことを特徴とする請求項1記載の情報処理装置。 - 前記特定部は、各々の前記ユーザによる発話音声に対して音声区間検出を行った結果に基づいて、当該ユーザによる発話の時期を特定する
ことを特徴とする請求項1記載の情報処理装置。 - 前記登録部は、前記発話内容データにおいて複数の前記ユーザによる発話が混在している期間がある場合には、当該期間における当該発話内容データと、当該期間において最も長い発話時間のユーザとを対応付けて登録する
ことを特徴とする請求項1記載の情報処理装置。 - 前記登録部は、前記発話内容データにおいて複数の前記ユーザによる発話が混在している期間において最も長い発話時間のユーザが複数いる場合には、当該期間における当該発話内容データと、当該複数のユーザとを対応付けて登録する
ことを特徴とする請求項4記載の情報処理装置。 - 前記登録部は、前記発話内容データに含まれる1人のユーザによる発話の時期と、特定された前記発話の時期との差分に基づいて、前記発話内容データと前記ユーザとを対応付けて登録する
ことを特徴とする請求項1記載の情報処理装置。 - 前記登録部により登録された内容を用いて処理を行う処理部を備える
ことを特徴とする請求項1記載の情報処理装置。 - 前記処理部は、前記発話内容データを音声認識した結果を前記ユーザごとに記録した発話記録を作成して出力する
ことを特徴とする請求項7記載の情報処理装置。 - 前記処理部は、前記発話内容データを音声認識した結果を前記ユーザごとに記録した発話記録に基づいて、対話形式で応答を行うアルゴリズムに入力するプロンプトを作成する
ことを特徴とする請求項7記載の情報処理装置。 - ユーザによって発話された内容を示す発話内容データを取得する取得ステップと、
各々の前記ユーザによる発話の時期と、発話した当該ユーザとを特定する特定ステップと、
前記発話内容データに基づいて特定される発話の時期と各々のユーザについて特定された前記発話の時期とを比較して、前記発話内容データと、当該発話内容データが示す内容を発話したユーザとを対応付けて登録する登録ステップと
を備えることを特徴とする情報処理方法。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2024/021262 WO2025257952A1 (ja) | 2024-06-12 | 2024-06-12 | 情報処理装置及び方法 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2024/021262 WO2025257952A1 (ja) | 2024-06-12 | 2024-06-12 | 情報処理装置及び方法 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025257952A1 true WO2025257952A1 (ja) | 2025-12-18 |
Family
ID=98050217
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2024/021262 Pending WO2025257952A1 (ja) | 2024-06-12 | 2024-06-12 | 情報処理装置及び方法 |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025257952A1 (ja) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2022016997A (ja) * | 2020-07-13 | 2022-01-25 | ソフトバンク株式会社 | 情報処理方法、情報処理装置及び情報処理プログラム |
| US20240054990A1 (en) * | 2022-08-10 | 2024-02-15 | Actionpower Corp. | Computing device for providing dialogues services |
| JP2024047807A (ja) * | 2022-09-27 | 2024-04-08 | 富士フイルムビジネスイノベーション株式会社 | プログラム及びウェブ会議システム |
-
2024
- 2024-06-12 WO PCT/JP2024/021262 patent/WO2025257952A1/ja active Pending
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2022016997A (ja) * | 2020-07-13 | 2022-01-25 | ソフトバンク株式会社 | 情報処理方法、情報処理装置及び情報処理プログラム |
| US20240054990A1 (en) * | 2022-08-10 | 2024-02-15 | Actionpower Corp. | Computing device for providing dialogues services |
| JP2024047807A (ja) * | 2022-09-27 | 2024-04-08 | 富士フイルムビジネスイノベーション株式会社 | プログラム及びウェブ会議システム |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20200219503A1 (en) | Method and apparatus for filtering out voice instruction | |
| WO2013122310A1 (en) | Method and apparatus for smart voice recognition | |
| CN108665895B (zh) | 用于处理信息的方法、装置和系统 | |
| US8374872B2 (en) | Dynamic update of grammar for interactive voice response | |
| JP7197992B2 (ja) | 音声認識装置、音声認識方法 | |
| JP2021101537A (ja) | グループ通話システム、グループ通話方法及びプログラム | |
| CN109754781A (zh) | 语音翻译终端、移动终端、翻译系统、翻译方法及其装置 | |
| CN110232553A (zh) | 会议支援系统以及计算机可读取的记录介质 | |
| US11323803B2 (en) | Earphone, earphone system, and method in earphone system | |
| CN117012214A (zh) | 多场景优化的对讲设备控制方法、装置、介质及设备 | |
| WO2025257952A1 (ja) | 情報処理装置及び方法 | |
| US11322145B2 (en) | Voice processing device, meeting system, and voice processing method for preventing unintentional execution of command | |
| JP2024115929A (ja) | 音声書き起こしシステム及び音声翻訳システム | |
| CN113889084B (zh) | 音频识别方法、装置、电子设备及存储介质 | |
| WO2023170470A1 (en) | Hearing aid for cognitive help using speaker | |
| JP7512254B2 (ja) | 音声対話システム、モデル生成装置、バージイン発話判定モデル及び音声対話プログラム | |
| WO2023003272A1 (ko) | 오디오 품질을 개선하는 방법 및 전자 디바이스 | |
| CN112435690B (zh) | 双工蓝牙翻译处理方法、装置、计算机设备和存储介质 | |
| JP7837804B2 (ja) | 遠隔会議制御装置 | |
| US20200043492A1 (en) | Speech recognition method and apparatus | |
| CN111582708A (zh) | 医疗信息的检测方法、系统、电子设备及计算机可读存储介质 | |
| CN110855832A (zh) | 一种辅助通话的方法、装置和电子设备 | |
| JP7571111B2 (ja) | 通信端末、情報処理装置、通信方法及びプログラム | |
| CN112752199B (zh) | 一种基于alsa框架的声卡左右声道独立控制装置及方法 | |
| JP2023125442A (ja) | 音声認識装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24943316 Country of ref document: EP Kind code of ref document: A1 |