WO2020253115A1 - 基于语音识别的产品推荐方法、装置、设备和存储介质 - Google Patents
基于语音识别的产品推荐方法、装置、设备和存储介质 Download PDFInfo
- Publication number
- WO2020253115A1 WO2020253115A1 PCT/CN2019/121198 CN2019121198W WO2020253115A1 WO 2020253115 A1 WO2020253115 A1 WO 2020253115A1 CN 2019121198 W CN2019121198 W CN 2019121198W WO 2020253115 A1 WO2020253115 A1 WO 2020253115A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- customer
- data stream
- sales
- voice
- text
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06Q—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
- G06Q30/00—Commerce
- G06Q30/06—Buying, selling or leasing transactions
- G06Q30/0601—Electronic shopping [e-shopping]
- G06Q30/0631—Recommending goods or services
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/26—Speech to text systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L17/00—Speaker identification or verification techniques
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
- G10L25/63—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for estimating an emotional state
Definitions
- This application relates to the field of e-commerce, and in particular to a method, device, equipment and storage medium for product recommendation based on voice recognition.
- the main purpose of this application is to provide a product recommendation method, device, equipment, and storage medium based on voice recognition, which aims to solve the technical problem of inaccurate customer demand analysis during current telemarketing.
- the present application provides a method for product recommendation based on voice recognition.
- the method for product recommendation based on voice recognition includes the following steps:
- the step of processing the voice information to generate a customer data stream and a sales data stream includes:
- the sales voice information is recognized through the preset voice processing model to obtain corresponding sales text data and sales voice feature data, and the sales text data and the sales voice feature data are sorted in a time sequence to generate a sales data stream.
- the present application also provides a voice recognition-based product recommendation device, and the voice recognition-based product recommendation device includes:
- the voice processing module is used to process the voice information to generate a customer data stream and a sales data stream when the voice information sent by the terminal is received;
- the detection and analysis module is configured to obtain the time node and the customer attention text corresponding to the positive emotional fluctuation when a positive emotional fluctuation is detected in the customer data stream;
- a retrospective acquisition module configured to trace the sales data stream according to the time node and the customer attention text, and acquire target sales text data in the sales data stream that causes the positive mood fluctuations;
- the acquiring and sending module is configured to acquire the product information corresponding to the target sales text data in the preset product database, and send the product information to the terminal, so that the sales personnel corresponding to the terminal can introduce the product information according to the product information.
- this application also provides a product recommendation device based on voice recognition
- the voice recognition-based product recommendation device includes: a memory, a processor, and computer-readable instructions stored on the memory and running on the processor, wherein:
- this application also provides a computer storage medium
- the computer storage medium stores computer readable instructions, and when the computer readable instructions are executed by a processor, the steps of the above-mentioned voice recognition-based product recommendation method are realized.
- the method, device, device, and storage medium for product recommendation based on voice recognition proposed in the embodiments of this application, when the server receives the voice information sent by the terminal, processes the voice information to generate a customer data stream and a sales data stream.
- the voice information is divided into customer data streams and sales data streams, and processed separately for customer data streams and sales data streams to achieve detailed analysis.
- the server when positive emotional fluctuations in the customer data stream are detected, the server obtains The time node corresponding to the positive emotion fluctuation and the customer attention text, and then the server traces the sales data stream according to the time node and the customer attention text to determine the target sales text data that causes the customer’s positive emotion fluctuation, And obtain the product information corresponding to the target sales text data in the preset product database, and send the product information to the terminal, so that the corresponding sales staff of the terminal can introduce the product information according to the product information, which realizes the accurate user Demand analysis, and effective product recommendation.
- FIG. 1 is a schematic diagram of the device structure of the hardware operating environment involved in the solution of the embodiment of the present application;
- FIG. 2 is a schematic flowchart of a first embodiment of a product recommendation method based on speech recognition in this application;
- FIG. 3 is a schematic diagram of functional modules of an embodiment of a product recommendation device based on voice recognition in this application.
- Figure 1 is the server of the hardware operating environment involved in the solution of the embodiment of the application (also called the product recommendation device based on voice recognition, where the product recommendation device based on voice recognition can be a separate voice recognition-based product recommendation device
- the product recommendation device is constituted by a combination of other devices and a product recommendation device based on voice recognition).
- the server in the embodiment of the present application refers to a computer that manages resources and provides services for users, and is generally divided into a file server, a database server, and an application-readable instruction server.
- the computer or computer system running the above software is also called a server.
- the server may include: a processor 1001, such as a central processing unit (Central Processing Unit, CPU), network interface 1004, user interface 1003, memory 1005, communication bus 1002, chipset, disk system, network and other hardware.
- the communication bus 1002 is used to implement connection and communication between these components.
- the user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface.
- the network interface 1004 may optionally include a standard wired interface and a wireless interface (such as WIreless-FIdelity, WIFI interface).
- the memory 1005 may be a high-speed random access memory (random access memory, RAM), or stable memory (non-volatile memory), such as disk storage.
- the memory 1005 may also be a storage device independent of the foregoing processor 1001.
- the server may also include a camera, RF (Radio Frequency, radio frequency) circuit, sensor, audio circuit, WiFi module; input unit, display screen, touch screen; network interface can be selected except WiFi, Bluetooth, probe, etc.
- RF Radio Frequency, radio frequency
- the server structure shown in FIG. 1 does not constitute a limitation on the server, and may include more or fewer components than shown in the figure, or a combination of certain components, or different component arrangements.
- the computer software product is stored in a storage medium (storage medium: also called computer storage medium, computer medium, readable medium, readable storage medium, computer readable storage medium, or directly called medium, etc., storage medium
- storage medium can be a non-volatile readable storage medium, such as RAM, magnetic disk, optical disk, and includes several instructions to make a terminal device (can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute this application
- the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and computer-readable instructions.
- the network interface 1004 is mainly used to connect to the back-end database and perform data communication with the back-end database;
- the user interface 1003 is mainly used to connect to the client (the client, also called the user terminal or the terminal, the embodiment of the application
- the terminal can be a fixed terminal or a mobile terminal, such as a PC, a smart phone, a tablet computer, an e-book reader, a portable computer, etc.
- the terminal contains sensors such as light sensors, motion sensors and other sensors, which will not be repeated here), Perform data communication with the client; and the processor 1001 may be used to call computer-readable instructions stored in the memory 1005, and execute the steps in the voice recognition-based product recommendation method provided in the following embodiments of the present application.
- the first embodiment of the present application provides a product recommendation method based on voice recognition, which is applied to the server shown in FIG. 1.
- the product recommendation method based on voice recognition includes:
- Step S10 When the voice information sent by the terminal is received, the voice information is processed to generate a customer data stream and a sales data stream.
- the sales staff conduct telephone product sales through the terminal.
- the terminal collects the voice information of the call.
- the voice information of the call includes: the sales voice information of the salesperson and the customer voice information of the customer.
- the terminal sends the collected voice information to the server, and the server receives
- the voice information sent by the terminal and the voice information received by the server processing specifically, include:
- Step S11 When the voice information sent by the terminal is received, the voice information is recognized through a preset voice processing model to obtain the voiceprint feature corresponding to the voice information, and the voice information is divided into Customer voice messages and sales voice messages.
- the server receives the voice information sent by the terminal, that is, the voice recognition model is preset in the server, and the preset voice recognition model is the voice recognition algorithm obtained by preset training.
- the voice recognition algorithm can realize the voice recognition of the voice information.
- the server uses the preset voice recognition model to extract voice feature data in the voice information, and determines the voice feature of the voice information based on the extracted voice feature data.
- the server converts the voice information according to the voice feature data. Divided into customer voice information and sales voice information.
- the server in this embodiment also divides the voice information into sales voice information or customer voice information according to other principles, which will not be repeated in this embodiment.
- the server divides the voice information according to the voice content of the voice information.
- the voice content is: I am the xxx server of the xxx company to determine that the voice information is sales voice information.
- Step S12 Recognize the customer voice information through the preset voice processing model, obtain corresponding customer text data and customer voice feature data, and sort the customer text data and the customer voice feature data in a time sequence to generate customer data flow.
- the server After the server divides the voice information into customer voice information and sales voice information, the server processes the customer voice information to generate a customer data stream; specifically, the server inputs the customer voice information into the preset voice recognition model, and the preset voice recognition model first The customer voice information is denoised, and then each frame of the customer voice information is recognized as a state, and the states are further combined into phonemes. Finally, each phoneme is combined into words to generate the customer text data corresponding to the customer voice information. The server converts the customer text The data is sorted by time to generate a customer text data stream.
- the preset voice processing model extracts the customer voice feature data from the customer voice information.
- the customer voice feature data includes the frequency, pitch, amplitude, etc. of the customer voice.
- the server sorts the customer voice feature data according to time to obtain the customer voice feature Data stream, the server combines the customer text data stream and the customer voice feature data stream in chronological order to obtain the customer data stream; it can be understood that the customer data stream includes the customer voice feature data stream and the customer text data stream.
- the customer can be The data stream is analogous to the music score.
- the customer voice feature data stream in the customer data stream is equivalent to the notes in the music score
- the customer text data stream in the customer data stream is equivalent to the lyrics in the music score.
- Step S13 Recognize the sales voice information through the preset voice processing model to obtain corresponding sales text data and sales voice feature data, and sort the sales text data and the sales voice feature data in a time sequence to generate sales data flow.
- the server processes the sales voice information to generate a sales data stream; specifically, the server inputs the sales voice information into a preset voice recognition model, and the preset voice recognition model first denoises the sales voice information, and then transfers the sales voice Each frame of the information is recognized as a state, and the states are further combined into phonemes, and finally each phoneme is combined into words to generate sales text data corresponding to the sales voice information.
- the server sorts the sales text data according to time to generate a sales text data stream;
- the preset voice processing model extracts the sales voice feature data in the sales voice information, where the sales voice feature data is the frequency, pitch, amplitude, etc. of the salesperson’s voice, and the server sorts the sales voice feature data according to time to obtain Sales voice feature data stream; the server combines the sales text data stream and the sales voice feature data stream in chronological order to obtain the sales data stream.
- voice information is processed to generate customer data streams and sales data streams to facilitate accurate and detailed analysis during subsequent voice analysis.
- voice information is processed to generate customer data streams and sales data streams to facilitate accurate and detailed analysis during subsequent voice analysis.
- step S20 when it is detected that a positive emotion fluctuation occurs in the customer data stream, a time node and a customer attention text corresponding to the positive emotion fluctuation are obtained.
- the server can analyze the customer data stream in different ways:
- method 1 The server performs analysis based on the client text data stream in the client data stream, specifically, including:
- Step a Compare the customer text data in the customer data stream with the target words in the preset target vocabulary.
- Step b When there is target customer text data matching the target word, it is determined that positive mood fluctuations occur in the customer data stream.
- Step c Use the target customer text data as customer attention text, and obtain the time node corresponding to the customer attention text in the customer data stream.
- the server obtains the client text data stream in the client data stream, and the server compares the client text data in the client text data stream with the target words in the preset target vocabulary, where the preset target vocabulary refers to the advance
- the database is set to store the target words.
- the target words in the preset target word database refer to words related to the product, for example, the target words are: performance, price, etc.
- the positive emotion fluctuations in this embodiment of the application refer to the customer
- the server takes the target customer text data corresponding node as the positive emotional fluctuation point
- the server takes the target customer text data as the customer attention text
- the server performs a combined analysis based on the customer voice feature data stream and the customer text data stream in the customer data stream, specifically, including:
- the server obtains each voice feature data in the customer voice feature data stream, and the server sets the voice feature data higher than the preset voice feature data (the preset voice feature data refers to the user's usual voice frequency, pitch, and timbre set in advance)
- the time point is regarded as the point of positive mood fluctuation.
- the positive mood fluctuation in the embodiment of this application refers to the change of voice feature data such as question tone appearing in the customer voice information when the customer is interested in the product, and the server obtains the customer voice The time node corresponding to the positive emotion fluctuation in the characteristic data stream, and then the server obtains the customer attention text of the time node in the customer text data stream.
- Step S30 Trace the sales data stream according to the time node and the customer attention text, and obtain target sales text data in the sales data stream that causes the positive mood swing.
- Step a Obtain a target sales text data stream for a preset time period before the time node in the sales text data stream, and compare the sales text data in the target sales text data stream with the customer attention text;
- Step b Obtain target sales text data matching the customer focus text.
- the server obtains the target sales text data stream for a preset time period before the time node in the sales text data stream, where the preset time period refers to a preset time interval, and the preset time period can be flexibly set according to specific scenarios, for example,
- the preset time period is set to 1 minute, that is, when the server determines that the time node of the positive mood swing is 15:40:30, the server obtains the target sales data between 15:39:30 and 15:40:30 Stream, the server obtains the target sales text data stream in the target sales data stream, and compares the sales text data in the target sales text data stream with the customer attention text; to determine whether there is a customer attention text match in the target sales text data stream If there is no target sales text data matching the customer focus text in the target sales text data stream, the server sends the customer focus text as prompt information to the terminal so that the sales staff corresponding to the terminal can understand the customer focus text.
- the server obtains the target sales text data matching the customer's attention text to perform product information query based on the target sales text data, specifically:
- Step S40 Obtain the product information corresponding to the target sales text data in the preset product database, and send the product information to the terminal, so that a salesperson corresponding to the terminal can introduce the product information according to the product information.
- the server recommends product information according to the target sales text data.
- the product database is preset in the server, and the product information is stored in the preset product database.
- the server queries the preset product database to obtain the target sales text data in the preset product database.
- the server sends the product information to the terminal for the terminal's corresponding sales staff to introduce the product information.
- the server converts voice information into sales data streams and customer data streams.
- the time nodes and customer attention texts corresponding to the positive emotional fluctuations are determined by the server based on the time nodes and The customer pays attention to the text to analyze the sales data flow, and obtains the corresponding product information according to the target sales data, realizes accurate user demand analysis, and effectively introduces products to avoid calls caused by sales staff not understanding user needs or product information The problem of difficult sales.
- This embodiment is a refinement of step S20 in the first embodiment of the present application.
- the method in which the server determines the positive mood fluctuations according to the customer voice feature data stream and the customer text data stream in the customer data stream is specifically explained.
- the product recommendation methods based on speech recognition include:
- Step S21 Acquire basic feature data in the customer voice feature data stream, and when the customer voice feature data stream is higher than the basic feature data, it is determined that positive mood fluctuations occur in the customer data stream.
- the server obtains the basic feature data in the customer's voice feature data stream, where the basic feature data is determined by the server according to the voice frequency, amplitude, and pitch in the customer's voice feature data stream. Take a voice feature parameter of frequency as an example for illustration. 10% of the voice feature data stream is less than 30 Hz, 80% of the frequency is 30-50 Hz, and 10% of the frequency is greater than 50 Hz.
- the server sets the basic feature data to 50 Hz, and the server sets the customer voice feature data stream higher than the basic feature data
- the target customer’s voice feature data is used as a positive mood swing.
- Step S22 Obtain the time node corresponding to the positive emotional fluctuation in the customer voice feature data stream, and obtain the customer attention text corresponding to the time node in the customer text data stream.
- the server obtains the time node corresponding to the positive mood fluctuation in the customer voice characteristic data stream, and the server determines the customer's attention point at this time node, that is, the server obtains the customer attention text corresponding to the time node in the customer text data stream.
- the server performs accurate customer demand analysis according to the customer data stream and the sales data stream, which improves the accuracy of customer demand analysis.
- This embodiment is a step after step S40 in the first implementation.
- the server after the server sends the product information to the terminal, the server detects the customer feedback voice information sent by the terminal, and updates the preset product database according to the customer feedback voice information.
- the product recommendation method based on voice recognition includes:
- Step S50 Acquire customer feedback voice information based on the product information, and extract questions of interest from the customer feedback voice information.
- the salesperson introduces the product information of the terminal.
- the terminal receives the customer feedback voice information and sends it to the server.
- the server receives the customer feedback voice information.
- the customer feedback voice information refers to the customer's announcement based on the salesperson.
- the server obtains the customer’s feedback voice information and extracts the question of interest, that is, the server obtains the product performance information, product price information or product logistics information in the customer’s feedback voice information as the question of interest.
- Step S60 Count the number of questions of each interest question, and when the number of questions of interest exceeds a preset threshold, add the interest question to the preset product database.
- the server counts the number of questions asked for each question of interest, that is, the server records the obtained questions of interest separately, and when there are repetitions, the server accumulates, and the server detects that the number of questions asked for the question of interest exceeds a preset threshold (the preset threshold is preset Set the number of times, the preset threshold can be set according to specific conditions, for example, when the preset threshold is set to 10 times), the question of interest is added to the preset product database.
- the server updates the preset product database according to the voice information received from the customer, so that the later product recommendation is more intelligent.
- This embodiment is a step after step S40 in the first embodiment.
- the server can score customers according to customer data streams to realize potential customer mining.
- the voice recognition-based product recommendation method includes :
- Step S70 When it is detected that the voice information transmission of the terminal is suspended, acquire the customer voice time and customer text data in the customer data stream.
- the server When the server detects that the terminal's voice information transmission is suspended, that is, when the server detects that a customer communication is completed, the server obtains the customer voice time and customer text data in the customer data stream, where the customer voice time refers to the total time of the customer voice information.
- Step S80 the customer data stream is scored according to the customer voice time and the customer text data, and when the score is higher than a preset score value, the customer corresponding to the customer data stream is regarded as a target customer and marked.
- the server scores the customer data stream according to the customer voice time and customer text data.
- the customer voice time in the customer data stream is more than 2 minutes
- the customer text data contains: Please introduce the xxx product, the customer data stream is scored 8 -10 points; the customer voice time in the customer data stream is 1 to 2 minutes, the customer data stream score is 4-7 points; the customer voice time in the customer data stream is less than 2 minutes, the customer data stream score is 0-3 points
- the server sets the customer data stream corresponding to the customer whose score is higher than the preset score value (the preset score value is a preset score value, and the preset score value can be set to 6 points) as target customers and marks them.
- the data stream identifies and identifies target customers who have a tendency to purchase the corresponding product for later follow-up, making telemarketing more intelligent.
- an embodiment of the present application also proposes a product recommendation device based on voice recognition, and the product recommendation device based on voice recognition includes:
- the voice processing module 10 is configured to process the voice information to generate a customer data stream and a sales data stream when the voice information sent by the terminal is received;
- the detection and analysis module 20 is configured to obtain the time node and customer attention text corresponding to the positive emotional fluctuation when a positive emotional fluctuation is detected in the customer data stream;
- the retrospective acquisition module 30 is configured to trace the sales data stream according to the time node and the customer attention text, and acquire the target sales text data in the sales data stream that causes the positive mood fluctuations;
- the obtaining and sending module 40 is used to obtain the product information corresponding to the target sales text data in the preset product database, and send the product information to the terminal, so that the corresponding sales staff of the terminal can introduce the product information according to the product information. .
- the voice processing module 10 includes:
- the voice receiving unit is configured to recognize the voice information through a preset voice processing model when receiving the voice information sent by the terminal, obtain the voiceprint characteristics corresponding to the voice information, and divide the voice information according to the voiceprint Features are divided into customer voice information and sales voice information;
- the first generating unit is configured to recognize the customer voice information through the preset voice processing model to obtain corresponding customer text data and customer voice feature data, and arrange the customer text data and the customer voice feature data in a time sequence Sort and generate customer data stream;
- the second generating unit is configured to recognize the sales voice information through the preset voice processing model to obtain corresponding sales text data and sales voice feature data, and arrange the sales text data and the sales voice feature data in a time sequence Sort to generate sales data stream.
- the detection and analysis module 20 includes:
- the information comparison unit is used to compare the customer text data in the customer data stream with the target words in the preset target vocabulary
- the comparison and determination unit is configured to determine that there is a positive mood swing in the customer data stream when there is target customer text data that matches the target word;
- the first acquiring unit is configured to use the target customer text data as customer attention text, and obtain the time node corresponding to the customer attention text in the customer data stream.
- the customer data stream includes a customer voice feature data stream and a customer text data stream;
- the detection and analysis module 20 includes:
- the information analysis unit is used to obtain basic feature data in the customer voice feature data stream, and determine that positive emotions appear in the customer data stream when the customer voice feature data stream is higher than the basic feature data fluctuation;
- the second acquiring unit is configured to acquire the time node corresponding to the positive mood fluctuation in the customer voice feature data stream, and acquire the customer attention text corresponding to the time node in the customer text data stream.
- the sales data stream includes a sales text data stream
- the retrospective acquisition module 30 includes:
- the information comparison unit is used to obtain the target sales text data stream of the preset time period before the time node in the sales text data stream, and compare the sales text data in the target sales text data stream with the customer attention text Compare
- the information acquisition unit is used to acquire target sales text data that matches the customer's attention text.
- the product recommendation device based on voice recognition includes:
- An acquisition and extraction module configured to acquire customer feedback voice information based on the product information, and extract questions of interest from the customer feedback voice information
- the statistical update module is used to count the number of questions asked for each question of interest, and when the number of questions asked for the question of interest exceeds a preset threshold, add the question of interest to the preset product database.
- the product recommendation device based on voice recognition further includes:
- the detection and acquisition module is configured to acquire the customer voice time and customer text data in the customer data stream when it is detected that the voice information transmission of the terminal is suspended;
- the evaluation marking module is used to score the customer data stream according to the customer voice time and the customer text data, and when the score is higher than a preset score value, the customer corresponding to the customer data stream is regarded as the target customer and marked .
- each functional module of the voice recognition-based product recommendation device can refer to the various embodiments of the voice recognition-based product recommendation method of the present application, which will not be repeated here.
- the embodiment of the present application also proposes a computer storage medium.
- the computer storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the operations in the voice recognition-based product recommendation method provided in the foregoing embodiments are implemented.
Landscapes
- Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Physics & Mathematics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Computational Linguistics (AREA)
- Business, Economics & Management (AREA)
- Finance (AREA)
- Accounting & Taxation (AREA)
- Signal Processing (AREA)
- Marketing (AREA)
- Hospice & Palliative Care (AREA)
- General Health & Medical Sciences (AREA)
- Child & Adolescent Psychology (AREA)
- Development Economics (AREA)
- Economics (AREA)
- Psychiatry (AREA)
- Strategic Management (AREA)
- General Business, Economics & Management (AREA)
- General Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Telephonic Communication Services (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
Abstract
Description
Claims (20)
- 一种基于语音识别的产品推荐方法,其特征在于,所述基于语音识别的产品推荐方法包括以下步骤:在接收到终端发送的语音信息时,将所述语音信息处理生成客户数据流和销售数据流;在检测到所述客户数据流中出现正向情绪波动时,获取所述正向情绪波动对应的时间节点和客户关注文本;按所述时间节点和所述客户关注文本追溯所述销售数据流,获取所述销售数据流中引起所述正向情绪波动的目标销售文本数据;获取预设产品数据库中所述目标销售文本数据对应的产品信息,将所述产品信息发送至所述终端,以供所述终端对应销售人员按所述产品信息进行介绍;其中,所述在接收到终端发送的语音信息时,将所述语音信息处理生成客户数据流和销售数据流的步骤,包括:在接收到终端发送的语音信息时,通过预设语音处理模型识别所述语音信息,得到所述语音信息对应的声纹特征,并将所述语音信息按所述声纹特征划分为客户语音信息和销售语音信息;通过所述预设语音处理模型识别所述客户语音信息,得到对应的客户文本数据与客户语音特征数据,将所述客户文本数据与所述客户语音特征数据按时间序列排序生成客户数据流;通过所述预设语音处理模型识别所述销售语音信息,得到对应的销售文本数据与销售语音特征数据,将所述销售文本数据与所述销售语音特征数据按时间序列排序生成销售数据流。
- 如权利要求1所述的基于语音识别的产品推荐方法,其特征在于,所述在检测到所述客户数据流中出现正向情绪波动时,获取所述正向情绪波动对应的时间节点和客户关注文本的步骤,包括:将所述客户数据流中的客户文本数据与预设目标词库中的目标词语进行比对;在存在与所述目标词语匹配的目标客户文本数据时,判定所述客户数据流中出现正向情绪波动;将所述目标客户文本数据作为客户关注文本,并获取所述客户关注文本在所述客户数据流中对应的时间节点。
- 如权利要求1所述的基于语音识别的产品推荐方法,其特征在于,所述客户数据流包括客户语音特征数据流和客户文本数据流;所述在检测到所述客户数据流中出现正向情绪波动时,获取所述正向情绪波动对应的时间节点和客户关注文本的步骤,包括:获取所述客户语音特征数据流中的基础特征数据,在所述客户语音特征数据流中出现高于所述基础特征数据时,判定所述客户数据流中出现正向情绪波动;获取所述客户语音特征数据流中所述正向情绪波动对应的时间节点,并获取所述客户文本数据流中所述时间节点对应的客户关注文本。
- 如权利要求1所述的基于语音识别的产品推荐方法,其特征在于,所述销售数据流中包括销售文本数据流;所述按所述时间节点和所述客户关注文本追溯所述销售数据流,获取所述销售数据流中引起所述正向情绪波动的目标销售文本数据的步骤,包括:获取所述销售文本数据流中所述时间节点之前预设时间段的目标销售文本数据流,将所述目标销售文本数据流中的销售文本数据与所述客户关注文本进行比对;获取与所述客户关注文本匹配的目标销售文本数据。
- 如权利要求1所述的基于语音识别的产品推荐方法,其特征在于,所述获取预设产品数据库中所述目标销售文本数据对应的产品信息,将所述产品信息发送至所述终端的步骤之后,包括:获取基于所述产品信息的客户反馈语音信息,从所述客户反馈语音信息中提取感兴趣问题;统计各感兴趣问题的提问次数,在感兴趣问题提问次数超过预设阈值时,将所述感兴趣问题添加到所述预设产品数据库。
- 如权利要求1所述的基于语音识别的产品推荐方法,其特征在于,所述获取预设产品数据库中所述目标销售文本数据对应的产品信息,将所述产品信息发送至所述终端的步骤之后,包括:在检测到所述终端的语音信息发送中止时,获取所述客户数据流中的客户语音时间和客户文本数据;按所述客户语音时间和所述客户文本数据对所述客户数据流进行评分,在评分高于预设评分值时,将所述客户数据流对应客户作为目标客户并标记。
- 一种基于语音识别的产品推荐装置,其特征在于,所述基于语音识别的产品推荐装置包括:语音处理模块,用于在接收到终端发送的语音信息时,将所述语音信息处理生成客户数据流和销售数据流;检测分析模块,用于在检测到所述客户数据流中出现正向情绪波动时,获取所述正向情绪波动对应的时间节点和客户关注文本;追溯获取模块,用于按所述时间节点和所述客户关注文本追溯所述销售数据流,获取所述销售数据流中引起所述正向情绪波动的目标销售文本数据;获取发送模块,用于获取预设产品数据库中所述目标销售文本数据对应的产品信息,将所述产品信息发送至所述终端,以供所述终端对应销售人员按所述产品信息进行介绍;其中,所述语音处理模块,包括:语音接收单元,用于在接收到终端发送的语音信息时,通过预设语音处理模型识别所述语音信息,得到所述语音信息对应的声纹特征,并将所述语音信息按所述声纹特征划分为客户语音信息和销售语音信息;第一生成单元,用于通过所述预设语音处理模型识别所述客户语音信息,得到对应的客户文本数据与客户语音特征数据,将所述客户文本数据与所述客户语音特征数据按时间序列排序生成客户数据流;第二生成单元,用于通过所述预设语音处理模型识别所述销售语音信息,得到对应的销售文本数据与销售语音特征数据,将所述销售文本数据与所述销售语音特征数据按时间序列排序生成销售数据流。
- 如权利要求7所述的基于语音识别的产品推荐装置,其特征在于,所述检测分析模块,包括:信息比对单元,用于将所述客户数据流中的客户文本数据与预设目标词库中的目标词语进行比对;比对判定单元,用于在存在与所述目标词语匹配的目标客户文本数据时,判定所述客户数据流中出现正向情绪波动;第一获取单元,用于将所述目标客户文本数据作为客户关注文本,并获取所述客户关注文本在所述客户数据流中对应的时间节点。
- 如权利要求7所述的基于语音识别的产品推荐装置,其特征在于,所述客户数据流包括客户语音特征数据流和客户文本数据流;所述检测分析模块,包括:信息分析单元,用于获取所述客户语音特征数据流中的基础特征数据,在所述客户语音特征数据流中出现高于所述基础特征数据时,判定所述客户数据流中出现正向情绪波动;第二获取单元,用于获取所述客户语音特征数据流中所述正向情绪波动对应的时间节点,并获取所述客户文本数据流中所述时间节点对应的客户关注文本。
- 如权利要求7所述的基于语音识别的产品推荐装置,其特征在于,所述销售数据流中包括销售文本数据流;所述追溯获取模块,包括:信息比对单元,用于获取所述销售文本数据流中所述时间节点之前预设时间段的目标销售文本数据流,将所述目标销售文本数据流中的销售文本数据与所述客户关注文本进行比对;信息获取单元,用于获取与所述客户关注文本匹配的目标销售文本数据。
- 如权利要求7所述的基于语音识别的产品推荐装置,其特征在于,所述的基于语音识别的产品推荐装置,包括:获取提取模块,用于获取基于所述产品信息的客户反馈语音信息,从所述客户反馈语音信息中提取感兴趣问题;统计更新模块,用于统计各感兴趣问题的提问次数,在感兴趣问题提问次数超过预设阈值时,将所述感兴趣问题添加到所述预设产品数据库。
- 如权利要求7所述的基于语音识别的产品推荐装置,其特征在于,所述的基于语音识别的产品推荐装置,还包括:检测获取模块,用于在检测到所述终端的语音信息发送中止时,获取所述客户数据流中的客户语音时间和客户文本数据;评价标记模块,用于按所述客户语音时间和所述客户文本数据对所述客户数据流进行评分,在评分高于预设评分值时,将所述客户数据流对应客户作为目标客户并标记。
- 一种基于语音识别的产品推荐设备,其特征在于,所述基于语音识别的产品推荐设备包括:存储器、处理器及存储在所述存储器上并可在所述处理器上运行的计算机可读指令,其中:所述计算机可读指令被所述处理器执行时实现如以下的步骤:在接收到终端发送的语音信息时,将所述语音信息处理生成客户数据流和销售数据流;在检测到所述客户数据流中出现正向情绪波动时,获取所述正向情绪波动对应的时间节点和客户关注文本;按所述时间节点和所述客户关注文本追溯所述销售数据流,获取所述销售数据流中引起所述正向情绪波动的目标销售文本数据;获取预设产品数据库中所述目标销售文本数据对应的产品信息,将所述产品信息发送至所述终端,以供所述终端对应销售人员按所述产品信息进行介绍;其中,所述在接收到终端发送的语音信息时,将所述语音信息处理生成客户数据流和销售数据流的步骤,包括:在接收到终端发送的语音信息时,通过预设语音处理模型识别所述语音信息,得到所述语音信息对应的声纹特征,并将所述语音信息按所述声纹特征划分为客户语音信息和销售语音信息;通过所述预设语音处理模型识别所述客户语音信息,得到对应的客户文本数据与客户语音特征数据,将所述客户文本数据与所述客户语音特征数据按时间序列排序生成客户数据流;通过所述预设语音处理模型识别所述销售语音信息,得到对应的销售文本数据与销售语音特征数据,将所述销售文本数据与所述销售语音特征数据按时间序列排序生成销售数据流。
- 如权利要求13所述的基于语音识别的产品推荐设备,其特征在于,所述在检测到所述客户数据流中出现正向情绪波动时,获取所述正向情绪波动对应的时间节点和客户关注文本的步骤,包括:将所述客户数据流中的客户文本数据与预设目标词库中的目标词语进行比对;在存在与所述目标词语匹配的目标客户文本数据时,判定所述客户数据流中出现正向情绪波动;将所述目标客户文本数据作为客户关注文本,并获取所述客户关注文本在所述客户数据流中对应的时间节点。
- 如权利要求13所述的基于语音识别的产品推荐设备,其特征在于,所述客户数据流包括客户语音特征数据流和客户文本数据流;所述在检测到所述客户数据流中出现正向情绪波动时,获取所述正向情绪波动对应的时间节点和客户关注文本的步骤,包括:获取所述客户语音特征数据流中的基础特征数据,在所述客户语音特征数据流中出现高于所述基础特征数据时,判定所述客户数据流中出现正向情绪波动;获取所述客户语音特征数据流中所述正向情绪波动对应的时间节点,并获取所述客户文本数据流中所述时间节点对应的客户关注文本。
- 如权利要求13所述的基于语音识别的产品推荐设备,其特征在于,所述销售数据流中包括销售文本数据流;所述按所述时间节点和所述客户关注文本追溯所述销售数据流,获取所述销售数据流中引起所述正向情绪波动的目标销售文本数据的步骤,包括:获取所述销售文本数据流中所述时间节点之前预设时间段的目标销售文本数据流,将所述目标销售文本数据流中的销售文本数据与所述客户关注文本进行比对;获取与所述客户关注文本匹配的目标销售文本数据。
- 如权利要求13所述的基于语音识别的产品推荐设备,其特征在于,所述获取预设产品数据库中所述目标销售文本数据对应的产品信息,将所述产品信息发送至所述终端的步骤之后,包括:获取基于所述产品信息的客户反馈语音信息,从所述客户反馈语音信息中提取感兴趣问题;统计各感兴趣问题的提问次数,在感兴趣问题提问次数超过预设阈值时,将所述感兴趣问题添加到所述预设产品数据库。
- 如权利要求13所述的基于语音识别的产品推荐设备,其特征在于,所述获取预设产品数据库中所述目标销售文本数据对应的产品信息,将所述产品信息发送至所述终端的步骤之后,包括:在检测到所述终端的语音信息发送中止时,获取所述客户数据流中的客户语音时间和客户文本数据;按所述客户语音时间和所述客户文本数据对所述客户数据流进行评分,在评分高于预设评分值时,将所述客户数据流对应客户作为目标客户并标记。
- 一种计算机存储介质,其特征在于,所述计算机存储介质上存储有计算机可读指令,所述计算机可读指令被处理器执行时实现以下的步骤:在接收到终端发送的语音信息时,将所述语音信息处理生成客户数据流和销售数据流;在检测到所述客户数据流中出现正向情绪波动时,获取所述正向情绪波动对应的时间节点和客户关注文本;按所述时间节点和所述客户关注文本追溯所述销售数据流,获取所述销售数据流中引起所述正向情绪波动的目标销售文本数据;获取预设产品数据库中所述目标销售文本数据对应的产品信息,将所述产品信息发送至所述终端,以供所述终端对应销售人员按所述产品信息进行介绍;其中,所述在接收到终端发送的语音信息时,将所述语音信息处理生成客户数据流和销售数据流的步骤,包括:在接收到终端发送的语音信息时,通过预设语音处理模型识别所述语音信息,得到所述语音信息对应的声纹特征,并将所述语音信息按所述声纹特征划分为客户语音信息和销售语音信息;通过所述预设语音处理模型识别所述客户语音信息,得到对应的客户文本数据与客户语音特征数据,将所述客户文本数据与所述客户语音特征数据按时间序列排序生成客户数据流;通过所述预设语音处理模型识别所述销售语音信息,得到对应的销售文本数据与销售语音特征数据,将所述销售文本数据与所述销售语音特征数据按时间序列排序生成销售数据流。
- 如权利要求19所述的计算机存储介质,其特征在于,所述在检测到所述客户数据流中出现正向情绪波动时,获取所述正向情绪波动对应的时间节点和客户关注文本的步骤,包括:将所述客户数据流中的客户文本数据与预设目标词库中的目标词语进行比对;在存在与所述目标词语匹配的目标客户文本数据时,判定所述客户数据流中出现正向情绪波动;将所述目标客户文本数据作为客户关注文本,并获取所述客户关注文本在所述客户数据流中对应的时间节点。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201910535455.6A CN110335596A (zh) | 2019-06-19 | 2019-06-19 | 基于语音识别的产品推荐方法、装置、设备和存储介质 |
| CN201910535455.6 | 2019-06-19 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020253115A1 true WO2020253115A1 (zh) | 2020-12-24 |
Family
ID=68142195
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/121198 Ceased WO2020253115A1 (zh) | 2019-06-19 | 2019-11-27 | 基于语音识别的产品推荐方法、装置、设备和存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN110335596A (zh) |
| WO (1) | WO2020253115A1 (zh) |
Families Citing this family (10)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110335596A (zh) * | 2019-06-19 | 2019-10-15 | 深圳壹账通智能科技有限公司 | 基于语音识别的产品推荐方法、装置、设备和存储介质 |
| CN110880330A (zh) * | 2019-10-28 | 2020-03-13 | 维沃移动通信有限公司 | 音频转换方法及终端设备 |
| CN111107400B (zh) * | 2019-12-30 | 2022-06-10 | 深圳Tcl数字技术有限公司 | 数据收集方法、装置、智能电视及计算机可读存储介质 |
| CN111626813B (zh) * | 2020-04-22 | 2023-09-29 | 北京水滴科技集团有限公司 | 产品推荐方法及其系统 |
| CN111680337B (zh) * | 2020-06-04 | 2021-07-06 | 宁波智讯联科科技有限公司 | Pdm系统产品设计需求信息获取方法及系统 |
| CN112633992A (zh) * | 2021-01-11 | 2021-04-09 | 上海明略人工智能(集团)有限公司 | 基于语音识别的销售管理方法及系统 |
| CN113657927A (zh) * | 2021-08-02 | 2021-11-16 | 上海明略人工智能(集团)有限公司 | 依据销售录音管理门店的方法及装置 |
| CN114974255A (zh) * | 2022-05-16 | 2022-08-30 | 上海华客信息科技有限公司 | 基于酒店场景的声纹识别方法、系统、设备及存储介质 |
| CN115063154A (zh) * | 2022-06-23 | 2022-09-16 | 平安银行股份有限公司 | 基于兴趣度的业务推荐方法、装置、终端及存储介质 |
| CN119848198A (zh) * | 2024-12-25 | 2025-04-18 | 内蒙古同远咨询集团股份有限公司 | 一种信息技术咨询服务管理方法及系统 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104125348A (zh) * | 2014-07-04 | 2014-10-29 | 北京智谷睿拓技术服务有限公司 | 通信控制方法、装置和智能终端 |
| CN108961072A (zh) * | 2018-06-07 | 2018-12-07 | 平安科技(深圳)有限公司 | 推送保险产品的方法、装置、计算机设备和存储介质 |
| WO2019071599A1 (en) * | 2017-10-13 | 2019-04-18 | Microsoft Technology Licensing, Llc | PROVIDING AN ANSWER IN A SESSION |
| CN110335596A (zh) * | 2019-06-19 | 2019-10-15 | 深圳壹账通智能科技有限公司 | 基于语音识别的产品推荐方法、装置、设备和存储介质 |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN103811009A (zh) * | 2014-03-13 | 2014-05-21 | 华东理工大学 | 一种基于语音分析的智能电话客服系统 |
| CN107341685A (zh) * | 2017-05-24 | 2017-11-10 | 百度在线网络技术(北京)有限公司 | 数据分析方法及装置 |
| US10311454B2 (en) * | 2017-06-22 | 2019-06-04 | NewVoiceMedia Ltd. | Customer interaction and experience system using emotional-semantic computing |
| CN109767791B (zh) * | 2019-03-21 | 2021-03-30 | 中国—东盟信息港股份有限公司 | 一种针对呼叫中心通话的语音情绪识别及应用系统 |
-
2019
- 2019-06-19 CN CN201910535455.6A patent/CN110335596A/zh active Pending
- 2019-11-27 WO PCT/CN2019/121198 patent/WO2020253115A1/zh not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104125348A (zh) * | 2014-07-04 | 2014-10-29 | 北京智谷睿拓技术服务有限公司 | 通信控制方法、装置和智能终端 |
| WO2019071599A1 (en) * | 2017-10-13 | 2019-04-18 | Microsoft Technology Licensing, Llc | PROVIDING AN ANSWER IN A SESSION |
| CN108961072A (zh) * | 2018-06-07 | 2018-12-07 | 平安科技(深圳)有限公司 | 推送保险产品的方法、装置、计算机设备和存储介质 |
| CN110335596A (zh) * | 2019-06-19 | 2019-10-15 | 深圳壹账通智能科技有限公司 | 基于语音识别的产品推荐方法、装置、设备和存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN110335596A (zh) | 2019-10-15 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2020253115A1 (zh) | 基于语音识别的产品推荐方法、装置、设备和存储介质 | |
| WO2020034526A1 (zh) | 保险录音的质检方法、装置、设备和计算机存储介质 | |
| WO2021051558A1 (zh) | 基于知识图谱的问答方法、装置和存储介质 | |
| CN112951275B (zh) | 语音质检方法、装置、电子设备及介质 | |
| WO2020207035A1 (zh) | 骚扰电话拦截方法、装置、设备及存储介质 | |
| WO2020015067A1 (zh) | 数据采集方法、装置、设备及存储介质 | |
| WO2020107761A1 (zh) | 广告文案处理方法、装置、设备及计算机可读存储介质 | |
| WO2020253112A1 (zh) | 测试策略的获取方法、装置、终端及可读存储介质 | |
| WO2020139058A1 (en) | Cross-device voiceprint recognition | |
| WO2021010744A1 (ko) | 음성 인식 기반의 세일즈 대화 분석 방법 및 장치 | |
| WO2016112558A1 (zh) | 智能交互系统中的问题匹配方法和系统 | |
| WO2020258657A1 (zh) | 异常检测方法、装置、计算机设备及存储介质 | |
| WO2020107762A1 (zh) | Ctr预估方法、装置及计算机可读存储介质 | |
| WO2015178600A1 (en) | Speech recognition method and apparatus using device information | |
| WO2020082766A1 (zh) | 输入法的联想方法、装置、设备及可读存储介质 | |
| WO2020119069A1 (zh) | 基于自编码神经网络的文本生成方法、装置、终端及介质 | |
| WO2020256204A1 (ko) | 텍스트의 내용 및 감정 분석에 기반한 답변 추천 시스템 및 방법 | |
| WO2021251539A1 (ko) | 인공신경망을 이용한 대화형 메시지 구현 방법 및 그 장치 | |
| WO2021017332A1 (zh) | 语音控制报错方法、电器及计算机可读存储介质 | |
| WO2014115952A1 (ko) | 유머 발화를 이용하는 음성 대화 시스템 및 그 방법 | |
| WO2019223547A1 (zh) | 教学分析方法、服务器及计算机可读存储介质 | |
| WO2020233055A1 (zh) | 基于活体检测的产品推广方法、装置、设备及存储介质 | |
| WO2018014593A1 (zh) | 基于大数据的风险预测方法、装置、服务器及存储介质 | |
| WO2019223543A1 (zh) | 教学分析方法、服务器及计算机可读存储介质 | |
| WO2021051557A1 (zh) | 基于语义识别的关键词确定方法、装置和存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19933824 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19933824 Country of ref document: EP Kind code of ref document: A1 |
|
| 32PN | Ep: public notification in the ep bulletin as address of the adressee cannot be established |
Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 30.03.2022) |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19933824 Country of ref document: EP Kind code of ref document: A1 |