WO2024178590A1 - 一种生成虚拟形象的方法和装置 - Google Patents

一种生成虚拟形象的方法和装置 Download PDF

Info

Publication number
WO2024178590A1
WO2024178590A1 PCT/CN2023/078643 CN2023078643W WO2024178590A1 WO 2024178590 A1 WO2024178590 A1 WO 2024178590A1 CN 2023078643 W CN2023078643 W CN 2023078643W WO 2024178590 A1 WO2024178590 A1 WO 2024178590A1
Authority
WO
WIPO (PCT)
Prior art keywords
information
user
virtual image
text content
characteristic information
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2023/078643
Other languages
English (en)
French (fr)
Inventor
潘倞燊
苏琪
聂为然
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Huawei Technologies Co Ltd
Original Assignee
Huawei Technologies Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Huawei Technologies Co Ltd filed Critical Huawei Technologies Co Ltd
Priority to CN202380089835.6A priority Critical patent/CN120457459A/zh
Priority to PCT/CN2023/078643 priority patent/WO2024178590A1/zh
Publication of WO2024178590A1 publication Critical patent/WO2024178590A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T13/00Animation
    • G06T13/80Two-dimensional [2D] animation, e.g. using sprites

Definitions

  • Embodiments of the present application relate to the field of human-computer interaction, and more specifically, to a method and device for generating a virtual image.
  • Virtual images are becoming more and more popular because they can provide users with diversified styles and intelligent interactive services.
  • the terminal device can provide users with a variety of style parameters for users to customize. Since there are many style parameters and there is no correlation between the parameters, users need to have a certain knowledge base on the setting of virtual images in order to set up a virtual image that meets their expectations. This will increase the user's learning cost and affect the user's setting efficiency and usage experience.
  • the present application provides a method and device for generating a virtual image, which helps to reduce the learning cost of users when setting up a virtual image, improves the efficiency of setting up the virtual image, and also helps to improve the user experience.
  • the method provided in the present application can be applied to mobile phones, tablet computers, wearable devices, augmented reality (AR)/virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA), vehicles and other terminal devices.
  • AR augmented reality
  • VR virtual reality
  • laptops laptops
  • ultra-mobile personal computers UMPC
  • netbooks netbooks
  • PDA personal digital assistants
  • the embodiments of the present application do not impose any restrictions on the specific types of terminal devices.
  • the vehicle is a vehicle in a broad sense, which can be a means of transportation (such as commercial vehicles, passenger cars, motorcycles, flying cars, trains, etc.), industrial vehicles (such as forklifts, trailers, tractors, etc.), engineering vehicles (such as excavators, bulldozers, cranes, etc.), agricultural equipment (such as mowers, harvesters, etc.), amusement equipment, toy vehicles, etc.
  • transportation such as commercial vehicles, passenger cars, motorcycles, flying cars, trains, etc.
  • industrial vehicles such as forklifts, trailers, tractors, etc.
  • engineering vehicles such as excavators, bulldozers, cranes, etc.
  • agricultural equipment such as mowers, harvesters, etc.
  • amusement equipment toy vehicles, etc.
  • toy vehicles etc.
  • the embodiments of the present application do not specifically limit the types of vehicles.
  • a method for generating a virtual image comprising: obtaining first text content; determining first personality characteristic information based on the first text content and a user's historical information record; and generating a virtual image corresponding to the first text content based on the first personality characteristic information.
  • the terminal device can generate a virtual image based on the text content input by the user and the user's historical information record.
  • a virtual image that better meets the user's expectations can be generated in combination with the user's historical information record, which also makes it easier for the user to resonate with the virtual image.
  • the user does not need to choose from a large number of unrelated style parameters, which helps to reduce the user's learning cost when setting the virtual image, improves the efficiency of setting the virtual image, and also helps to improve the user's experience.
  • the first text content includes a person's name or nickname.
  • the user's historical information records include the user's search records for news in news applications, the use of efficiency applications (such as work plan applications, speech-to-text applications, etc.), the search records for movies or TV series in video applications, the launch of games, etc., in the past period of time (for example, one month).
  • efficiency applications such as work plan applications, speech-to-text applications, etc.
  • the search records for movies or TV series in video applications the launch of games, etc., in the past period of time (for example, one month).
  • the user's historical information record also includes the duration and frequency of use of different applications. For example, a user who frequently uses news applications may prefer a serious format text; a user who frequently uses video applications may prefer a novel format text; a user who frequently uses efficiency applications may prefer a concise and brief format text.
  • a correspondence between personality characteristic information and a virtual image is stored in a terminal device, and a virtual image corresponding to the first text content is generated according to the first personality characteristic information, including: generating the virtual image according to the first personality characteristic and the correspondence.
  • determining the first personality characteristic information based on the first text content and the user's historical information record includes: when the first text content is included in the historical information record, determining, based on the first text content, text co-occurrence words corresponding to the first text content; and determining the first personality characteristic information based on the text co-occurrence words.
  • the first personality feature information can be determined by the text co-occurrence words corresponding to the first text content.
  • the personality feature information of the virtual image can be more in line with the user's expectations, and the final generated virtual image can be more in line with the user's expectations, which helps to improve the user's experience.
  • determining the first personality characteristic information based on the co-occurring words in the text includes: determining the first personality characteristic information corresponding to the first text content based on text similarities between the co-occurring words in the text and words related to different personality dimensions.
  • determining the first personality characteristic information based on the first text content and the user's historical information record includes: when the first text content is not included in the historical information record, determining the similarity between the first text content and each of one or more names in the historical information record; and determining the first personality characteristic information based on the similarity of each name and the personality characteristic information corresponding to each name.
  • the personality feature information corresponding to the first text content can be obtained by calculating the similarity with each name in the historical information record.
  • the personality feature information of the virtual image can be made more in line with the user's expectations, and the final generated virtual image can also be made more in line with the user's expectations, which helps to improve the user's experience.
  • the method further includes: when it is detected that the user issues a voice command or a text input command, controlling the virtual image to make a voice reply or a text reply through a second text content, and the second text content is determined by the historical information record.
  • generating a virtual image corresponding to the first text content based on the first personality characteristic information includes: prompting a user with multiple personality characteristic information based on the first text content and the historical information record; and determining the first personality characteristic information from the multiple personality characteristic information based on user input of the multiple personality characteristic information.
  • generating a virtual image corresponding to the first text content according to the first personality characteristic information includes: determining first style characteristic information according to the first personality characteristic information; and generating the virtual image according to the first style characteristic information.
  • a correspondence between personality feature information and style feature information is stored in a terminal device, and determining the first style feature information according to the first personality feature information includes: determining the first style feature information according to the first personality feature information and the correspondence.
  • the style feature information required for generating a virtual image can be quickly obtained and the corresponding virtual image can be generated.
  • the user does not need to select from a large number of unrelated style parameters, which helps to reduce the learning cost of the user when setting the virtual image, improves the efficiency of setting the virtual image, and also helps to improve the user's experience.
  • determining first style feature information based on the first personality feature information includes: inputting the first personality feature information into a prediction model to obtain the first style feature information; wherein the prediction model is trained by sample training data, the sample training data includes sample names, sample personality feature information, and sample style feature information, and the sample style feature information includes at least one of an avatar, a body shape, a timbre, a speaking speed, a tone, and an expression corresponding to the sample name.
  • the prediction model is trained through sample names, sample personality feature information and sample style feature information corresponding to the sample names, so that the prediction model can include different style feature dimensions with names as alignment labels, which helps to ensure the personality consistency of various style parameters of the virtual image, and can also make the final virtual image more in line with the user's expectations, which helps to improve the user experience.
  • a virtual image corresponding to the first text content is generated based on the first style feature information, including: prompting the user with style feature information of one or more dimensions; and generating the virtual image based on the user's input of the style feature information of the one or more dimensions.
  • the user can participate in the generation process of the virtual image, and the virtual image can be more in line with the user's expectations, which helps to improve the user's usage experience.
  • determining the first style feature information according to the first personality feature information includes: determining the style feature information of the one or more dimensions according to the first personality feature information.
  • the style feature information of one or more dimensions may be a style parameter list of one or more dimensions.
  • the one or more dimensions of style feature information include at least one of avatar information, shape information, timbre information, speaking speed information, intonation information, and expression information.
  • the avatar information may be an avatar list consisting of multiple avatars
  • the body information may be a body list consisting of multiple bodies
  • the timbre information may be a timbre list consisting of multiple timbres
  • the speaking rate information may be a speaking rate list consisting of multiple speaking rates
  • the tone information may be a tone list consisting of multiple tones
  • the expression information may be an expression list consisting of multiple expressions.
  • generating a virtual image corresponding to the first text content according to the first personality characteristic information includes: prompting a user with multiple virtual images according to the first personality characteristic information; According to the user's input to the multiple virtual images, a virtual image corresponding to the first text content is determined from the multiple virtual images.
  • the user can participate in the generation process of the virtual image, and the virtual image can be more in line with the user's expectations, which helps to improve the user's usage experience.
  • a device for generating a virtual image comprising: an acquisition unit for acquiring a first text content; a determination unit for determining first personality characteristic information based on the first text content and a user's historical information record; and a virtual image generation unit for generating a virtual image corresponding to the first text content based on the first personality characteristic information.
  • the determination unit is used to: when the first text content is included in the historical information record, determine, based on the first text content, text co-occurrence words corresponding to the first text content; and determine the first personality characteristic information based on the text co-occurrence words.
  • the determination unit is used to: when the first text content is not included in the historical information record, determine the similarity between the first text content and each of the one or more names in the historical information record; and determine the first personality characteristic information based on the similarity of each name and the personality characteristic information corresponding to each name.
  • the device also includes a detection unit and a control unit, the detection unit being used to detect that a user issues a voice command or a text input command; the control unit being used to control the virtual image to make a voice reply or a text reply through a second text content, and the second text content is determined by the historical information record.
  • the device also includes: a first prompting unit, used to prompt a user with multiple personality characteristic information based on the first text content and the historical information record; and the determining unit, used to determine the first personality characteristic information from the multiple personality characteristic information based on the user's input of the multiple personality characteristic information.
  • the determination unit is further used to determine first style feature information based on the first personality feature information; and the virtual image generation unit is used to generate the virtual image based on the first style feature information.
  • the determination unit is used to: input the first personality characteristic information into a prediction model to obtain the first style characteristic information; wherein the prediction model is trained by sample training data, the sample training data includes sample names, sample personality characteristic information, and sample style characteristic information, and the sample style characteristic information includes at least one of an avatar, shape, timbre, speaking speed, intonation, and expression corresponding to the sample name.
  • the device also includes: a second prompting unit, used to prompt the user with style feature information of one or more dimensions; and the virtual image generation unit, used to generate the virtual image based on the user's input of the style feature information of the one or more dimensions.
  • the one or more dimensions of style feature information include at least one of avatar information, shape information, timbre information, speaking speed information, intonation information, and expression information.
  • the device also includes: a third prompting unit, used to prompt a plurality of virtual images to the user based on the first personality characteristic information; and the virtual image generating unit, used to determine the virtual image corresponding to the first text content from the plurality of virtual images based on the user's input to the plurality of virtual images.
  • the present application provides a device for generating a virtual image, the device comprising a processing unit and a storage unit, wherein the storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit so that the device executes any possible method in the first aspect.
  • the present application provides a terminal device, which includes any possible device in the second aspect, or includes the device described in the third aspect.
  • the terminal device is a vehicle, a computer, or a mobile phone.
  • the present application provides a computer program product, comprising: a computer program code, when the computer program code is run on a computer, the computer executes any possible method in the first aspect above.
  • the above-mentioned computer program code can be stored in whole or in part on the first storage medium, wherein the first storage medium can be packaged together with the processor or separately packaged with the processor, and the embodiments of the present application do not specifically limit this.
  • the present application provides a computer-readable medium storing a program code, and when the computer program code runs on a computer, the computer executes any possible method in the first aspect.
  • the present application provides a chip, comprising a circuit, wherein the circuit is used to execute any possible method in the first aspect above.
  • FIG1 is a functional block diagram of a terminal device provided in an embodiment of the present application.
  • FIG. 2 is a schematic diagram of the distribution of display screens in a vehicle cabin provided in an embodiment of the present application.
  • FIG3 is a schematic flowchart of a method for generating a virtual image provided in an embodiment of the present application.
  • FIG. 4 is a graphical user interface GUI provided in an embodiment of the present application.
  • FIG. 5 is another set of GUIs provided by an embodiment of the present application.
  • FIG6 is a schematic flow chart of a method for generating a personality-style database provided in an embodiment of the present application.
  • FIG. 7 is a schematic diagram of using names as alignment labels of different style feature dimensions provided in an embodiment of the present application.
  • FIG. 8 is another set of GUIs provided by an embodiment of the present application.
  • FIG. 9 is another set of GUIs provided by an embodiment of the present application.
  • FIG. 10 is another GUI provided by an embodiment of the present application.
  • FIG. 11 is another schematic flowchart of the method for generating a virtual image provided in an embodiment of the present application.
  • FIG. 12 is a schematic flowchart of an apparatus for generating a virtual image provided in an embodiment of the present application.
  • prefixes such as “first” and “second” are used only to distinguish different description objects, and have no limiting effect on the position, order, priority, quantity or content of the described objects.
  • the use of prefixes such as ordinal numbers to distinguish description objects in the embodiments of the present application does not constitute a limitation on the described objects.
  • the meaning of "multiple" is two or more.
  • virtual images are becoming more and more popular because they can provide users with diversified styles and intelligent interactive services.
  • the terminal device can provide users with a variety of style parameters for users to customize. The more full the virtual image is, the more its style parameters increase dramatically. Since there are many style parameters and there is no correlation between the parameters, users need to have a certain knowledge base on the setting of virtual images in order to set up a virtual image that meets their expectations, which will increase the user's learning cost and affect the user's setting efficiency and user experience.
  • the embodiment of the present application provides a method and device for generating a virtual image, which generates a virtual image that better meets the user's expectations by combining the user's historical information records, so that the user is more likely to resonate with the virtual image.
  • the user does not need to choose from a large number of unrelated style parameters, which helps to reduce the user's learning cost when setting the virtual image, improves the efficiency of setting the virtual image, and also helps to improve the user's experience.
  • FIG1 is a functional block diagram of a terminal device 100 provided in an embodiment of the present application.
  • the terminal device 100 may include a text input unit 110, a display device 120, and a computing platform 130, wherein the text input unit 110 is used to obtain text content input by a user.
  • the text input unit 110 may be used to obtain a nickname (or name) of a virtual image input by a user.
  • the computing platform 130 may include one or more processors, such as processors 131 to 13n (n is a positive integer).
  • the processor is a circuit with signal processing capabilities.
  • the processor may be a circuit with instruction reading and execution capabilities, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP); in another implementation, the processor can implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable, such as a processor that is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as a field programmable gate array (FPGA).
  • ASIC application-specific integrated circuit
  • PLD programmable logic device
  • the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units.
  • the processor can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc.
  • the computing platform 130 can also include a memory, the memory is used to store instructions, and some or all of the processors 131 to 13n can call the instructions in the memory and execute the instructions to achieve the corresponding functions.
  • the computing platform 130 can obtain the text input unit
  • the user receives the nickname (or name) of the virtual image sent by 110, and generates the virtual image according to the nickname (or name) of the virtual image and the historical information record of the user.
  • the display device 120 in the vehicle cockpit is mainly divided into two categories.
  • the first category is the vehicle display screen;
  • the second category is a projection display screen, such as a head-up display (HUD).
  • the vehicle display screen is a physical display screen and an important part of the vehicle infotainment system.
  • There can be multiple display screens in the cockpit such as a digital instrument display screen, a central control screen, a display screen in front of the passenger in the co-pilot seat (also called the front passenger), a display screen in front of the left rear passenger, and a display screen in front of the right rear passenger.
  • Even the window can be used as a display screen for display.
  • Head-up display also known as a head-up display system.
  • HUD includes, for example, a combiner-HUD (C-HUD) system, a windshield-HUD (W-HUD) system, and an augmented reality HUD (AR-HUD) system.
  • C-HUD combiner-HUD
  • W-HUD windshield-HUD
  • AR-HUD augmented reality HUD
  • the computing platform 130 when it generates a virtual image, it can display the virtual image to the user through the display device 120.
  • the above display device 130 is described by taking a vehicle-mounted display screen and a projection display screen as examples, and the embodiments of the present application are not limited thereto.
  • the display device 130 can also be a light display screen or a projection screen.
  • Fig. 2 shows a schematic diagram of an exemplary display screen distribution in a vehicle cabin provided by an embodiment of the present application.
  • the vehicle cabin may include a display screen 201 (or, it may also be referred to as a central control screen), a display screen 202 (or, it may also be referred to as a co-pilot entertainment screen), a display screen 203 (or, it may also be referred to as an entertainment screen in the left area of the second row), a display screen 204 (or, it may also be referred to as an entertainment screen in the right area of the second row) and an instrument screen.
  • a display screen 201 or, it may also be referred to as a central control screen
  • a display screen 202 or, it may also be referred to as a co-pilot entertainment screen
  • a display screen 203 or, it may also be referred to as an entertainment screen in the left area of the second row
  • a display screen 204 or, it may also be referred to as an entertainment screen in the right area of the second row
  • FIG3 shows a schematic flow chart of a method 300 for generating a virtual image provided by an embodiment of the present application.
  • the method 300 can be executed by the terminal device 100, or the method 300 can be executed by the computing platform 130, or the method 300 can be executed by a system on a chip (SoC) in the computing platform 130, or the method 300 can be executed by a processor in the computing platform 130.
  • SoC system on a chip
  • the following is an example of the method 300 being executed by a terminal device.
  • the method 300 includes:
  • the first text content may be a name, given name or nickname of the virtual image.
  • the first text content may be the above-mentioned "Name A”.
  • FIG4 shows a graphical user interface (GUI) provided in an embodiment of the present application.
  • the GUI is a virtual image generation interface displayed by a display screen 201, and the virtual image generation interface includes a text input box 401, a virtual image generation control 402, and a cancel control 403.
  • the terminal device can obtain the first text content "name A”.
  • FIG4 is an example of a GUI displayed on a display screen 201 in a vehicle, but the present application is not limited thereto.
  • the GUI may also be a display interface for generating a game character displayed on a mobile phone or computer, or may also be a display interface of an interactive robot.
  • S320 Determine first personality characteristic information according to the first text content and the user's historical information record.
  • the historical information record of the user includes the record of the user using the terminal device in the past period of time.
  • the user's historical information records may include one or more of the following: records of the user searching for news in news applications, records of using efficiency applications (for example, applications for specifying work plans, voice-to-text applications, etc.), records of searching for movies or TV series in video applications, records of launching game applications, records of reading novels or jokes, records of searching in browsers, and records of listening to songs in music applications over the past period of time (for example, one month).
  • efficiency applications for example, applications for specifying work plans, voice-to-text applications, etc.
  • records of searching for movies or TV series in video applications records of launching game applications, records of reading novels or jokes, records of searching in browsers, and records of listening to songs in music applications over the past period of time (for example, one month).
  • the user's historical information record may include the usage duration and frequency of different applications. For example, a user who frequently uses news applications may prefer a serious format text; a user who frequently uses video applications may prefer a novel format text; a user who frequently uses efficiency applications may prefer a concise and brief format text.
  • the user's historical information record may be stored locally on the terminal device, or may be stored in other devices (e.g., a cloud server). If the user's historical information record is stored in the cloud server, when the first text content is obtained, the terminal device may request the user's historical information record from the cloud server. Thus, the terminal device may determine the first personality feature information based on the first text content and the historical information record.
  • the first text content is "Name A”
  • the user's historical information records include a record of launching a game application and using a game character named "Name A” in the game application.
  • the terminal device can determine the personality characteristic information of the game character named "Name A" as the first personality characteristic information corresponding to the first text content.
  • determining the first personality characteristic information based on the first text content and the user's historical information record includes: when the first text content is included in the historical information record, determining text co-occurrence words corresponding to the first text content based on the first text content; and determining the first personality characteristic information based on the text co-occurrence words.
  • the terminal device can search the user's historical information record. For example, the terminal device can determine that the user has searched for "Mulan” twice through the browser in the past period of time and clicked on the search results related to the literary image in the search history. For example, the information in the search results is as follows:
  • determining the first personality characteristic information according to the text co-occurring words includes: determining the first personality characteristic information corresponding to the first text content according to text similarities between the text co-occurring words and words related to different personality dimensions.
  • the embedding layer of the bidirectional encoder representation from transformers (BERT) model can be used to extract word vectors for different words, or the word2Vec bag-of-words model can be used to extract word vectors.
  • the following is an example of extracting word vectors using the embedding layer.
  • V A (a 1 , a 2 ..., a i , ..., a N )
  • V A represents the mathematical representation of word A in the word vector V A space
  • V B represents the mathematical representation of word B in the word vector V B space.
  • the similarity between word A and labeled word B can be calculated as shown in formula (1):
  • Similarity AB is the similarity between word A and calibrated word B.
  • n is the number of similarities calculated between the co-occurring words and the labeled words
  • VnameA is the personality characteristic parameter of the name A
  • VnameA can be used as the characteristic information of the name A in the personality dimension.
  • Table 1 shows a Big Five personality dimension table.
  • the calibrated words related to different personality dimensions may include “extrovert”, “talkative”, “quiet”, “shy”, “introvert”, “confident”, “dominant”, etc.
  • determining the first personality characteristic information based on the first text content and the user's historical information record includes: when the first text content is not included in the historical information record, determining the similarity between the first text content and each of one or more names in the historical information record; and determining the first personality characteristic information based on the similarity of each name and the personality characteristic information corresponding to each name.
  • the terminal device cannot extract the text co-occurrence related to the first text content. At this time, the terminal device can determine the first personality feature information in combination with the names of people similar to the first text content in the historical information record.
  • the first personality characteristic information corresponding to the first text content can be determined by the following formula (3):
  • V the first text content is the first personality characteristic information
  • Similarity i is the similarity between the first text content and each person's name in the historical information record
  • Vi is the personality characteristic parameter of each person's name
  • N is the number of names in the historical information record.
  • the first text content is "Hua Tielan", and "Hua Tielan” has not appeared in the historical information record, but includes names similar to "Hua Tielan", such as “Hua Mulan”, “Hua Rong”, “Hua Wuque” and “Temuzin”.
  • the terminal device can calculate the similarity between "Hua Tielan” and the names that have appeared in the historical information record.
  • the similarity of the names can be calculated by cosine similarity using the Embedding word vector, and the similarity is used as the weight.
  • Table 2 shows the calculated similarity between "Hua Tielan” and the names that have appeared in the historical information record.
  • VHuaTieLan is the personality characteristic parameter of "HuaTieLan” ( VHuaTieLan can be the personality characteristic information corresponding to "HuaTieLan”)
  • VHuaMuLan is the personality characteristic parameter of "HuaMuLan”
  • VHuaRong is the personality characteristic parameter of "HuaRong”
  • VHuaWuQue is the personality characteristic parameter of "HuaWuQue” personality characteristic parameters
  • V Temujin is the personality characteristic parameter of “Temujin”
  • Similarity 1 is the similarity between “Hua Tielan” and “Hua Mulan” (for example, 0.4)
  • Similarity 2 is the similarity between “Hua Tielan” and “Hua Rong” (for example, 0.2)
  • Similarity 3 is the similarity between “Hua Tielan” and “Hua Wuque” (for example, 0.2)
  • Similarity 4 is the similarity between “Hua Tielan” and “Te
  • V Hua Mulan , V Hua Rong , V Hua Wu Que and V Temujin can refer to the calculation process of the personality characteristic parameter of name A in the above formula (2), which will not be repeated here.
  • the above personality characteristic parameters may be a specific representation of the above personality characteristic information, and the embodiments of the present application are not limited thereto.
  • the personality characteristic information may also be text content describing personality characteristics (e.g., "extrovert”, “talkative”, “introvert”, etc.).
  • S330 Generate a virtual image corresponding to the first text content according to the first personality characteristic information.
  • the terminal device may store a correspondence between personality characteristic information and a virtual image, and generate a virtual image corresponding to the first text content according to the first personality characteristic information, including: generating the virtual image according to the personality characteristic information and the correspondence.
  • Table 3 shows a correspondence between personality feature information and a virtual image.
  • the first text content is "Name A”.
  • “Name A” Through “Name A” and the user's historical information record, it is determined that the personality characteristic information corresponding to Name A is extroversion and confidence. According to the corresponding relationship shown in Table 3 above, a virtual image 1 corresponding to "Name A" can be generated.
  • determining the first personality characteristic information based on the first text content and the user's historical information records includes: prompting the user with multiple personality characteristic information based on the first text content and the historical information records; and determining the first personality characteristic information from the multiple personality characteristic information based on the user's input of the multiple personality characteristic information.
  • FIG. 5 shows another GUI provided by an embodiment of the present application.
  • the personality characteristic information corresponding to the "name A” can be determined based on the "name A” and the user's historical information record, wherein the historical information record indicates that the user has started Game 1 and used the character corresponding to the "name A", started the browser and browsed the search records related to the "name A", and searched for songs about the "name A” through the music application in the past week.
  • the vehicle can display multiple personality characteristic information corresponding to the "name A” (for example, personality characteristic information 501-503) and the control corresponding to each personality characteristic information through the GUI based on the above historical information records.
  • the personality characteristic information 501 corresponding to the "name A” in Game 1 is “extroverted, talkative and firm and confident”
  • the personality characteristic information 502 corresponding to the "name A” in the browser's search record is "modest and magnanimous”
  • the personality characteristic information 503 corresponding to the "name A” in the music application is "rude and suspicious”.
  • the above user inputs the multiple personality characteristics information by clicking the control on the display screen as an example.
  • the embodiments of the present application are not limited thereto.
  • the user after seeing the personality characteristic information 501-503, the user can also issue a voice command "extrovert, talkative, and assertive".
  • the vehicle After the vehicle obtains the voice command, it can determine that the personality characteristic information selected by the user is "extrovert, talkative, and assertive" according to the user's voice command, and thus can generate a corresponding virtual image according to the personality characteristic information.
  • generating a virtual image corresponding to the first text content according to the first personality characteristic information includes: determining first style characteristic information according to the first personality characteristic information; and generating the virtual image according to the first style characteristic information.
  • the terminal device stores a correspondence between personality feature information and style feature information, and determining the first style feature information according to the first personality feature information includes: determining the first style feature information according to the first personality feature information and the correspondence.
  • the style feature information may be classified into different categories.
  • Table 4 shows a classification method of the style feature information.
  • Table 5 shows a correspondence between personality feature information and style feature information.
  • the style characteristic information of the virtual image can be determined to be facial features 1, body shape 1, clothing 1, timbre 1, speaking speed 1, intonation 1, rhythm 1, response content 1, expression 1, gesture 1 and posture 1 based on the corresponding relationship shown in Table 5 above, so that a corresponding virtual image can be generated based on the style characteristic information.
  • determining the first style characteristic information based on the first personality characteristic information includes: inputting the first personality characteristic information into a prediction model to obtain the first style characteristic information; wherein the prediction model is trained by sample training data, the sample training data includes a sample name, sample personality characteristic information and sample style characteristic information, and the sample style characteristic information includes at least one of an avatar, a body shape, a timbre, a speaking speed, a tone and an expression corresponding to the sample name.
  • the above embodiment introduces the process of obtaining the personality characteristic parameters of a person's name through the text co-occurrence words of the name.
  • the style features of other dimensions associated with the name can be used to label the personality characteristics of the name.
  • the word vector of "Name A” and the personality characteristic parameters corresponding to the name are concatenated as input, and the shape and sound characteristics of "Name A" in "Movie 1" are used as labels to train the prediction model.
  • data of multiple style characteristic dimensions with personality characteristic parameters as labels can be obtained. This type of data of multiple style characteristic dimensions can be used to obtain different style templates based on the style transfer method and store them in the prediction model.
  • the above prediction model can also be understood as a personality-style database.
  • FIG6 shows a schematic flow chart of a method 600 for generating a personality-style database provided in an embodiment of the present application.
  • the method 600 may be executed by a device (e.g., a cloud server) including a model training device.
  • the method 600 includes:
  • the personality characteristic parameter of "name A” can be calculated according to the above formula (2).
  • the personality characteristic parameters of "name A” appearing in different categories of historical information records may be different.
  • Table 6 shows the correspondence between the text co-occurrence words and personality characteristic parameters of "name A" appearing in different categories of historical information records.
  • FIG. 7 shows an example of an alignment method using names as alignment marks for different style feature dimensions provided by an embodiment of the present application.
  • the avatar features of person A in “Movie 1” include avatar 1 and avatar 2
  • the avatar features in “Game 1” include avatar 3 and avatar 4.
  • the voice features of person A in “Movie 1” include voice feature 1, and the voice features in “Game 1” are voice feature 2.
  • An association relationship can be established between person A, personality feature parameter 1, avatar 1, avatar 2, and voice feature 1, and an association relationship can be established between person A, personality feature parameter 2, avatar 3, avatar 4, and voice feature 2.
  • generating a virtual image corresponding to the first text content according to the first style feature information includes: prompting the user with style feature information of one or more dimensions; and generating the virtual image according to the user's input of the style feature information of the one or more dimensions.
  • style feature information of one or more dimensions is a style parameter list of one or more dimensions.
  • FIG. 8 shows another set of GUIs provided by an embodiment of the present application.
  • the vehicle can determine the personality feature information corresponding to the “name A” according to the “name A” and the user’s historical information record, wherein the user’s historical information record indicates that the user frequently opens game 1 and uses the character corresponding to the “name A” in the past week.
  • the vehicle can determine the style parameter list of one or more dimensions corresponding to the “name A” according to the personality feature information corresponding to the “name A”, wherein the parameter list of the one or more dimensions includes an avatar list, a timbre list, and a tone list.
  • the avatar list includes avatar 3 and avatar 4.
  • the tone list includes voice 1 and voice 2, wherein the timbre of voice 1 is a timbre corresponding to the game character “name A”, and the timbre of voice 2 is another timbre corresponding to the game character “name A”.
  • the tone list includes voice 3 and voice 4, wherein the tone of voice 3 is a tone corresponding to the game character “name A”, and the tone of voice 4 is another tone corresponding to the game character “name A”. Users can select their favorite style features from a list of style parameters in different dimensions.
  • the vehicle when it is detected that the user has selected avatar 3, voice 1 and voice 3 and clicked the control 801, the vehicle can generate a virtual image 1 based on the avatar 3, the timbre corresponding to voice 1 and the pitch corresponding to voice 3, and display the virtual image 1 through the display screen.
  • the avatar of the virtual image 1 is avatar 3 and the timbre and pitch of the voice signal "Hello, I am your virtual image Xiao A" emitted by the virtual image 1 match the timbre in the above-mentioned voice 1 and the pitch in voice 3 respectively.
  • FIG8 is an example of allowing a user to manually select style parameters of different dimensions, but the embodiments of the present application are not limited thereto.
  • the terminal device can automatically select the style parameter with the highest recommendation degree from the style parameter list of different dimensions as the style feature parameter of the virtual image, thereby automatically generating the corresponding virtual image.
  • the terminal device can automatically randomly select a style parameter from the style parameter list of different dimensions as the style feature parameter of the virtual image, thereby automatically generating the corresponding virtual image. In this way, the process of allowing the user to select the style parameter of the virtual image is avoided, which helps to improve the user experience.
  • the method 300 further includes: when it is detected that the user issues a voice command or a text input command, controlling the virtual image to make a voice reply or a text reply through a second text content, and the second text content is determined by the historical information record.
  • FIG. 9 shows another set of GUIs provided by an embodiment of the present application.
  • the user can perform voice interaction with the virtual image 1.
  • the vehicle can detect the voice command “open the window” issued by the user.
  • the vehicle can control the virtual image 1
  • the user is responded to by voice through the lines in "Game 1".
  • the lines of the game character "Name A" in "Game 1" include “Let me show you advanced operations”.
  • the vehicle can control the virtual image to send a voice signal "Let me show you advanced operations” to respond to the user's voice command.
  • using lines to replace the conventional response method during the conversation can give the user a more appropriate virtual image experience.
  • generating a virtual image corresponding to the first text content according to the first personality characteristic information includes: prompting a plurality of virtual images to the user according to the first personality characteristic information; and determining the virtual image corresponding to the first text content from the plurality of virtual images according to the user's input to the plurality of virtual images.
  • prompting a user with multiple virtual images based on the first personality characteristic information includes: determining first style characteristic information based on the first personality characteristic information; and prompting the user with multiple virtual images based on the first style characteristic information.
  • FIG10 shows another GUI provided by an embodiment of the present application.
  • the vehicle can control the central control screen to display the virtual image 1001 and the virtual image 1002 corresponding to "name A" according to the personality characteristic information corresponding to "name A”; or the vehicle can determine the style characteristic information corresponding to "name A” according to the personality characteristic information corresponding to "name A” and control the central control screen to display the virtual image 1001 and the virtual image 1002 corresponding to "name A” according to the style characteristic information.
  • the central control screen can determine the style characteristic information corresponding to "name A” according to the personality characteristic information corresponding to "name A” and control the central control screen to display the virtual image 1001 and the virtual image 1002 corresponding to "name A” according to the style characteristic information.
  • the user can also issue a voice command "select the virtual image on the far left". After the vehicle detects the voice command issued by the user, it can determine that the virtual image corresponding to the "name A" that the user wants to generate is virtual image 1001 according to the voice command.
  • FIG11 shows a schematic flow chart of a method 1100 for generating a virtual image provided by an embodiment of the present application.
  • the method 1100 may be executed by the terminal device 100, or the method 1100 may be executed by the computing platform 130, or the method 1100 may be executed by the SoC in the computing platform 130, or the method 1100 may be executed by the processor in the computing platform 130.
  • the following description is made by taking the method 1100 executed by the terminal device as an example.
  • the method 1100 includes:
  • the terminal device when it is detected that the user inputs “name A” in the text input box 401 and clicks on the operation of generating a virtual image control 402 , the terminal device can obtain the information of “name A”.
  • determining the personality characteristic parameters corresponding to the name based on the name and the user's historical information record includes: when the name is included in the user's historical information record, determining the personality characteristic parameters corresponding to the name based on text co-occurrence words of the name.
  • the personality characteristic parameter can be calculated by the above formula (2).
  • the personality characteristic parameter can be represented in the form of a vector.
  • determining personality characteristic parameters corresponding to the name based on the name and the user's historical information record includes: when the name is not included in the user's historical information record, determining the similarity between the name and each of one or more names in the historical information record; and determining the personality characteristic parameters of the name based on the similarity of each name and the personality characteristic parameters corresponding to each name.
  • the personality characteristic parameter may be input into the above prediction model (or personality-style database) to obtain the style characteristic parameter corresponding to the name.
  • the terminal device can generate a virtual image based on the name input by the user and the user's historical information record. In this way, a virtual image that better meets the user's expectations can be generated based on the user's historical information record, and the user can resonate with the virtual image, which helps to improve the user's experience.
  • Figure 12 shows a schematic block diagram of a device 1200 for generating a virtual image provided by an embodiment of the present application.
  • the device 1200 includes: an acquisition unit 1210, used to acquire a first text content; a determination unit 1220, used to determine first personality feature information based on the first text content and the user's historical information record; and a virtual image generation unit 1230, used to generate a virtual image corresponding to the first text content based on the first personality feature information.
  • the determination unit 1220 is used to: when the first text content is included in the historical information record, determine, based on the first text content, text co-occurrence words corresponding to the first text content; and determine the first personality characteristic information based on the text co-occurrence words.
  • the determination unit 1220 is used to: when the first text content is not included in the historical information record, determine the similarity between the first text content and each of the one or more names in the historical information record; and determine the first personality characteristic information based on the similarity of each name and the personality characteristic information corresponding to each name.
  • the device 1200 also includes a detection unit and a control unit, the detection unit is used to detect that the user issues a voice command or a text input command; the control unit is used to control the virtual image to make a voice reply or a text reply through a second text content, and the second text content is determined by the historical information record.
  • the detection unit is used to detect that the user issues a voice command or a text input command
  • the control unit is used to control the virtual image to make a voice reply or a text reply through a second text content, and the second text content is determined by the historical information record.
  • the device 1200 also includes: a first prompting unit, used to prompt the user with multiple personality characteristic information based on the first text content and the historical information record; the determining unit 1220, used to determine the first personality characteristic information from the multiple personality characteristic information based on the user's input of the multiple personality characteristic information.
  • a first prompting unit used to prompt the user with multiple personality characteristic information based on the first text content and the historical information record
  • the determining unit 1220 used to determine the first personality characteristic information from the multiple personality characteristic information based on the user's input of the multiple personality characteristic information.
  • the determination unit 1220 is further configured to determine first style characteristic information according to the first personality characteristic information; and the virtual image generation unit is configured to generate the virtual image according to the first style characteristic information.
  • the determination unit 1220 is used to: input the first personality characteristic information into a prediction model to obtain the first style characteristic information; wherein the prediction model is trained by sample training data, the sample training data includes a sample name, sample personality characteristic information and sample style characteristic information, and the sample style characteristic information includes at least one of an avatar, shape, timbre, speaking speed, intonation and expression corresponding to the sample name.
  • the device 1200 also includes: a second prompting unit, used to prompt the user with the style feature information of the one or more dimensions; and the virtual image generating unit 1230, used to generate the virtual image according to the user's input of the style feature information of the one or more dimensions.
  • a second prompting unit used to prompt the user with the style feature information of the one or more dimensions
  • the virtual image generating unit 1230 used to generate the virtual image according to the user's input of the style feature information of the one or more dimensions.
  • the one or more dimensions of style feature information include at least one of avatar information, shape information, timbre information, speaking speed information, intonation information, and expression information.
  • the device 1200 also includes: a third prompting unit, used to prompt a plurality of virtual images to the user according to the first personality characteristic information; and the virtual image generating unit 1230, used to determine a virtual image corresponding to the first text content from the plurality of virtual images according to the user's input to the plurality of virtual images.
  • a third prompting unit used to prompt a plurality of virtual images to the user according to the first personality characteristic information
  • the virtual image generating unit 1230 used to determine a virtual image corresponding to the first text content from the plurality of virtual images according to the user's input to the plurality of virtual images.
  • the acquisition unit 1210 may be the computing platform in FIG. 1 or a processing circuit or a processing circuit in the computing platform. Taking the acquisition unit 1210 as the processor 131 in the computing platform as an example, the processor 131 can acquire the first text content input by the user.
  • the determination unit 1220 is the computing platform in Figure 1 or a processing circuit, processor or controller in the computing platform. Taking the determination unit 1220 as the processor 132 in the computing platform as an example, the processor 132 can determine the first personality feature information based on the first text content obtained by the processor 131 and the historical information record of the user.
  • the virtual image generation unit 1230 is the computing platform in Figure 1 or a processing circuit, processor or controller in the computing platform. Taking the virtual image generation unit 1230 as the processor 133 in the computing platform as an example, the processor 133 can generate a virtual image corresponding to the first text content according to the first personality feature information determined by the processor 132.
  • the functions implemented by the above-mentioned acquisition unit 1210, the functions implemented by the determination unit 1220 and the functions implemented by the virtual image generation unit 1230 can be implemented by different processors, or, can be implemented by the same processor, or, part of the functions can be implemented by the same processor, and the embodiments of the present application are not limited to this.
  • the division of the units in the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated.
  • the units in the device can be implemented in the form of a processor calling software; for example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory.
  • the processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of the units of the device, wherein the processor is, for example, a general-purpose processor, such as a CPU or a microprocessor, and the memory is a memory in the device or a memory outside the device.
  • the units in the device can be implemented in the form of hardware circuits, and the functions of some or all of the units can be realized by designing the hardware circuits.
  • the hardware circuit can be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are realized by designing the logical relationship of the components in the circuit; for example, in another implementation, the hardware circuit can be realized by PLD.
  • FPGA as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured through the configuration file, so as to realize the functions of some or all of the above units. All units of the above device may be implemented entirely in the form of a processor calling software, or entirely in the form of a hardware circuit, or partially in the form of a processor calling software and the rest in the form of a hardware circuit.
  • Each unit in the above device may be one or more processors (or processing circuits) configured to implement the above method, such as a CPU, a GPU, an NPU, a TPU, a DPU, a microprocessor, a DSP, an ASIC, an FPGA, or a combination of at least two of these processor forms.
  • processors or processing circuits configured to implement the above method, such as a CPU, a GPU, an NPU, a TPU, a DPU, a microprocessor, a DSP, an ASIC, an FPGA, or a combination of at least two of these processor forms.
  • the SoC may include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device.
  • the type of the at least one processor may be different, for example, including CPU and FPGA, CPU and artificial intelligence processor, CPU and GPU, etc.
  • An embodiment of the present application also provides a device, which includes a processing unit and a storage unit, wherein the storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit so that the device executes the method or steps executed by the above embodiment.
  • the processing unit may be the processor 131 - 13n shown in FIG. 1 .
  • An embodiment of the present application further provides a system, which includes a computing platform and a display device, wherein the computing platform may include the above-mentioned device 1200.
  • the display device may be a display screen, such as the display screens 201 - 204 described above.
  • the embodiment of the present application further provides a terminal device, which may include the above-mentioned device 1200, or, Including the above system.
  • the terminal device may be a mobile phone, a computer or a vehicle.
  • the embodiment of the present application further provides a computer program product, which includes: a computer program code, and when the computer program code is executed on a computer, the computer executes the method in the above embodiment.
  • the embodiment of the present application further provides a computer-readable medium, wherein the computer-readable medium stores a program code.
  • the computer program code runs on a computer, the computer executes the method in the above embodiment.
  • An embodiment of the present application further provides a chip, wherein the chip includes a circuit, and the circuit is used to execute the method in the above embodiment.
  • each step of the above method can be completed by an integrated logic circuit of hardware in a processor or an instruction in the form of software.
  • the method disclosed in conjunction with the embodiment of the present application can be directly embodied as a hardware processor for execution, or a combination of hardware and software modules in a processor for execution.
  • the software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or a power-on erasable programmable memory, a register, etc.
  • the storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it is not described in detail here.
  • the memory may include a read-only memory and a random access memory, and provide instructions and data to the processor.
  • the size of the serial numbers of the above-mentioned processes does not mean the order of execution.
  • the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
  • the disclosed systems, devices and methods can be implemented in other ways.
  • the device embodiments described above are only schematic.
  • the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
  • Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
  • the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
  • each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
  • the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
  • the part that makes technical contribution or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.
  • the aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

Landscapes

  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

本申请提供了一种生成虚拟形象的方法和装置,该方法可以应用于人机交互领域。该方法包括:获取用户输入的人名;根据该人名和用户的历史信息记录,确定该人名对应的人格特征信息;根据该人格特征信息,生成该人名对应的虚拟形象。本申请实施例可以应用于包括智能汽车或者电动汽车在内的终端设备中,通过结合用户输入的人名和用户的历史信息记录生成更符合用户期望的虚拟形象,使得用户更易对虚拟形象产生共鸣,有助于提升用户的使用体验。

Description

一种生成虚拟形象的方法和装置 技术领域
本申请实施例涉及人机交互领域,并且更具体地,涉及一种生成虚拟形象的方法和装置。
背景技术
由于可以为用户提供多元化风格以及智能化交互服务,虚拟形象越来越受到人们的青睐。当前在设置虚拟形象时,终端设备可以为用户提供多种风格参数供用户自定义选择。由于风格参数众多且参数之间并无关联关系,导致用户需要对虚拟形象的设置有一定的知识基础才能设置得到符合用户期望的虚拟形象,这样会提高用户的学习成本,且影响用户的设置效率和使用体验。
发明内容
本申请提供了一种生成虚拟形象的方法和装置,有助于降低用户在设置虚拟形象时的学习成本,提升虚拟形象的设置效率,也有助于提升用户的使用体验。
本申请提供的方法可以应用于手机、平板电脑、可穿戴设备、增强现实(augmented reality,AR)/虚拟现实(virtual reality,VR)设备、笔记本电脑、超级移动个人计算机(ultra-mobile personal computer,UMPC)、上网本、个人数字助理(personal digital assistant,PDA)、车辆等终端设备上,本申请实施例对终端设备的具体类型不作任何限制。例如,该车辆为广义概念上的车辆,可以是交通工具(如商用车、乘用车、摩托车、飞行车、火车等),工业车辆(如:叉车、挂车、牵引车等),工程车辆(如挖掘机、推土车、吊车等),农用设备(如割草机、收割机等),游乐设备,玩具车辆等,本申请实施例对车辆的类型不作具体限定。
第一方面,提供了一种生成虚拟形象的方法,该方法包括:获取第一文本内容;根据该第一文本内容和用户的历史信息记录,确定第一人格特征信息;根据该第一人格特征信息,生成该第一文本内容对应的虚拟形象。
基于上述技术方案,终端设备可以根据用户输入的文本内容和用户的历史信息记录,生成虚拟形象。这样,可以结合用户的历史信息记录生成更符合用户期望的虚拟形象,也使得用户更易对虚拟形象产生共鸣。同时,无需用户在参数众多且无关联关系的风格参数中进行选择,有助于降低用户在设置虚拟形象时的学习成本,提升虚拟形象的设置效率,也有助于提升用户的使用体验。
在一些可能的实现方式中,该第一文本内容包括人名或者昵称。
在一些可能的实现方式中,该用户的历史信息记录包括用户在过去一段时间(例如,一个月)内,在新闻类应用中搜索新闻的记录、使用效率类应用(例如,制定工作计划的应用、语音转文字类应用等)的记录、在视频应用中搜索电影或者电视剧的记录、启动游 戏应用的记录、阅读小说或者笑话的记录、在浏览器中的搜索记录、在音乐应用中收听歌曲的记录中的一项或者多项。
在一些可能的实现方式中,该用户的历史信息记录还包括不同应用的使用时长和频率。例如,使用新闻类应用频率比较高的用户可能偏好于严正的格式文本;使用视频类应用频率比较高的用户可能偏好新奇的格式文本;使用效率类应用频率比较高的用户可能偏好精炼简短的格式文本。
在一些可能的实现方式中,终端设备中保存有人格特征信息与虚拟形象的对应关系,根据该第一人格特征信息,生成该第一文本内容对应的虚拟形象,包括:根据该第一人格特征和该对应关系,生成该虚拟形象。
结合第一方面,在第一方面的某些实现方式中,根据该第一文本内容和用户的历史信息记录,确定该第一人格特征信息,包括:在该历史信息记录中包括该第一文本内容时,根据该第一文本内容,确定该第一文本内容对应的文本共现词;根据该文本共现词,确定该第一人格特征信息。
基于上述技术方案,在用户的历史信息记录中包括该第一文本内容时,可以通过第一文本内容对应的文本共现词,确定该第一人格特征信息。这样,通过提取第一文本内容上下文的文本共现词来确定第一人格特征信息,可以使得虚拟形象的人格特征信息更符合用户的期望,也使得最终生成的虚拟形象更符合用户的期望,有助于提升用户的使用体验。
在一些可能的实现方式中,该根据该文本共现词,确定该第一人格特征信息,包括:根据该文本共现词和不同人格维度相关词的文本相似度,确定该第一文本内容对应的第一人格特征信息。
结合第一方面,在第一方面的某些实现方式中,该根据该第一文本内容和用户的历史信息记录,确定该第一人格特征信息,包括:在该历史信息记录中不包括该第一文本内容时,确定该第一文本内容与该历史信息记录中一个或者多个人名中每个人名的相似度;根据该每个人名的相似度以及该每个人名对应的人格特征信息,确定该第一人格特征信息。
基于上述技术方案,在历史信息记录中不包括第一文本内容的情况下,通过和历史信息记录中每个人名的相似度的计算,可以得到该第一文本内容对应的人格特征信息。这样,通过结合历史信息记录中与该第一文本内容相似的人名,可以使得虚拟形象的人格特征信息更符合用户的期望,也使得最终生成的虚拟形象更符合用户的期望,有助于提升用户的使用体验。
结合第一方面,在第一方面的某些实现方式中,该方法还包括:在检测到用户发出语音指令或文本输入指令时,控制该虚拟形象通过第二文本内容进行语音回复或者文本回复,该第二文本内容由该历史信息记录确定。
基于上述技术方案,在与虚拟形象对话或者文本交互的过程中,通过控制虚拟形象使用历史信息记录中的内容来替换常规的应答方式,可以给用户更贴切的虚拟形象体验。
结合第一方面,在第一方面的某些实现方式中,根据该第一人格特征信息,生成该第一文本内容对应的虚拟形象,包括:根据该第一文本内容和该历史信息记录,向用户提示多个人格特征信息;根据用户对该多个人格特征信息的输入,从该多个人格特征信息中确定该第一人格特征信息。
基于上述技术方案,通过向用户提示多个人格特征信息,可以使得用户参与到虚拟形 象的生成过程中,也使得虚拟形象更符合用户的预期,有助于提升用户的使用体验。
结合第一方面,在第一方面的某些实现方式中,根据该第一人格特征信息,生成该第一文本内容对应的虚拟形象,包括:根据该第一人格特征信息,确定第一风格特征信息;根据该第一风格特征信息,生成该虚拟形象。
在一些可能的实现方式中,终端设备中保存有人格特征信息与风格特征信息的对应关系,根据该第一人格特征信息,确定第一风格特征信息,包括:根据该第一人格特征信息和该对应关系,确定该第一风格特征信息。
基于上述技术方案,通过在终端设备中保存人格特征信息和风格特征信息的对应关系,在确定人格特征信息后可以快速获取到生成虚拟形象所需的风格特征信息并生成相应的虚拟形象。这样,无需用户在参数众多且无关联关系的风格参数中进行选择,有助于降低用户在设置虚拟形象时的学习成本,提升虚拟形象的设置效率,也有助于提升用户的使用体验。
结合第一方面,在第一方面的某些实现方式中,根据该第一人格特征信息,确定第一风格特征信息,包括:将该第一人格特征信息输入预测模型中,得到该第一风格特征信息;其中,该预测模型由样本训练数据训练得到,该样本训练数据包括样本人名、样本人格特征信息以及样本风格特征信息,该样本风格特征信息包括与该样本人名对应的头像、形体、音色、语速、语调、表情中的至少一项。
基于上述技术方案,通过样本人名、样本人格特征信息以及与该样本人名对应的样本风格特征信息对预测模型进行训练,可以使得预测模型中包括以人名为对齐标签的不同风格特征维度,有助于保证虚拟形象的多种风格参数的人格一致性,也可以使得最终得到的虚拟形象更符合用户的预期,有助于提升用户的使用体验。
结合第一方面,在第一方面的某些实现方式中,根据该第一风格特征信息,生成该第一文本内容对应的虚拟形象,包括:向用户提示一个或者多个维度的风格特征信息;根据用户对该一个或者多个维度的风格特征信息的输入,生成该虚拟形象。
基于上述技术方案,通过向用户提示一个或者多个维度的风格特征信息,可以使得用户参与到虚拟形象的生成过程中,也使得虚拟形象更符合用户的预期,有助于提升用户的使用体验。
在一些可能的实现方式中,根据该第一人格特征信息,确定第一风格特征信息,包括:根据该第一人格特征信息,确定该一个或者多个维度的风格特征信息。
在一些可能的实现方式中,该一个或者多个维度的风格特征信息可以为一个或者多个维度的风格参数列表。
结合第一方面,在第一方面的某些实现方式中,该一个或者多个维度的风格特征信息包括头像信息、形体信息、音色信息、语速信息、语调信息、表情信息中的至少一项。
例如,该头像信息可以为由多个头像组成的头像列表,该形体信息可以为由多个形体组成的形体列表,该音色信息可以为由多个音色组成的音色列表,该语速信息可以为由多个语速组成的语速列表,该语调信息可以为由多个语调组成的语调列表,该表情信息可以为由多个表情组成的表情列表。
结合第一方面,在第一方面的某些实现方式中,根据该第一人格特征信息,生成该第一文本内容对应的虚拟形象,包括:根据该第一人格特征信息,向用户提示多个虚拟形象; 根据用户对该多个虚拟形象的输入,从该多个虚拟形象中确定该第一文本内容对应的虚拟形象。
基于上述技术方案,通过提示用户从多个虚拟形象中进行选择,可以使得用户参与到虚拟形象的生成过程中,也使得虚拟形象更符合用户的预期,有助于提升用户的使用体验。
第二方面,提供了一种生成虚拟形象的装置,该装置包括:获取单元,用于获取第一文本内容;确定单元,用于根据该第一文本内容和用户的历史信息记录,确定第一人格特征信息;虚拟形象生成单元,用于根据该第一人格特征信息,生成该第一文本内容对应的虚拟形象。
结合第二方面,在第二方面的某些实现方式中,该确定单元,用于:在该历史信息记录中包括该第一文本内容时,根据该第一文本内容,确定该第一文本内容对应的文本共现词;根据该文本共现词,确定该第一人格特征信息。
结合第二方面,在第二方面的某些实现方式中,该确定单元,用于:在该历史信息记录中不包括该第一文本内容时,确定该第一文本内容与该历史信息记录中一个或者多个人名中每个人名的相似度;根据该每个人名的相似度以及该每个人名对应的人格特征信息,确定该第一人格特征信息。
结合第二方面,在第二方面的某些实现方式中,该装置还包括检测单元和控制单元,该检测单元,用于检测到用户发出语音指令或文本输入指令;该控制单元,用于控制该虚拟形象通过第二文本内容进行语音回复或者文本回复,该第二文本内容由所述历史信息记录确定。
结合第二方面,在第二方面的某些实现方式中,该装置还包括:第一提示单元,用于根据该第一文本内容和该历史信息记录,向用户提示多个人格特征信息;该确定单元,用于根据用户对该多个人格特征信息的输入,从该多个人格特征信息中确定该第一人格特征信息。
结合第二方面,在第二方面的某些实现方式中,该确定单元,还用于根据该第一人格特征信息,确定第一风格特征信息;该虚拟形象生成单元,用于根据该第一风格特征信息,生成该虚拟形象。
结合第二方面,在第二方面的某些实现方式中,该确定单元,用于:将该第一人格特征信息输入预测模型中,得到该第一风格特征信息;其中,该预测模型由样本训练数据训练得到,该样本训练数据包括样本人名、样本人格特征信息以及样本风格特征信息,该样本风格特征信息包括与该样本人名对应的头像、形体、音色、语速、语调、表情中的至少一项。
结合第二方面,在第二方面的某些实现方式中,该装置还包括:第二提示单元,用于向用户提示一个或者多个维度的风格特征信息;该虚拟形象生成单元,用于根据用户对该一个或者多个维度的风格特征信息的输入,生成该虚拟形象。
结合第二方面,在第二方面的某些实现方式中,该一个或者多个维度的风格特征信息包括头像信息、形体信息、音色信息、语速信息、语调信息、表情信息中的至少一项。
结合第二方面,在第二方面的某些实现方式中,该装置还包括:第三提示单元,用于根据该第一人格特征信息,向用户提示多个虚拟形象;该虚拟形象生成单元,用于根据用户对该多个虚拟形象的输入,从该多个虚拟形象中确定该第一文本内容对应的虚拟形象。
第三方面,本申请提供了一种生成虚拟形象的装置,该装置包括处理单元和存储单元,其中存储单元用于存储指令,处理单元执行存储单元所存储的指令,以使该装置执行第一方面中任一种可能的方法。
第四方面,本申请提供了一种终端设备,该终端设备包括第二方面中任一种可能的装置,或者,包括第三方面所述的装置。
结合第四方面,在第四方面的某些实现方式中,该终端设备为车辆、电脑或者手机。
第五方面,本申请提供了一种计算机程序产品,所述计算机程序产品包括:计算机程序代码,当所述计算机程序代码在计算机上运行时,使得计算机执行上述第一方面中任一种可能的方法。
需要说明的是,上述计算机程序代码可以全部或者部分存储在第一存储介质上,其中第一存储介质可以与处理器封装在一起的,也可以与处理器单独封装,本申请实施例对此不作具体限定。
第六方面,本申请提供了一种计算机可读介质,所述计算机可读介质存储有程序代码,当所述计算机程序代码在计算机上运行时,使得计算机执行上述第一方面中任一种可能的方法。
第七方面,本申请提供了一种芯片,该芯片包括电路,该电路用于执行上述第一方面中任一种可能的方法。
附图说明
图1是本申请实施例提供的终端设备的功能框图示意。
图2是本申请实施例提供的车辆座舱内显示屏分布的示意图。
图3是本申请实施例提供的生成虚拟形象的方法的示意性流程图。
图4是本申请实施例提供的图像用户界面GUI。
图5是本申请实施例提供的另一组GUI。
图6是本申请实施例提供的生成人格-风格数据库的方法的示意性流程图。
图7是本申请实施例提供的利用人名作为不同风格特征维度的对齐标签的示意图。
图8是本申请实施例提供的另一组GUI。
图9是本申请实施例提供的另一组GUI。
图10是本申请实施例提供的另一GUI。
图11是本申请实施例提供的生成虚拟形象的方法的另一示意性流程图。
图12是本申请实施例提供的生成虚拟形象的装置的示意性流程图。
具体实施方式
下面将结合本申请实施例中的附图,对本申请实施例中的技术方案进行描述。其中,在本申请实施例的描述中,除非另有说明,“/”表示或的意思,例如,A/B可以表示A或B;本文中的“和/或”仅仅是一种描述关联对象的关联关系,表示可以存在三种关系,例如,A和/或B,可以表示:单独存在A,同时存在A和B,单独存在B这三种情况。“至少一项”是指一项或一项以上。例如,“A和B中的至少一项”,类似于“A和/或B”,描述关联对象的关联关系,表示可以存在三种关系,例如,A和B中的至少一项,可以表示:单独存在 A,同时存在A和B,单独存在B这三种情况。
本申请实施例中采用诸如“第一”、“第二”的前缀词,仅仅为了区分不同的描述对象,对被描述对象的位置、顺序、优先级、数量或内容等没有限定作用。本申请实施例中对序数词等用于区分描述对象的前缀词的使用不对所描述对象构成限制,对所描述对象的陈述参见权利要求或实施例中上下文的描述,不应因为使用这种前缀词而构成多余的限制。此外,在本实施例的描述中,除非另有说明,“多个”的含义是两个或两个以上。
如前所述,由于可以为用户提供多元化风格以及智能化交互服务,虚拟形象越来越受到人们的青睐。当前在设置虚拟形象时,终端设备可以为用户提供多种风格参数供用户自定义选择。虚拟形象越丰满的同时,其风格参数也急剧增加。由于风格参数众多且参数之间并无关联关系,导致用户需要对虚拟形象的设置有一定的知识基础才能设置得到符合用户期望的虚拟形象,这样会提高用户的学习成本,且影响用户的设置效率和使用体验。
当前还可以通过文本生成图像的方式来设置虚拟形象。但是由于认知不同,用户对于同一个名字理解不同。例如对于10岁小学生群体,人名A可能更多的是游戏角色;而对于50岁中老年人群体,人名A可能更多的来自文学戏剧形象。如果一个50岁的中老年用户在终端设备上输入虚拟形象的昵称“人名A”,终端设备可能会生成类似于游戏角色中“人名A”的虚拟形象,这样会和用户的预期有较大的偏差,从而会影响用户的使用体验。
本申请实施例提供了一种生成虚拟形象的方法和装置,通过结合用户的历史信息记录生成更符合用户期望的虚拟形象,使得用户更易对虚拟形象产生共鸣。同时,无需用户在参数众多且无关联关系的风格参数中进行选择,有助于降低用户在设置虚拟形象时的学习成本,提升虚拟形象的设置效率,也有助于提升用户的使用体验。
图1是本申请实施例提供的终端设备100的一个功能框图示意。终端设备100可以包括文本输入单元110、显示装置120和计算平台130,其中,文本输入单元110用于获取用户输入的文本内容。本申请实施例中,文本输入单元110可以用于获取用户输入的虚拟形象的昵称(或者人名)。
计算平台130可包括一个或多个处理器,例如处理器131至13n(n为正整数),处理器是一种具有信号的处理能力的电路,在一种实现中,处理器可以是具有指令读取与运行能力的电路,例如中央处理单元(central processing unit,CPU)、微处理器、图形处理器(graphics processing unit,GPU)(可以理解为一种微处理器)、或数字信号处理器(digital signal processor,DSP)等;在另一种实现中,处理器可以通过硬件电路的逻辑关系实现一定功能,该硬件电路的逻辑关系是固定的或可以重构的,例如处理器为专用集成电路(application-specific integrated circuit,ASIC)或可编程逻辑器件(programmable logic device,PLD)实现的硬件电路,例如现场可编程门阵列(field programmable gate array,FPGA)。在可重构的硬件电路中,处理器加载配置文档,实现硬件电路配置的过程,可以理解为处理器加载指令,以实现以上部分或全部单元的功能的过程。此外,处理器还可以是针对人工智能设计的硬件电路,其可以理解为一种ASIC,例如神经网络处理单元(neural network processing unit,NPU)、张量处理单元(tensor processing unit,TPU)、深度学习处理单元(deep learning processing unit,DPU)等。此外,计算平台130还可以包括存储器,存储器用于存储指令,处理器131至13n中的部分或全部处理器可以调用存储器中的指令,执行指令,以实现相应的功能。本申请实施例中,计算平台130可以获取文本输入单元 110发送的虚拟形象的昵称(或者人名),并根据虚拟形象的昵称(或者人名)和用户的历史信息记录,生成虚拟形象。
以终端设备100是车辆为例,车辆座舱内的显示装置120主要分为两类,第一类是车载显示屏;第二类是投影显示屏,例如抬头显示装置(head up display,HUD)。车载显示屏是一种物理显示屏,是车载信息娱乐系统的重要组成部分,座舱内可以设置有多块显示屏,如数字仪表显示屏,中控屏,副驾驶位上的乘客(也称为前排乘客)面前的显示屏,左侧后排乘客面前的显示屏以及右侧后排乘客面前的显示屏,甚至是车窗也可以作为显示屏进行显示。抬头显示,也称平视显示系统。主要用于在驾驶员前方的显示设备(例如挡风玻璃)上显示例如时速、导航等驾驶信息。以降低驾驶员视线转移时间,避免因驾驶员视线转移而导致的瞳孔变化,提升行驶安全性和舒适性。HUD例如包括组合型抬头显示(combiner-HUD,C-HUD)系统、风挡型抬头显示(windshield-HUD,W-HUD)系统、增强现实型抬头显示系统(augmented reality HUD,AR-HUD)。应理解,HUD也可以随着技术演进出现其他类型的系统,本申请对此不作限定。本申请实施例中,计算平台130在生成虚拟形象时,可以通过显示装置120向用户展示该虚拟形象。
以上显示装置130是以车载显示屏和投影显示屏为例进行说明的,本申请实施例并不限于此。例如,该显示装置130还可以为光显示屏或者投影幕布。
图2示出了本申请实施例提供的一种示例性车辆座舱内显示屏分布的示意图。如图2所示,车辆的座舱内可以包括显示屏201(或者,也可以称为中控屏)、显示屏202(或者,也可以称为副驾娱乐屏)、显示屏203(或者,也可以称为二排左侧区域的娱乐屏)、显示屏204(或者,也可以称为二排右侧区域的娱乐屏)以及仪表屏。
图3示出了本申请实施例提供的生成虚拟形象的方法300的示意性流程图。该方法300可以由上述终端设备100执行,或者,该方法300可以由上述计算平台130执行,或者,该方法300可以由上述计算平台130中的片上系统(chip on system,SoC)执行,或者,该方法300可以由上述计算平台130中的处理器执行。以下以方法300由终端设备执行为例进行说明。如图3所示,该方法300包括:
S310,获取第一文本内容。
示例性的,该第一文本内容可以为虚拟形象的人名、名字或者昵称。如该第一文本内容可以为上述“人名A”。
以该终端设备是车辆为例,图4示出了本申请实施例提供的图像用户界面(graphical user interface,GUI)。该GUI为通过显示屏201显示的虚拟形象生成界面,该虚拟形象生成界面上包括文本输入框401、生成虚拟形象控件402以及取消控件403。在检测到用户在文本输入框401中输入“人名A”且点击生成虚拟形象控件402的操作时,终端设备可以获取到第一文本内容“人名A”。
图4是以该GUI为车辆中的显示屏201显示的GUI为例进行说明的,本申请实施例并不限于此。例如,该GUI也可以为手机或者电脑上显示的生成游戏角色的显示界面,或者,还可以是交互机器人的显示界面。
S320,根据该第一文本内容和用户的历史信息记录,确定第一人格特征信息。
一个实施例中,该用户的历史信息记录包括用户在过去一段时间内使用终端设备时的记录。
示例性的,该用户的历史信息记录可以包括用户在过去一段时间(例如,一个月)内,在新闻类应用中搜索新闻的记录、使用效率类应用(例如,指定工作计划的应用、语音转文字类应用等)的记录、在视频应用中搜索电影或者电视剧的记录、启动游戏应用的记录、阅读小说或者笑话的记录、在浏览器中的搜索记录、在音乐应用中收听歌曲的记录中的一项或者多项。
示例性的,该用户的历史信息记录可以包括不同应用的使用时长和频率。例如,使用新闻类应用频率比较高的用户可能偏好于严正的格式文本;使用视频类应用频率比较高的用户可能偏好新奇的格式文本;使用效率类应用频率比较高的用户可能偏好精炼简短的格式文本。
一个实施例中,该用户的历史信息记录可以是保存在终端设备本地的,或者,也可以是保存在其他设备(例如,云端服务器)中。若该用户的历史信息记录保存在云端服务器中,在获取到该第一文本内容时,终端设备可以向云端服务器请求该用户的历史信息记录。从而终端设备可以根据该第一文本内容和该历史信息记录,确定该第一人格特征信息。
示例性的,该第一文本内容为“人名A”,用户的历史信息记录包括启动游戏应用并使用游戏应用中的游戏角色名称为“人名A”的记录,终端设备可以将该游戏角色名称为“人名A”的人格特征信息确定为第一文本内容对应的第一人格特征信息。
可选地,根据该第一文本内容和用户的历史信息记录,确定该第一人格特征信息,包括:在该历史信息记录中包括该第一文本内容时,根据该第一文本内容,确定该第一文本内容对应的文本共现词;根据该文本共现词,确定该第一人格特征信息。
示例性的,以该第一文本内容为“花木兰”为例,在获取到该第一文本内容时,终端设备可以对用户的历史信息记录进行搜索。例如,终端设备可以确定用户在过去的一段时间内通过浏览器搜索了2次“木兰”并在搜索记录中点击了与文学形象相关的搜索结果。例如,搜索结果中的信息如下:
“木兰既是奇女子又是普通人,既是帼国英雄又是平民少女,既是矫健的勇士又是娇美的女儿。她勤劳善良又坚毅勇敢,淳厚质朴又机敏活泼,热爱亲人又报效祖国,不慕高的官厚禄而热爱和平生活”。
通过命名实体分析方法或者词条分析方法,可以将与人名“木兰”共现的相关形容词“奇”、“普通”、“矫健的”、“娇美的”、“勤劳善良”、“坚毅勇敢”、“醇厚质朴”和“机敏活泼”等提取出来,这些词可以作为与人名“木兰”对应的文本共现词。
可选地,根据该文本共现词,确定该第一人格特征信息,包括:根据该文本共现词和不同人格维度相关词的文本相似度,确定该第一文本内容对应的第一人格特征信息。
示例性的,可以使用来自变形的双向编码表示(bidirectional encoder representation from transformers,BERT)模型的嵌入(embedding)层提取不同词汇的词向量,或者,也可以使用word2Vec词袋模型提取词向量。下面以嵌入层提取词向量为例进行说明。
若人名A的文本共现词中包括词A且词A的词向量为VA=(a1,a2…,ai,…,aN),VA表示词A在词向量VA空间内的数理表示;不同人格维度相关词中包括标定词B且标定词B的词向量为VB=(b1,b2…,bi,…,bN),VB表示词B在词向量VB空间内的数理表示。词A和词B的含义越相似,在词向量空间的向量方向越相近。示例性的,词A和标定词B的相似度的计算方式可以如公式(1)所示:
其中,SimilarityAB为词A和标定词B的相似度。
同样的,还可以计算词A和其他人格维度相关词中每个词的相似度,以及文本共现词中除词A以外的其他文本共现词与不同人格维度相关词的相似度。该人名A的人格特征参数的计算方式可以如公式(2)所示:
其中,n为共现词与标定词计算的相似度的个数,V人名A为该人名A的人格特征参数,或者,V人名A可以作为该人名A在该人格维度的特征信息。
示例性的,上述不同人格维度相关词可以取自大五人格维度表。如表1示出了一种大五人格维度表。
表1大五人格维度表

示例性的,不同人格维度相关标定词可以包括“外向”、“健谈”、“安静”、“害羞”、“内向”、“坚定自信”、“主导地位”等。
可选地,根据该第一文本内容和用户的历史信息记录,确定该第一人格特征信息,包括:在该历史信息记录中不包括该第一文本内容时,确定该第一文本内容与该历史信息记录中一个或者多个人名中每个人名的相似度;根据该每个人名的相似度以及该每个人名对应的人格特征信息,确定该第一人格特征信息。
在该历史信息记录中不包括该第一文本内容时,终端设备无法提取与该第一文本内容相关的文本共现。此时终端设备可以结合该历史信息记录中与该第一文本内容相似的人名,确定该第一人格特征信息。
一个实施例中,在该历史信息记录中不包括该第一文本内容时,可以通过如下公式(3)确定该第一文本内容对应的该第一人格特征信息:
其中,V第一文本内容为该第一人格特征信息,Similarityi为第一文本内容与历史信息记录中每个人名的相似度,Vi为每个人名的人格特征参数,N为历史信息记录中人名的个数。
示例性的,该第一文本内容为“花铁兰”,该历史信息记录中未出现过“花铁兰”,而是包括与“花铁兰”相似的人名“花木兰”、“花荣”、“花无缺”和“铁木真”等。终端设备可以分别计算“花铁兰”与历史信息记录中出现过的人名的相似度。人名相似度可以用Embedding词向量计算余弦相似度,并以相似度作为权重。示例性的,表2示出了计算得到的“花铁兰”与历史信息记录中出现过的人名的相似度。
表2
示例性的,可以通过如下公式(4)计算“花铁兰”的人格特征参数:
V花铁兰=Similarity1*V花木兰+Similarity2*V花荣+Similarity3*V花无缺+Similarity4*V铁木真  (4)
其中,V花铁兰为“花铁兰”的人格特征参数(V花铁兰可以为“花铁兰”对应的人格特征信息),V花木兰为“花木兰”的人格特征参数,V花荣为“花荣”的人格特征参数,V花无缺为“花无缺”的人 格特征参数,V铁木真为“铁木真”的人格特征参数,Similarity1为“花铁兰”与“花木兰”相似度(例如,0.4),Similarity2为“花铁兰”与“花荣”相似度(例如,0.2),Similarity3为“花铁兰”与“花无缺”相似度(例如,0.2),Similarity4为“花铁兰”与“铁木真”相似度(例如,0.2)。
以上V花木兰、V花荣、V花无缺、V铁木真的计算过程可以参考上述公式(2)中人名A的人格特征参数的计算过程,此处不再赘述。
以上人格特征参数可以为上述人格特征信息的一种具体的表征方式,本申请实施例并不限于此。例如,该人格特征信息还可以为描述人格特征的文本内容(例如,“外向”、“健谈”、“内向”等)。
S330,根据该第一人格特征信息,生成该第一文本内容对应的虚拟形象。
一个实施例中,终端设备可以保存有人格特征信息与虚拟形象的对应关系,根据该第一人格特征信息,生成该第一文本内容对应的虚拟形象,包括:根据该人格特征信息和该对应关系,生成该虚拟形象。
示例性的,表3示出了一种人格特征信息与虚拟形象的对应关系。
表3
示例性的,该第一文本内容为“人名A”。通过“人名A”和用户的历史信息记录确定出该人名A对应的人格特征信息为外向和自信,可以根据上述表3所示的对应关系,生成该“人名A”对应的虚拟形象1。
可选地,根据该第一文本内容和用户的历史信息记录,确定第一人格特征信息,包括:根据该第一文本内容和该历史信息记录,向用户提示多个人格特征信息;根据用户对该多个人格特征信息的输入,从该多个人格特征信息中确定该第一人格特征信息。
图5示出了本申请实施例提供的另一GUI。
如图5中的(a)所示,车辆在检测到用户点击控件402的操作时,可以根据“人名A”和用户的历史信息记录,确定该“人名A”对应的人格特征信息,其中该历史信息记录指示用户在过去一周内启动游戏1且使用“人名A”对应的角色、启动浏览器且浏览与“人名A”相关的搜索记录、通过音乐应用搜索关于“人名A”的歌曲。车辆可以根据上述历史信息记录,通过GUI显示“人名A”对应的多个人格特征信息(例如,人格特征信息501-503)以及每个人格特征信息对应的控件。例如,游戏1中“人名A”对应的人格特征信息501为“外向、健谈和坚定自信”,浏览器的搜索记录中“人名A”对应的人格特征信息502为“谦逊礼让、宽宏大量”,音乐应用中“人名A”对应的人格特征信息503为“粗鲁、怀疑”。
如图5中的(b)所示,在检测到用户点击人格特征信息501对应的控件504且点击确定控件505的操作时,车辆可以控制显示屏显示生成的虚拟形象并控制该虚拟形象发出的语音信号“你好,我是你的虚拟形象小A”。
以上用户对该多个人格特征信息的输入是以通过点击显示屏上的控件为例进行说明 的,本申请实施例并不限于此。例如,用户在看到人格特征信息501-503后,还可以发出语音指令“外向、健谈和坚定自信”。车辆在获取到该语音指令后,可以根据用户的语音指令确定用户选择的人格特征信息为“外向、健谈和坚定自信”,从而可以根据该人格特征信息生成对应的虚拟形象。
可选地,该根据该第一人格特征信息,生成该第一文本内容对应的虚拟形象,包括:根据该第一人格特征信息,确定第一风格特征信息;根据该第一风格特征信息,生成该虚拟形象。
可选地,该终端设中保存有人格特征信息与风格特征信息的对应关系,该根据该第一人格特征信息,确定第一风格特征信息,包括:根据该第一人格特征信息和该对应关系,确定该第一风格特征信息。
示例性的,风格特征信息可以分为不同的类别,表4示出了一种风格特征信息的分类方式。
表4
示例性的,表5示出了一种人格特征信息与风格特征信息的对应关系。
表5

例如,在根据第一文本内容和用户的历史信息记录确定该第一人格特征信息为外向且自信时,可以根据上述表5所示的对应关系,确定虚拟形象的风格特征信息为五官1、身材1、服装1、音色1、语速1、语调1、韵律1、应答内容1、表情1、手势1且姿态1,从而可以根据该风格特征信息生成对应的虚拟形象。
可选地,根据该第一人格特征信息,确定第一风格特征信息,包括:将该第一人格特征信息输入预测模型中,得到该第一风格特征信息;其中,该预测模型由样本训练数据训练得到,该样本训练数据包括样本人名、样本人格特征信息以及样本风格特征信息,该样本风格特征信息包括与该样本人名对应的头像、形体、音色、语速、语调、表情中的至少一项。
上述实施例中介绍了通过人名的文本共现词得到该人名的人格特征参数的过程。同样的,可以利用与该人名关联的其他维度的风格特征,进行该人名的人格特征标注。例如,将“人名A”的词向量和其人名对应的人格特征参数拼接起来作为输入,以“人名A”在《电影1》中的形体和声音特征作为标签,进行预测模型的训练。最终可以得到人格特征参数作为标签的多个风格特征维度的数据。可以利用该类多个风格特征维度的数据,基于风格迁移的方法得到不同的风格模板,将其存入预测模型中。
以上预测模型也可以理解为人格-风格数据库。
图6示出了本申请实施例提供的生成人格-风格数据库的方法600的示意性流程图。该方法600可以由包含模型训练装置的设备(例如,云端服务器)执行。该方法600包括:
S610,确定人名A的人格特征参数。
示例性的,在用户的历史信息记录中出现“人名A”时,可以根据上述公式(2)计算“人名A”的人格特征参数。不同类别的历史信息记录中出现的“人名A”的人格特征参数可能不同。示例性的,表6示出了不同类别的历史信息记录中出现的“人名A”的文本共现词以及人格特征参数的对应关系。
表6
S620,利用人名A作为不同风格特征维度的对齐标签,通过风格迁移的方式生成不同的参数模板,填充到人格-风格数据库中。
示例性的,图7示出了本申请实施例提供的利用人名作为不同风格特征维度的对齐标 签的示意图。如图7所示,人名A在《电影1》中的头像特征包括头像1和头像2且在《游戏1》中的头像特征包括头像3和头像4。人名A在《电影1》中的声音特征包括声音特征1且在《游戏1》中的声音特征为声音特征2。可以建立人名A、人格特征参数1、头像1、头像2和声音特征1之间的关联关系,以及建立人名A、人格特征参数2、头像3、头像4和声音特征2之间的关联关系。
可选地,该根据该第一风格特征信息,生成该第一文本内容对应的虚拟形象,包括:向用户提示一个或者多个维度的风格特征信息;根据用户对该一个或者多个维度的风格特征信息的输入,生成该虚拟形象。
下面结合图8,以该一个或者多个维度的风格特征信息是一个或者多个维度的风格参数列表为例进行说明。
图8示出了本申请实施例提供的另一组GUI。
如图8中的(a)所示,在检测到用户点击控件402的操作时,车辆可以根据“人名A”和用户的历史信息记录,确定该“人名A”对应的人格特征信息,其中该用户的历史信息记录指示用户在过去一周内经常打开游戏1且使用“人名A”对应的角色。车辆可以根据该“人名A”对应的人格特征信息,确定该“人名A”对应的一个或者多个维度的风格参数列表,该一个或者多个维度的参数列表中包括头像列表、音色列表和语调列表。其中,头像列表中包括头像3和头像4。音色列表中包括语音1和语音2,其中语音1的音色为游戏角色“人名A”对应的一种音色,语音2的音色为游戏角色“人名A”对应的另一种音色。语调列表中包括语音3和语音4,其中语音3的音调为游戏角色“人名A”对应的一种音调,语音4的音调为游戏角色“人名A”对应的另一种音调。用户可以从不同维度的风格参数列表中选择自己喜欢的风格特征。
如图8中的(b)所示,在检测到用户选择了头像3、语音1和语音3且点击控件801的操作时,车辆可以根据该头像3、语音1对应的音色和语音3对应的音调生成虚拟形象1并通过显示屏显示该虚拟形象1,该虚拟形象1的头像为头像3且该虚拟形象1发出的语音信号“你好,我是你的虚拟形象小A”的音色和音调分别与上述语音1中音色、语音3中的音调相匹配。
图8中是以让用户手动选择不同维度的风格参数为例进行说明的,本申请实施例并不限于此。例如,终端设备可以自动在不同维度的风格参数列表中选择推荐度最高的风格参数作为虚拟形象的风格特征参数,从而自动生成对应的虚拟形象。又例如,终端设备可以自动在不同维度的风格参数列表中随机选择某个风格参数作为虚拟形象的风格特征参数,从而自动生成对应的虚拟形象。这样,避免了让用户选择虚拟形象的风格参数的过程,有助于提升用户的使用体验。
可选地,该方法300还包括:在检测到用户发出语音指令或文本输入指令时,控制该虚拟形象通过第二文本内容进行语音回复或者文本回复,该第二文本内容由该历史信息记录确定。
图9示出了本申请实施例提供的另一组GUI。
如图9中的(a)所示,以该终端设备是车辆为例,在生成虚拟形象1后,用户可以与该虚拟形象1进行语音交互。例如,车辆可以检测到用户发出的语音指令“打开车窗”。
如图9中的(b)所示,在检测到用户发出的语音指令时,车辆可以控制虚拟形象1 通过《游戏1》中的台词对用户进行语音回复,例如《游戏1》中游戏角色“人名A”的台词包括“让我来给你展示高端操作”。这样,车辆可以控制虚拟形象发出语音信号“让我来给你展示高端操作”,从而对用户的语音指令进行应答。这样,在对话的过程中使用台词来替换常规的应答方式,可以给用户更贴切的虚拟形象体验。
可选地,根据该第一人格特征信息,生成该第一文本内容对应的虚拟形象,包括:根据该第一人格特征信息,向用户提示多个虚拟形象;根据用户对该多个虚拟形象的输入,从该多个虚拟形象中确定该第一文本内容对应的虚拟形象。
可选地,根据该第一人格特征信息,向用户提示多个虚拟形象,包括:根据该第一人格特征信息,确定第一风格特征信息;根据该第一风格特征信息,向用户提示多个虚拟形象。
图10示出了本申请实施例提供的另一GUI。车辆可以根据“人名A”对应的人格特征信息,控制中控屏显示“人名A”对应的虚拟形象1001和虚拟形象1002;或者,车辆可以根据“人名A”对应的人格特征信息,确定“人名A”对应的风格特征信息且根据该风格特征信息控制中控屏显示“人名A”对应的虚拟形象1001和虚拟形象1002。在检测到用户选择虚拟形象1001对应的操作时,可以确定用户希望生成的“人名A”对应的虚拟形象为虚拟形象1001。
以上是以检测到用户点击多个虚拟形象中某个虚拟形象后,从多个虚拟形象中确定该第一文本对应的虚拟形象为例进行说明的,本申请实施例并不限于此。例如,用户还可以发出语音指令“选择最左侧的虚拟形象”。车辆在检测到用户发出的语音指令后,可以根据该语音指令确定用户希望生成的“人名A”对应的虚拟形象为虚拟形象1001。
图11示出了本申请实施例提供的生成虚拟形象的方法1100的示意性流程图。该方法1100可以由上述终端设备100执行,或者,该方法1100可以由上述计算平台130执行,或者,该方法1100可以由上述计算平台130中的SoC执行,或者,该方法1100可以由上述计算平台130中的处理器执行。以下以方法1100由终端设备执行为例进行说明。如图11所示,该方法1100包括:
S1110,获取用户输入的人名。
示例性的,如图4所示,在检测到用户在文本输入框401中输入“人名A”且点击生成虚拟形象控件402的操作时,终端设备可以获取到“人名A”的信息。
S1120,根据该人名和用户的历史信息记录,确定该人名对应的人格特征参数。
可选地,根据该人名和用户的历史信息记录,确定该人名对应的人格特征参数,包括:在用户的历史信息记录中包括该人名时,根据该人名的文本共现词,确定该人名对应的人格特征参数。
示例性的,该人格特征参数可以通过上述公式(2)计算得到。该人格特征参数可以通过向量的形式表征。
可选地,根据该人名和用户的历史信息记录,确定该人名对应的人格特征参数,包括:在用户的历史信息记录中不包括该人名时,确定该人名与该历史信息记录中一个或者多个人名中每个人名的相似度;根据该每个人名的相似度以及该每个人名对应的人格特征参数,确定该人名的人格特征参数。
应理解,S1120的实现过程可以参考上述S320的实现过程,此处不再赘述。
S1130,根据该人格特征参数,确定该人名对应的风格特征参数。
示例性的,可以将该人格特征参数输入到上述预测模型(或者,人格-风格数据库)中,得到该人名对应的风格特征参数。
对于预测模型的描述可以参考上述实施例的描述,此处不再赘述。
S1140,根据该风格特征参数,生成该人名对应的虚拟形象。
本申请实施例中,终端设备可以根据用户输入的人名和用户的历史信息记录,生成虚拟形象。这样,可以通过用户的历史信息记录生成更符合用户期望的虚拟形象,也使得用户对虚拟形象产生共鸣,有助于提升用户的使用体验。
图12示出了本申请实施例提供的生成虚拟形象的装置1200的示意性框图。如图12所示,该装置1200包括:获取单元1210,用于获取第一文本内容;确定单元1220,用于根据该第一文本内容和用户的历史信息记录,确定第一人格特征信息;虚拟形象生成单元1230,用于根据该第一人格特征信息,生成该第一文本内容对应的虚拟形象。
可选地,该确定单元1220,用于:在该历史信息记录中包括该第一文本内容时,根据该第一文本内容,确定该第一文本内容对应的文本共现词;根据该文本共现词,确定该第一人格特征信息。
可选地,该确定单元1220,用于:在该历史信息记录中不包括该第一文本内容时,确定该第一文本内容与该历史信息记录中一个或者多个人名中每个人名的相似度;根据该每个人名的相似度以及该每个人名对应的人格特征信息,确定该第一人格特征信息。
可选地,该装置1200还包括检测单元和控制单元,该检测单元,用于检测到用户发出语音指令或文本输入指令;该控制单元,用于控制该虚拟形象通过第二文本内容进行语音回复或者文本回复,该第二文本内容由所述历史信息记录确定。
可选地,该装置1200还包括:第一提示单元,用于根据该第一文本内容和该历史信息记录,向用户提示多个人格特征信息;所述确定单元1220,用于根据用户对该多个人格特征信息的输入,从该多个人格特征信息中确定该第一人格特征信息。
可选地,该确定单元1220,还用于根据该第一人格特征信息,确定第一风格特征信息;该虚拟形象生成单元,用于根据该第一风格特征信息,生成该虚拟形象。
可选地,该确定单元1220,用于:将该第一人格特征信息输入预测模型中,得到该第一风格特征信息;其中,该预测模型由样本训练数据训练得到,该样本训练数据包括样本人名、样本人格特征信息以及样本风格特征信息,该样本风格特征信息包括与该样本人名对应的头像、形体、音色、语速、语调、表情中的至少一项。
可选地,该装置1200还包括:第二提示单元,用于向用户提示该一个或者多个维度的风格特征信息;该虚拟形象生成单元1230,用于根据用户对该一个或者多个维度的风格特征信息的输入,生成该虚拟形象。
可选地,该一个或者多个维度的风格特征信息包括头像信息、形体信息、音色信息、语速信息、语调信息、表情信息中的至少一项。
可选地,该装置1200还包括:第三提示单元,用于根据该第一人格特征信息,向用户提示多个虚拟形象;所述虚拟形象生成单元1230,用于根据用户对该多个虚拟形象的输入,从该多个虚拟形象中确定该第一文本内容对应的虚拟形象。
例如,该获取单元1210可以是图1中的计算平台或者计算平台中的处理电路、处理 器或者控制器。以获取单元1210为计算平台中的处理器131为例,处理器131可以获取用户输入的第一文本内容。
又例如,确定单元1220以是图1中的计算平台或者计算平台中的处理电路、处理器或者控制器。以确定单元1220为计算平台中的处理器132为例,处理器132可以根据处理器131获取的第一文本内容和用户的历史信息记录,确定第一人格特征信息。
又例如,虚拟形象生成单元1230以是图1中的计算平台或者计算平台中的处理电路、处理器或者控制器。以虚拟形象生成单元1230为计算平台中的处理器133为例,处理器133可以根据处理器132确定的第一人格特征信息,生成该第一文本内容对应的虚拟形象。
以上获取单元1210所实现的功能、确定单元1220所实现的功能和虚拟形象生成单元1230所实现的功能可以由不同的处理器实现,或者,也可以由相同的处理器实现,或者,还可以是部分功能由相同的处理器实现,本申请实施例对此不作限定。
应理解以上装置中各单元的划分仅是一种逻辑功能的划分,实际实现时可以全部或部分集成到一个物理实体上,也可以物理上分开。此外,装置中的单元可以以处理器调用软件的形式实现;例如装置包括处理器,处理器与存储器连接,存储器中存储有指令,处理器调用存储器中存储的指令,以实现以上任一种方法或实现该装置各单元的功能,其中处理器例如为通用处理器,例如CPU或微处理器,存储器为装置内的存储器或装置外的存储器。或者,装置中的单元可以以硬件电路的形式实现,可以通过对硬件电路的设计实现部分或全部单元的功能,该硬件电路可以理解为一个或多个处理器;例如,在一种实现中,该硬件电路为ASIC,通过对电路内元件逻辑关系的设计,实现以上部分或全部单元的功能;再如,在另一种实现中,该硬件电路为可以通过PLD实现,以FPGA为例,其可以包括大量逻辑门电路,通过配置文件来配置逻辑门电路之间的连接关系,从而实现以上部分或全部单元的功能。以上装置的所有单元可以全部通过处理器调用软件的形式实现,或全部通过硬件电路的形式实现,或部分通过处理器调用软件的形式实现,剩余部分通过硬件电路的形式实现。
以上装置中的各单元可以是被配置成实施以上方法的一个或多个处理器(或处理电路),例如:CPU、GPU、NPU、TPU、DPU、微处理器、DSP、ASIC、FPGA,或这些处理器形式中至少两种的组合。
此外,以上装置中的各单元可以全部或部分可以集成在一起,或者可以独立实现。在一种实现中,这些单元集成在一起,以SoC的形式实现。该SoC中可以包括至少一个处理器,用于实现以上任一种方法或实现该装置各单元的功能,该至少一个处理器的种类可以不同,例如包括CPU和FPGA,CPU和人工智能处理器,CPU和GPU等。
本申请实施例还提供了一种装置,该装置包括处理单元和存储单元,其中存储单元用于存储指令,处理单元执行存储单元所存储的指令,以使该装置执行上述实施例执行的方法或者步骤。
可选地,若该装置位于终端设备中,上述处理单元可以是图1所示的处理器131-13n。
本申请实施例还提供了一种系统,该系统包括计算平台和显示装置,其中,计算平台可以包括上述装置1200。
示例性的,该显示装置可以为显示屏,如上述显示屏201-204。
本申请实施例还提供了一种终端设备,该终端设备可以包括上述装置1200,或者, 包括上述系统。
示例性的,该终端设备可以为手机、电脑或者车辆。
本申请实施例还提供了一种计算机程序产品,所述计算机程序产品包括:计算机程序代码,当所述计算机程序代码在计算机上运行时,使得计算机执行上述实施例中的方法。
本申请实施例还提供了一种计算机可读介质,所述计算机可读介质存储有程序代码,当所述计算机程序代码在计算机上运行时,使得计算机执行上述实施例中的方法。
本申请实施例还提供了一种芯片,所述芯片包括电路,所述电路用于执行上述实施例中的方法。
在实现过程中,上述方法的各步骤可以通过处理器中的硬件的集成逻辑电路或者软件形式的指令完成。结合本申请实施例所公开的方法可以直接体现为硬件处理器执行完成,或者用处理器中的硬件及软件模块组合执行完成。软件模块可以位于随机存储器,闪存、只读存储器,可编程只读存储器或者上电可擦写可编程存储器、寄存器等本领域成熟的存储介质中。该存储介质位于存储器,处理器读取存储器中的信息,结合其硬件完成上述方法的步骤。为避免重复,这里不再详细描述。
应理解,本申请实施例中,该存储器可以包括只读存储器和随机存取存储器,并向处理器提供指令和数据。
还应理解,在本申请的各种实施例中,上述各过程的序号的大小并不意味着执行顺序的先后,各过程的执行顺序应以其功能和内在逻辑确定,而不应对本申请实施例的实施过程构成任何限定。
本领域普通技术人员可以意识到,结合本文中所公开的实施例描述的各示例的单元及算法步骤,能够以电子硬件、或者计算机软件和电子硬件的结合来实现。这些功能究竟以硬件还是软件方式来执行,取决于技术方案的特定应用和设计约束条件。专业技术人员可以对每个特定的应用来使用不同方法来实现所描述的功能,但是这种实现不应认为超出本申请的范围。
所属领域的技术人员可以清楚地了解到,为描述的方便和简洁,上述描述的系统、装置和单元的具体工作过程,可以参考前述方法实施例中的对应过程,在此不再赘述。
在本申请所提供的几个实施例中,应该理解到,所揭露的系统、装置和方法,可以通过其它的方式实现。例如,以上所描述的装置实施例仅仅是示意性的,例如,所述单元的划分,仅仅为一种逻辑功能划分,实际实现时可以有另外的划分方式,例如多个单元或组件可以结合或者可以集成到另一个系统,或一些特征可以忽略,或不执行。另一点,所显示或讨论的相互之间的耦合或直接耦合或通信连接可以是通过一些接口,装置或单元的间接耦合或通信连接,可以是电性,机械或其它的形式。
所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部单元来实现本实施例方案的目的。
另外,在本申请各个实施例中的各功能单元可以集成在一个处理单元中,也可以是各个单元单独物理存在,也可以两个或两个以上单元集成在一个单元中。
所述功能如果以软件功能单元的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本申请的技术方案本质上或者说对现 有技术做出贡献的部分或者该技术方案的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行本申请各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(read-only memory,ROM)、随机存取存储器(random access memory,RAM)、磁碟或者光盘等各种可以存储程序代码的介质。
以上所述,仅为本申请的具体实施方式,但本申请的保护范围并不局限于此,任何熟悉本技术领域的技术人员在本申请揭露的技术范围内,可轻易想到变化或替换,都应涵盖。在本申请的保护范围之内。因此,本申请的保护范围应以所述权利要求的保护范围为准。

Claims (24)

  1. 一种生成虚拟形象的方法,其特征在于,包括:
    获取第一文本内容;
    根据所述第一文本内容和用户的历史信息记录,确定第一人格特征信息;
    根据所述第一人格特征信息,生成所述第一文本内容对应的虚拟形象。
  2. 如权利要求1所述的方法,其特征在于,所述根据所述第一文本内容和用户的历史信息记录,确定所述第一人格特征信息,包括:
    在所述历史信息记录中包括所述第一文本内容时,根据所述第一文本内容,确定所述第一文本内容对应的文本共现词;
    根据所述文本共现词,确定所述第一人格特征信息。
  3. 如权利要求1所述的方法,其特征在于,所述根据所述第一文本内容和用户的历史信息记录,确定所述第一人格特征信息,包括:
    在所述历史信息记录中不包括所述第一文本内容时,确定所述第一文本内容与所述历史信息记录中一个或者多个人名中每个人名的相似度;
    根据所述每个人名的相似度以及所述每个人名对应的人格特征信息,确定所述第一人格特征信息。
  4. 如权利要求1至3中任一项所述的方法,其特征在于,所述方法还包括:
    在检测到用户发出语音指令或文本输入指令时,控制所述虚拟形象通过第二文本内容进行语音回复或者文本回复,所述第二文本内容由所述历史信息记录确定。
  5. 如权利要求1至4中任一项所述的方法,其特征在于,所述根据所述第一文本内容和用户的历史信息记录,确定第一人格特征信息,包括:
    根据所述第一文本内容和所述历史信息记录,向用户提示多个人格特征信息;
    根据用户对所述多个人格特征信息的输入,从所述多个人格特征信息中确定所述第一人格特征信息。
  6. 如权利要求1至5中任一项所述的方法,其特征在于,所述根据所述第一人格特征信息,生成所述第一文本内容对应的虚拟形象,包括:
    根据所述第一人格特征信息,确定第一风格特征信息;
    根据所述第一风格特征信息,生成所述虚拟形象。
  7. 如权利要求6所述的方法,其特征在于,所述根据所述第一人格特征信息,确定第一风格特征信息,包括:
    将所述第一人格特征信息输入预测模型中,得到所述第一风格特征信息;
    其中,所述预测模型由样本训练数据训练得到,所述样本训练数据包括样本人名、样本人格特征信息以及样本风格特征信息,所述样本风格特征信息包括与所述样本人名对应的头像、形体、音色、语速、语调、表情中的至少一项。
  8. 如权利要求6或7所述的方法,其特征在于,所述根据所述第一风格特征信息,生成所述虚拟形象,包括:
    向用户提示一个或者多个维度的风格特征信息;
    根据用户对所述一个或者多个维度的风格特征信息的输入,生成所述虚拟形象。
  9. 如权利要求8所述的方法,其特征在于,所述一个或者多个维度的风格特征信息包括头像信息、形体信息、音色信息、语速信息、语调信息、表情信息中的至少一项。
  10. 如权利要求1至9中任一项所述的方法,其特征在于,所述根据所述第一人格特征信息,生成所述第一文本内容对应的虚拟形象,包括:
    根据所述第一人格特征信息,向用户提示多个虚拟形象;
    根据用户对所述多个虚拟形象的输入,从所述多个虚拟形象中确定所述第一文本内容对应的虚拟形象。
  11. 一种生成虚拟形象的装置,其特征在于,包括:
    获取单元,用于获取第一文本内容;
    确定单元,用于根据所述第一文本内容和用户的历史信息记录,确定第一人格特征信息;
    虚拟形象生成单元,用于根据所述第一人格特征信息,生成所述第一文本内容对应的虚拟形象。
  12. 如权利要求11所述的装置,其特征在于,所述确定单元,用于:
    在所述历史信息记录中包括所述第一文本内容时,根据所述第一文本内容,确定所述第一文本内容对应的文本共现词;
    根据所述文本共现词,确定所述第一人格特征信息。
  13. 如权利要求11所述的装置,其特征在于,所述确定单元,用于:
    在所述历史信息记录中不包括所述第一文本内容时,确定所述第一文本内容与所述历史信息记录中一个或者多个人名中每个人名的相似度;
    根据所述每个人名的相似度以及所述每个人名对应的人格特征信息,确定所述第一人格特征信息。
  14. 如权利要求11至13中任一项所述的装置,其特征在于,所述装置还包括检测单元和控制单元,
    所述检测单元,用于检测到用户发出语音指令或文本输入指令;
    所述控制单元,用于控制所述虚拟形象通过第二文本内容进行语音回复或者文本回复,所述第二文本内容由所述历史信息记录确定。
  15. 如权利要求11至14中任一项所述的装置,其特征在于,所述装置还包括:
    第一提示单元,用于根据所述第一文本内容和所述历史信息记录,向用户提示多个人格特征信息;
    所述确定单元,用于根据用户对所述多个人格特征信息的输入,从所述多个人格特征信息中确定所述第一人格特征信息。
  16. 如权利要求11至15中任一项所述的装置,其特征在于,
    所述确定单元,还用于根据所述第一人格特征信息,确定第一风格特征信息;
    所述虚拟形象生成单元,用于根据所述第一风格特征信息,生成所述虚拟形象。
  17. 如权利要求16所述的装置,其特征在于,所述确定单元,用于:
    将所述第一人格特征信息输入预测模型中,得到所述第一风格特征信息;
    其中,所述预测模型由样本训练数据训练得到,所述样本训练数据包括样本人名、样 本人格特征信息以及样本风格特征信息,所述样本风格特征信息包括与所述样本人名对应的头像、形体、音色、语速、语调、表情中的至少一项。
  18. 如权利要求16或17所述的装置,其特征在于,所述装置还包括:
    第二提示单元,用于向用户提示一个或者多个维度的风格特征信息;
    所述虚拟形象生成单元,用于根据用户对所述一个或者多个维度的风格特征信息的输入,生成所述虚拟形象。
  19. 如权利要求18所述的装置,其特征在于,所述一个或者多个维度的风格特征信息包括头像信息、形体信息、音色信息、语速信息、语调信息、表情信息中的至少一项。
  20. 如权利要求11至19中任一项所述的装置,其特征在于,所述装置还包括:
    第三提示单元,用于根据所述第一人格特征信息,向用户提示多个虚拟形象;
    所述虚拟形象生成单元,用于根据用户对所述多个虚拟形象的输入,从所述多个虚拟形象中确定所述第一文本内容对应的虚拟形象。
  21. 一种生成虚拟形象的装置,其特征在于,包括:
    存储器,用于存储计算机程序;
    处理器,用于执行所述存储器中存储的计算机程序,以使得所述装置执行如权利要求1至10中任一项所述的方法。
  22. 一种终端设备,其特征在于,包括如权利要求11至21中任一项所述的装置。
  23. 一种计算机可读存储介质,其特征在于,其上存储有计算机程序,所述计算机程序被计算机执行时,以使得实现如权利要求1至10中任一项所述的方法。
  24. 一种芯片,其特征在于,所述芯片包括电路,所述电路用于执行如权利要求1至10中任一项所述的方法。
PCT/CN2023/078643 2023-02-28 2023-02-28 一种生成虚拟形象的方法和装置 Ceased WO2024178590A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
CN202380089835.6A CN120457459A (zh) 2023-02-28 2023-02-28 一种生成虚拟形象的方法和装置
PCT/CN2023/078643 WO2024178590A1 (zh) 2023-02-28 2023-02-28 一种生成虚拟形象的方法和装置

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2023/078643 WO2024178590A1 (zh) 2023-02-28 2023-02-28 一种生成虚拟形象的方法和装置

Publications (1)

Publication Number Publication Date
WO2024178590A1 true WO2024178590A1 (zh) 2024-09-06

Family

ID=92589075

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2023/078643 Ceased WO2024178590A1 (zh) 2023-02-28 2023-02-28 一种生成虚拟形象的方法和装置

Country Status (2)

Country Link
CN (1) CN120457459A (zh)
WO (1) WO2024178590A1 (zh)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN108510437A (zh) * 2018-04-04 2018-09-07 科大讯飞股份有限公司 一种虚拟形象生成方法、装置、设备以及可读存储介质
US10354256B1 (en) * 2014-12-23 2019-07-16 Amazon Technologies, Inc. Avatar based customer service interface with human support agent
CN111339938A (zh) * 2020-02-26 2020-06-26 广州腾讯科技有限公司 信息交互方法、装置、设备及存储介质
US20200321020A1 (en) * 2015-10-29 2020-10-08 True Image Interactive, Inc. Systems And Methods For Machine-Generated Avatars
CN113536007A (zh) * 2021-07-05 2021-10-22 北京百度网讯科技有限公司 一种虚拟形象生成方法、装置、设备以及存储介质

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR102080878B1 (ko) * 2019-02-07 2020-02-24 류경희 가상 현실 공간에서 서비스 제공을 위한 캐릭터 생성 및 학습 시스템
CN110812843B (zh) * 2019-10-30 2023-09-15 腾讯科技(深圳)有限公司 基于虚拟形象的交互方法及装置、计算机存储介质
CN111309886B (zh) * 2020-02-18 2023-03-21 腾讯科技(深圳)有限公司 一种信息交互方法、装置和计算机可读存储介质
CN115317924B (zh) * 2022-07-07 2025-06-10 网易(杭州)网络有限公司 游戏角色的生成方法、装置和电子设备
CN115222857A (zh) * 2022-07-27 2022-10-21 北京中电慧声科技有限公司 生成虚拟形象的方法、装置、电子设备和计算机可读介质

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10354256B1 (en) * 2014-12-23 2019-07-16 Amazon Technologies, Inc. Avatar based customer service interface with human support agent
US20200321020A1 (en) * 2015-10-29 2020-10-08 True Image Interactive, Inc. Systems And Methods For Machine-Generated Avatars
CN108510437A (zh) * 2018-04-04 2018-09-07 科大讯飞股份有限公司 一种虚拟形象生成方法、装置、设备以及可读存储介质
CN111339938A (zh) * 2020-02-26 2020-06-26 广州腾讯科技有限公司 信息交互方法、装置、设备及存储介质
CN113536007A (zh) * 2021-07-05 2021-10-22 北京百度网讯科技有限公司 一种虚拟形象生成方法、装置、设备以及存储介质

Also Published As

Publication number Publication date
CN120457459A (zh) 2025-08-08

Similar Documents

Publication Publication Date Title
US12462441B2 (en) Iterative image generation from text
US20250342828A1 (en) Digital assistant control of applications
US20190163768A1 (en) Automatically curated image searching
CN113330455A (zh) 使用有条件的生成对抗网络查找互补的数字图像
WO2024022437A1 (zh) 一种显示方法、装置和移动载体
CN115408611A (zh) 菜单推荐方法、装置、计算机设备和存储介质
CN113486260A (zh) 互动信息的生成方法、装置、计算机设备及存储介质
CN115858850A (zh) 内容推荐方法、装置、车辆及计算机可读存储介质
CN118276746A (zh) 用于图像编辑的方法、装置、设备、介质和程序产品
JP7289756B2 (ja) 生成装置、生成方法および生成プログラム
WO2024178590A1 (zh) 一种生成虚拟形象的方法和装置
CN120804187A (zh) 视觉数据生成方法、装置、电子设备及可读存储介质
CN116166823A (zh) 基于用户偏好的多媒体信息展示方法及装置、存储介质
WO2025036359A1 (zh) 图像处理方法、装置、计算机设备及存储介质
WO2025039700A1 (zh) 一种搜索结果排序方法、装置及系统
CN118409685A (zh) 车机界面的显示方法、装置、电子设备、车辆及存储介质
CN116797322A (zh) 提供商品对象信息的方法及电子设备
CN112734949B (zh) Vr内容的属性修改方法、装置、计算机设备及存储介质
CN114741602A (zh) 对象推荐方法、目标模型的训练方法、装置及设备
CN114816038A (zh) 虚拟现实内容生成方法、装置及计算机可读存储介质
US20240119489A1 (en) Product score unique to user
WO2025102351A1 (zh) 语音交互方法和装置
KR20260029177A (ko) 보조 미디어 컨텐츠를 제공하는 방법 및 이를 수행하는 전자 장치
CN121329550A (zh) 商品推荐方法及系统
CN109167723B (zh) 图像的处理方法、装置、存储介质及电子设备

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23924565

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 202380089835.6

Country of ref document: CN

WWP Wipo information: published in national office

Ref document number: 202380089835.6

Country of ref document: CN

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 23924565

Country of ref document: EP

Kind code of ref document: A1