WO2020007185A1 - 图像处理方法、装置、存储介质和计算机设备 - Google Patents
图像处理方法、装置、存储介质和计算机设备 Download PDFInfo
- Publication number
- WO2020007185A1 WO2020007185A1 PCT/CN2019/091359 CN2019091359W WO2020007185A1 WO 2020007185 A1 WO2020007185 A1 WO 2020007185A1 CN 2019091359 W CN2019091359 W CN 2019091359W WO 2020007185 A1 WO2020007185 A1 WO 2020007185A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- audio data
- virtual object
- attribute information
- scene image
- attribute
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N23/00—Cameras or camera modules comprising electronic image sensors; Control thereof
- H04N23/80—Camera processing pipelines; Components thereof
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T19/00—Manipulating three-dimensional [3D] models or images for computer graphics
- G06T19/006—Mixed reality
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/78—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
- G06F16/783—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
- G06F16/7834—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content using audio features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/78—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
- G06F16/783—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content
- G06F16/7837—Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using metadata automatically derived from the content using objects detected or recognised in the video content
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/25—Fusion techniques
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T3/00—Geometric image transformations in the plane of the image
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/40—Extraction of image or video features
- G06V10/44—Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/764—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/20—Scenes; Scene-specific elements in augmented reality scenes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
- G06V40/16—Human faces, e.g. facial parts, sketches or expressions
- G06V40/168—Feature extraction; Face representation
- G06V40/171—Local features and components; Facial parts ; Occluding parts, e.g. glasses; Geometrical relationships
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/10—Human or animal bodies, e.g. vehicle occupants or pedestrians; Body parts, e.g. hands
- G06V40/16—Human faces, e.g. facial parts, sketches or expressions
- G06V40/172—Classification, e.g. identification
Definitions
- the present application relates to the field of image processing technologies, and in particular, to an image processing method, apparatus, storage medium, and computer equipment.
- the user when recording a video, the user can freely select virtual objects in the recording interface of the client, and add these virtual objects to the corresponding position of the image frame corresponding to the video, so that the virtual object can follow the movement of the moving target in the image frame and move.
- the virtual object can only move following the movement of the moving target, and the interaction is poor.
- an image processing method, device, storage medium, and computer device are provided, which can solve the technical problem that a virtual object can only move following the movement of a moving target, making the interaction poor.
- An image processing method which is applied to an image processing system, and the method includes:
- An image processing device comprising:
- An audio data acquisition module configured to acquire audio data corresponding to a real-time scene image collected in real time
- An attribute information determining module configured to dynamically determine attribute information of a virtual object according to the audio data, and the attribute information is used to determine a visual state of the virtual object;
- a target object determination module configured to determine a target object from the real scene image
- a fusion position determination module configured to determine, according to the target object, a fusion position of a virtual object determined according to the attribute information in the real scene image
- a fusion module is configured to fuse a virtual object determined according to the attribute information to the real scene image according to the fusion position; the virtual object presents different visual states when the attribute information is different.
- a storage medium is characterized in that a computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute the steps of the image processing method described above.
- a computer device includes a memory and a processor.
- the memory stores a computer program.
- the processor causes the processor to perform the following operations:
- the above-mentioned image processing method, device, storage medium and computer equipment acquire audio data corresponding to real-time scene images collected in real time, and dynamically determine attributes of virtual objects through the audio data, thereby implementing control of the attributes of virtual objects according to the audio data.
- the fusion position of the virtual object in the real scene image is determined, and the virtual object determined based on the attribute information is fused into the real scene image according to the fusion position. Since the attribute information of the virtual object is controlled by audio data, When the audio data changes, the attribute information of the virtual object integrated into the real scene image also changes, which improves the interactivity.
- FIG. 1 is a system structural diagram of an image processing method according to an embodiment
- FIG. 2 is a schematic flowchart of an image processing method according to an embodiment
- FIG. 3 is a schematic diagram of fusing a virtual object into a real scene image in an embodiment
- FIG. 4 is a schematic diagram of fusing a virtual object into a real scene image in an embodiment
- FIG. 5 is a schematic diagram of determining attribute information according to audio data and fusing a virtual object having the attribute information into a real scene image according to an embodiment
- FIG. 6 is a schematic flowchart of a step of determining attribute information of a virtual object according to parameter values of audio data in an embodiment
- FIG. 7 is a schematic diagram of sampling, quantizing, and encoding audio data in an embodiment
- FIG. 8 is a schematic flowchart of a step of determining a frequency value according to encoded audio data in an embodiment
- FIG. 9 is a schematic flowchart of a step of adjusting a virtual object according to an attribute adjustment amount and fusing the adjusted virtual object into a real scene image according to an embodiment
- FIG. 10 is a schematic flowchart of a step of adjusting a virtual object according to an attribute change target value and fusing the adjusted virtual object into a real scene image;
- FIG. 11 is a schematic flowchart of a step of determining a fusion position of a virtual object in a real scene image according to the characteristics of the target object according to an embodiment
- FIG. 12 is a schematic diagram of a facial feature point of a target object in an embodiment
- FIG. 13 is a schematic flowchart of an image processing method in another embodiment
- FIG. 14 is a structural block diagram of an image processing apparatus according to an embodiment
- FIG. 15 is a structural block diagram of a computer device in one embodiment.
- FIG. 1 is an application environment diagram of an image processing method in an embodiment.
- the image processing method is applied to an image processing system.
- the image processing system may be a terminal or a combination of multiple terminals, and the terminal may be a smartphone, a computer, or other equipment that can support AR (Augmented Reality, Augmented Reality) technology.
- the image processing system may include: a camera, a scene generator, an image synthesizer, and a display; of which:
- the camera is configured to obtain a real scene image corresponding to the environment of the target object, and send the obtained real scene image to the image synthesizer to perform a synthesis operation with the virtual object of the augmented reality model.
- the scene generator is used to determine the fusion position of the virtual object according to the position information of the target object in the real scene image, such as determining the fusion position of the virtual object by analyzing the characteristics of the target object, and then sending the virtual object to the image synthesizer.
- An image synthesizer is used to fuse real scene images and virtual objects about a target object according to a fusion position, and output the fusion result to a display screen.
- the display is used to display the fused image output by the image synthesizer to form a display effect of the target object and the virtual object used for the augmented reality model.
- an image processing method is provided. This embodiment is described by using the method applied to the terminal in FIG. 1 as an example. Referring to FIG. 2, the image processing method specifically includes the following steps:
- the realistic scene may refer to a realistic picture viewed by a user through a certain medium, and the realistic picture includes at least one of the following: a person, a natural landscape, a human landscape, and a human intelligent work.
- Human wisdom works refer to works created by humans through labor and wisdom.
- the real scene is a picture of people and nature viewed by the user through the naked eye, or a picture of a stereoscopic movie viewed by the user through 3D glasses.
- the real scene image may be an image about a real scene collected through a terminal.
- the real scene image is an image of a real scene collected in real time through a camera in FIG. 1. After the terminal collects multiple real scene images, the multiple real scene images are combined according to the collection time to obtain a video.
- the audio data is an audio signal in the time domain.
- the audio data carries information about the frequency and amplitude of regular sound waves of speech, music, and sound effects. According to the characteristics of sound waves, audio data can be classified into regular audio and irregular audio. Regular audio can be divided into speech, music and sound effects.
- the source of audio data may be obtained by the terminal from the outside, or it may be obtained by reading from the background audio with the image of the real scene.
- S202 may specifically include: when acquiring a real-world scene image in real time, the terminal collects real-time audio data corresponding to the real-world scene image from the current environment; or, the terminal acquires real-time audio from the background audio of the real-world scene image, Read the audio data corresponding to the timestamp corresponding to the real scene image.
- the terminal when collecting images of a real scene in real time, the terminal collects music or speaking sounds in the current environment through a microphone. Or, during the development process, the developer sets the background music to be automatically played when the real scene image is collected in real time.
- the time stamp corresponding to the real scene image is generated and read in the automatically played background music. Audio data corresponding to the timestamp corresponding to the real scene image.
- the music client is playing music at the same time, and the terminal will generate a real-time time stamp corresponding to the real scene image when the real-time scene image is collected in real time. Audio data corresponding to the timestamp.
- the virtual object may include an image material, such as a static sticker or a dynamic sticker.
- the virtual object may further include a virtual prop for enhancing the display effect of the target object.
- the virtual props may be various virtual pendants and virtual backgrounds used to dress up a target object.
- the virtual object may include one or more attributes, and each attribute has attribute information, which is used to determine a visual state of the virtual object.
- the attribute information of the virtual object may include at least one of the following: an attribute adjustment amount and an attribute change target value. Attribute adjustments include scaling, rotation, and offset of virtual objects.
- the attribute change target value includes the color RGB value of the virtual object.
- the attribute adjustment amount is used to indicate the adjustment range of the attributes of the virtual object. According to the attribute adjustment amount, the corresponding attributes of the virtual object can be adjusted to determine the adjusted attributes.
- the attribute change target value is used to indicate the target attribute value of the attribute of the virtual object, and the corresponding attribute of the virtual object can be adjusted according to the attribute change target value, so that the adjusted attribute is the attribute change target value.
- the attribute information of the determined virtual object is the current attribute information of real-time scene image collection, and the attribute information of the determined virtual object may be different for different real scene images. For example, suppose a user shoots a video with AR effect. If the i-th frame of the real scene image in the video is the current real scene image, the attribute information of the virtual object at the time corresponding to the i-th frame of the real scene image is the current attribute information.
- the i-1 frame of the real scene image in the video is the real scene image of the previous moment, and the attribute information of the virtual object at the time corresponding to the i-1 frame of the real scene image is the attribute information of the last moment.
- i is a positive integer greater than or equal to 1.
- the terminal dynamically determines the attribute information of the virtual object according to the audio data, adjusts the virtual object according to the determined attribute information, and obtains the virtual object determined according to the attribute information.
- the audio data has parameter values such as a volume value, a frequency value, and a tone color.
- the terminal dynamically determines the attribute information of the virtual object according to the audio data.
- the terminal updates the original attribute information of the virtual object to the determined attribute information, adjustment of the virtual object can be achieved. For example, the terminal adjusts the zoom ratio of the virtual object or adjusts the color RGB value of the virtual object according to the frequency value of the audio data.
- the target object may be a living object in nature, such as a human, an animal, or a plant.
- S206 may specifically include: identifying a biological feature from a real scene image; and when the biological feature meets a preset condition, determining a biological object corresponding to the biological feature in the real scene image as a target object.
- the biological feature may be a contour feature of the creature, or may be a detailed feature of the creature, such as a person's face feature.
- the preset condition may include a preset biometric or a preset completeness threshold of the biometric.
- the terminal recognizes the biometric feature from the image of the real scene; when the biometric feature meets the preset biometric feature and / or the integrity of the biometric feature reaches a preset integrity threshold, it determines the object corresponding to the biometric feature in the real scene image as the target Object.
- an object corresponding to the biological feature in the real scene image is determined as the target object.
- an object corresponding to the biological feature in the real scene image is determined as the target object.
- an object corresponding to the biological feature in the real scene image is determined as the target object.
- an object corresponding to the biological feature in the real scene image is determined as the target object.
- the fused position may refer to a position of a virtual object's center point or key position point in the real scene image when fused to the real scene image.
- the fusion position can be one position point or multiple position points.
- different fusion positions can be determined.
- different virtual objects can be fused in different parts of the target object, and different parts of the same virtual object can be fused in different parts of the target object.
- the terminal detects the features of the target object, selects features matching the virtual object from the features of the target object, and determines the fusion position of the virtual object determined by the attribute information in the real scene image according to the selected features.
- the characteristics of the virtual object matching may be determined according to the type of the virtual object, and different types of virtual objects may match the same feature or different features.
- multiple types of virtual objects can be set, and matching characteristics can be set for each type of virtual object, so as to establish a matching relationship between the virtual object and the feature, and then the virtual object matching can be determined according to the matching relationship.
- feature For example, the features matched by AR diving glasses are eye features, and the features matched by virtual props dressed by AR rabbits are mouth features.
- determining the fusion position of the virtual object determined by the attribute information in the real scene image may include: determining the position of the selected feature as the fusion position of the virtual object in the real scene image.
- the terminal detects the characteristics of the target object, determines the eye characteristics from the detected characteristics, and determines the fusion position of the virtual object according to the eye characteristics. Is the eye position of the user.
- the terminal detects the characteristics of the target object and determines the mouth from the detected features. Feature, and the fusion position of the virtual object is determined as the mouth position of the user according to the mouth feature.
- the terminal detects the features of the target object, determines the head features (such as hair) from the detected features, and determines the fusion position of the virtual object as the head position of the user according to the head features.
- the virtual object may include multiple parts, and different parts may match different features.
- the virtual object is a raincoat virtual prop
- the head matching feature in the raincoat virtual prop is a head feature
- the arm matching feature in the raincoat virtual prop is an arm feature.
- the terminal adjusts the virtual object according to the attribute information, so that the virtual object can be scaled in size, changed in color RGB value, or changed in rotation angle.
- the virtual object adjusted by the terminal according to the attribute information is a reference to the virtual object in the fused real scene image of the previous frame as a reference, so as to scale the current virtual object, change the color RGB value, or change the rotation angle. And so on.
- the virtual object determined according to the attribute information is fused to the real scene image according to the fusion position; the virtual object presents different visual states when the attribute information is different.
- the terminal determines a center point or a key position point of the virtual object, and fuses the center point or the key position point of the virtual object to an area matching the fusion position, thereby fusing the virtual object into a real scene image.
- the center point or key position point is used to determine the area to be fused by the virtual object, and is used to match the corresponding fusion position, that is, when the fusion is performed, the center point of the virtual object is placed at the corresponding fusion position, or the The key position points are placed at the corresponding fusion positions.
- p represents the key position points of the virtual object
- the key position points are the three positions of the virtual props of the AR diving glasses.
- q represents a key position point of the virtual object, which is respectively located in an upper part of a rabbit tooth virtual prop dressed up by an AR rabbit, and a lower part of an ear virtual prop dressed up by two AR rabbits.
- the AR rabbit-dressed virtual props are fused to the target's lips and head according to key positions, so as to achieve accurate fusion of AR rabbit-dressed virtual props, specifically The effect is shown in Figure 4 (b).
- the virtual object presents different visual states when the attribute information is different. For example, suppose a user shoots a video with AR effects. If the i-th frame of the real scene image in the video is the current real scene image, the i-th frame of the real scene image The attribute information of the virtual object at the corresponding time is the current attribute information.
- the i-1 frame of the real scene image in the video is the real scene image of the previous moment, and the attribute information of the virtual object at the time corresponding to the i-1 frame of the real scene image is the attribute information of the last moment.
- i is a positive integer greater than or equal to 1.
- the attribute information of the virtual object changes as the audio data changes. If the volume value or frequency value of the audio data changes, the size, color RGB value, or orientation of the virtual object changes accordingly.
- FIG. 5 (a) is a real scene image after the previous fusion
- m is a virtual object of an original size.
- the parameter value of the acquired audio data changes, such as the volume level changes, or the frequency value changes
- the size of the corresponding virtual object also changes.
- the changed virtual object is shown as n in Figure 5 (b).
- n is a virtual object with an enlarged size.
- audio data corresponding to a real-time scene image collected in real time is acquired, and the attributes of the virtual object are dynamically determined through the audio data, thereby implementing control of the attributes of the virtual object according to the audio data.
- the fusion position of the virtual object in the real scene image is determined, and the virtual object determined based on the attribute information is fused into the real scene image according to the fusion position. Since the attribute information of the virtual object is controlled by audio data, When the audio data changes, the attribute information of the virtual object integrated into the real scene image also changes, which improves the interactivity.
- S204 may specifically include:
- the parameter values of the audio data include a volume value, a frequency value, a tone color, and the like of the audio data.
- the volume value can be any of the following: average frequency value, maximum volume value, or minimum volume value.
- the frequency value can be any of the following: average frequency value, maximum frequency value, or minimum frequency value.
- the terminal obtains parameter values such as a volume value, a frequency value, and a tone color in the audio data by analyzing the audio data.
- S602 may specifically include: sampling audio data; quantizing and encoding the sampled result to obtain encoded audio data; and determining parameter values of the audio data according to the obtained encoded audio data.
- the terminal uses PCM (Pulse Code Modulation, Pulse Code Modulation) to sample, quantize, and encode the collected time-domain continuous time-domain audio data to obtain binary encoded audio data.
- PCM Pulse Code Modulation, Pulse Code Modulation
- the terminal according to the obtained
- the encoded audio data determines parameter values of the audio data, such as determining the volume value of the audio data.
- the audio data U (t) is sampled to discretize temporally continuous time-domain audio data.
- the discretized audio data is quantized to obtain quantized audio data in M base, where M is a positive integer greater than 2.
- the quantized audio data is encoded to obtain binary encoded audio data.
- different types of parameter values correspond to different preset mapping relationships
- the terminal determines a preset mapping relationship between the parameter value and attribute information of the virtual object according to the type of the parameter value.
- the type includes volume type, frequency type and tone type.
- the parameter values include volume value, frequency value and tone color.
- the terminal determines a preset mapping relationship between the volume and the scaling ratio of the virtual object.
- the preset mapping relationship may be a functional relationship, as shown below:
- x is the volume value of the audio data, and the range can be 0 to 120 decibels (db), and f (x) is the scaling ratio of the virtual object.
- S606 Map the parameter value to the attribute information of the virtual object according to the preset mapping relationship.
- a parameter value is input as a variable into the preset mapping relationship to obtain corresponding attribute information of the virtual object.
- the volume value is 40 db
- the volume value 40 is input into the function f (x)
- the obtained attribute information is a scaling ratio of 1, that is, the original virtual object is not subjected to any enlargement or reduction processing.
- the volume value is 120 db
- the volume value 120 is input into the function f (x)
- the obtained attribute information is a scaling ratio of 4, that is, the original virtual object is enlarged to the original 4 times.
- a preset mapping relationship between a parameter value and attribute information of a virtual object is determined.
- the attribute information of the corresponding virtual object can be obtained through the preset mapping relationship, thereby The adjustment of the virtual object is realized, so that the virtual object presents different visual states, and the diversity of the virtual object is improved.
- the attribute information of the virtual object may be determined by a parameter value of the audio data.
- the parameter value may be a frequency value or a volume value.
- the step of encoding audio data to determine parameter values of the audio data may specifically include:
- the encoded audio data is a discretized audio signal.
- the terminal converts the encoded audio data into frequency domain audio data according to the discrete Fourier transform.
- the frequency-domain audio data includes the amplitude value (that is, the volume value), the frequency value, and the phase of the audio data.
- S804 Segment frequency-domain audio data to obtain multiple sub-frequency-domain audio data.
- the terminal segments frequency-domain audio data according to a set step size to obtain multiple sub-frequency-domain audio data. For example, when using a 512-point Fourier transform, the frequency band from 0 to the cutoff frequency (at the sampling rate of 48kHz, the cutoff frequency is 24kHz) can be equally divided into 256 bands, and then S806 is performed to determine the amplitude in each band .
- the terminal divides the frequency domain audio data into a plurality of frequency bands of unequal length in an unequal manner to obtain a plurality of sub-frequency domain audio data.
- S806 Determine the amplitude of the audio data in each sub-frequency domain.
- each sub-frequency-domain audio data includes amplitude, frequency value and phase
- the terminal determines the amplitude in each sub-frequency-domain audio data to obtain the volume value in each sub-frequency-domain audio data.
- a large amplitude of the audio data indicates that the power of the audio data is large.
- the corresponding useful signals are large, and when the power of the audio data is small, the corresponding useful signals are small.
- the terminal collects audio data through a microphone. When the power of the collected audio data is low, it indicates that the currently collected audio data may be a noise signal. Therefore, the sub-frequency domain audio data with the largest amplitude can be selected.
- the amplitudes of the audio data in each sub-frequency domain are compared to obtain the audio data in the sub-frequency domain with the largest amplitude.
- the terminal arranges the sub-frequency domain audio data according to the amplitude, and selects the sub-frequency domain audio data with the largest amplitude from the arranged sub-frequency domain audio data.
- the frequency domain audio data is segmented, and the frequency value can be determined from the sub-frequency domain audio data obtained after the segmentation.
- the frequency value can be used to adjust the virtual object; on the other hand, the segmentation
- the subsequent frequency domain audio data can reduce the calculation amount and speed up the calculation rate during the calculation process.
- the attribute information of the virtual object may be determined by a parameter value of the audio data.
- the parameter value may be a frequency value or a volume value.
- the audio data is determined according to the obtained encoded audio data.
- the step of the parameter value may specifically include: determining a volume value according to the obtained encoded audio data; or converting the time-domain encoded audio data into frequency-domain audio data; and determining the volume value according to the frequency-domain audio data.
- the amplitude of the time-domain encoded audio data may represent the volume value of the audio data.
- the terminal determines the amplitude of the time-domain encoded audio data as the volume value of the audio data.
- the encoded audio data is a discretized audio signal.
- the terminal converts the encoded audio data into frequency domain audio data according to the discrete Fourier transform.
- the frequency-domain audio data includes the amplitude (that is, the volume value), the frequency value, and the phase of the waveform corresponding to the audio data.
- the terminal determines the amplitude in the frequency domain audio data as the volume value of the audio data.
- the terminal after the terminal converts the encoded audio data into frequency domain audio data, the terminal segments the frequency domain audio data according to a set step size to obtain multiple sub-frequency domain audio data.
- the terminal determines the corresponding amplitude value according to the audio data in each sub-frequency domain, determines the maximum amplitude value as the volume value of the audio data, or determines the average amplitude value as the volume value of the audio data.
- the volume value is determined according to the obtained encoded audio data or the frequency domain audio data converted from the time-domain encoded audio data, and attribute information for determining the visual state of the virtual object is obtained.
- the volume value can be adjusted for virtual objects.
- the attribute information of the virtual object may include at least one of the following: an attribute adjustment amount and an attribute change target value.
- the attribute adjustment amount may include a zoom ratio, a rotation angle, and an offset of the virtual object, and the attribute change target value may include a color RGB value of the virtual object.
- the attribute information is the attribute adjustment amount
- the virtual object determined by the attribute information is a virtual object adjusted by the attribute according to the attribute adjustment amount.
- S210 may specifically include:
- S902 Determine an attribute of the virtual object and corresponding to the attribute adjustment amount.
- attributes refer to attributes owned by the virtual object, including scaling, color, rotation, and offset.
- the attributes corresponding to the attribute adjustment amount include scaling, rotation, and offset.
- the attribute adjustment amount is a specific value corresponding to the attribute.
- the terminal determines, according to the parameter value of the audio data, an attribute of the virtual object and corresponding to the attribute adjustment amount.
- the terminal after determining the attribute adjustment amount corresponding to the parameter value of the audio data according to the mapping relationship, the terminal adjusts the virtual object according to the attribute adjustment amount, changes the attributes of the virtual object, and obtains the virtual object after adjusting the attribute.
- the size of the virtual object is adjusted according to the zoom ratio to obtain the adjusted virtual object.
- the attributes of the virtual object are adjusted by the attribute adjustment amount, and the virtual object after the adjustment of the attribute is fused to the real scene image according to the fusion position to obtain the virtual object that changes with the parameters of the audio data, so as to adjust the virtual object according to the audio data
- the diversity of virtual objects has been improved and the user experience has been enhanced.
- the attribute information of the virtual object may include at least one of the following: an attribute adjustment amount and an attribute change target value.
- the attribute adjustment amount may include a zoom ratio, a rotation angle, and an offset of the virtual object
- the attribute change target value may include a color of the virtual object.
- the virtual object determined according to the attribute information is a virtual object after the corresponding attribute is changed to the attribute change target value.
- S210 may specifically include:
- S1002 Determine an attribute of the virtual object and corresponding to the attribute change target value.
- the attribute corresponding to the attribute change target value includes the color of the virtual object.
- the attribute change target value is the specific value corresponding to the attribute, such as the color RGB value.
- the terminal determines, according to the parameter value of the audio data, an attribute of the virtual object and corresponding to the attribute change target value.
- S1004 Change the attributes of the virtual object to the attribute change target value to obtain the virtual object after the attribute change.
- the terminal after determining the attribute change target value corresponding to the parameter value of the audio data according to the mapping relationship, the terminal adjusts the virtual object according to the attribute change target value, causes the attribute of the virtual object to change, and obtains the virtual object after adjusting the attribute. .
- the display color of the virtual object is adjusted according to the target color RGB value, so that the original display color of the virtual object is adjusted to the color corresponding to the target color RGB value. If the original is red, pass the target After adjusting the color RGB value, a blue virtual object is obtained.
- the attributes of the virtual object are adjusted by the target value of the attribute change, and the virtual object after adjusting the attribute is merged into the real scene image according to the fusion position to obtain the virtual object that changes with the parameters of the audio data, so as to adjust the virtual object according to the audio data , Enhances the diversity of virtual objects, and enhances the user experience.
- the virtual object determined by the attribute information is the first attribute adjusted by the attribute adjustment amount
- the The virtual object after the corresponding second attribute changes to the attribute change target value.
- S210 may specifically include: determining a first attribute possessed by the virtual object and corresponding to the attribute adjustment amount, determining a second attribute possessed by the virtual object and corresponding to the attribute change target value, and adjusting the virtual object according to the attribute adjustment amount.
- the first attribute is to change the second attribute of the virtual object to the attribute change target value, to obtain the virtual object after the attribute change, and fuse the virtual object after the attribute change to the real scene image according to the fusion position.
- S208 may specifically include:
- the terminal detects the features of the target object through feature point detection methods, such as cascade regression CNN (Convolutional Neural Networks, Convolutional Neural Networks), or feature point detection methods such as Dlib, Libfacedetect, or Seetaface.
- feature point detection methods such as cascade regression CNN (Convolutional Neural Networks, Convolutional Neural Networks), or feature point detection methods such as Dlib, Libfacedetect, or Seetaface.
- each facial feature point obtained by digital mark recognition is used.
- 1 to 17 shown in FIG. 12 indicate the target.
- the feature points of the subject's face edge, 18 to 22 represent the left eyebrow feature points of the target object
- 23 to 27 represent the right eyebrow feature points of the target object
- 28 to 36 represent the nose feature points of the target object
- 37 to 42 represent the target
- the subject's left-eye feature points, 43-48 represent the target's right-eye feature points
- 49-68 represent the target's lip feature points.
- S1104 Find a feature that matches a virtual object with attributes among the detected features.
- the fusion position in the real scene image corresponding to different virtual objects is also different.
- the virtual prop of the AR diving glasses corresponds to the fusion position in the real scene image, and should be the eye position of the target object.
- the ears of the virtual props dressed as AR rabbits correspond to the fusion position in the real scene image, which should be the head position of the target object; and the rabbit teeth of the virtual props dressed as AR rabbits correspond to the fusion position in the real scene image.
- the AR kitten dress corresponding to the fusion position in the image of the real scene should be the positions of the faces on both sides of the target object.
- the terminal determines the function or purpose of the virtual object, determines the location where the virtual object is to be mounted on the target object according to the function or purpose, and then determines the relationship between the virtual object and the characteristics of the target object. Among the detected features, the terminal obtains features that match the virtual object with attributes according to the determined relationship.
- S1106 Determine the fusion position of the virtual object with attributes in the real scene image according to the matched features.
- the fusion position of the virtual object in the real scene image is determined by the feature points of the target object, so that the virtual object is fused in the real scene image according to the fusion position, and the virtual object with a changed visual state is obtained, which improves The diversity of virtual objects has changed.
- the method may further include: extracting audio characteristics of the audio data; when the audio characteristics meet the first trigger condition, performing at least one of the following: adding a virtual object; switching a virtual object; switching a type of visual state .
- the audio characteristics may include at least one of the following: a volume value, a frequency value, a tone color, a logarithmic power spectrum, and a Mel frequency cepstrum coefficient of audio data.
- the logarithmic power spectrum and Mel frequency cepstrum coefficient can reflect the power value of the audio data as well as the characteristics of the speaker's style and voice expression.
- Speech expressiveness can be characterized by the tone, severity, and rhythm of speech.
- the corresponding first trigger condition may include that the volume value reaches a preset volume threshold, or the frequency value reaches a preset frequency threshold, or the tone color meets the tone condition, or the power value reaches the power threshold, or the speaker's style characteristics meet the style characteristic conditions, or The speaker's speech expressiveness meets expressiveness conditions and the like.
- the type of the visual state may be a display size, a display color, a motion track, and the like of the virtual object.
- the terminal performs frame framing and windowing on the audio data in the time domain to obtain audio data of each frame.
- the terminal performs Fourier transform on the audio data of each frame to obtain the corresponding frequency spectrum.
- the terminal calculates a power spectrum according to the spectrum of each frame, and then performs a logarithmic operation on the power spectrum to obtain a logarithmic power spectrum.
- the terminal may determine the logarithmic power spectrum as a speech feature, or determine a result obtained by performing a discrete cosine transform on the logarithmic power spectrum as a speech feature.
- the signal expression of the collected speech is x (n)
- the windowed speech x' (n) x (n) ⁇ h (n) performs discrete Fourier transform, and the corresponding spectrum signal is:
- N represents the number of points of the discrete Fourier transform.
- the terminal When obtaining the frequency spectrum of each frame of speech, the terminal calculates the corresponding power spectrum and obtains the logarithmic value of the power spectrum to obtain the logarithmic power spectrum, thereby obtaining the corresponding speech characteristics.
- the terminal inputs the logarithmic power spectrum into a Mel-scale triangular filter, and obtains a Mel frequency cepstrum coefficient after discrete cosine transform.
- the obtained Mel frequency cepstrum coefficient is:
- the L order refers to the order of the Mel frequency cepstrum coefficient, and the value can be 12-16.
- M refers to the number of triangular filters.
- a new virtual object can be added to the original virtual object, or the original virtual object can be switched to another virtual object, or
- the original visual state is switched to diversify the virtual objects fused in the image of the real scene and the visual state to be diversified, thereby improving the interaction between the user and the virtual object.
- the method may further include: identifying according to audio data to obtain a recognition result; determining a dynamic effect type matching the recognition result; determining a visual state presented by the virtual object according to the dynamic effect type and attribute information; and a visual state Matches the type of dynamic effect.
- the recognition result may refer to the audio type and the text features of keywords in the audio data.
- Audio type can refer to the type of music, such as light music, rock music, and jazz music.
- Text features can refer to the light and accented keywords. Stress can be expressed by increasing the intensity or pitch.
- Dynamic effects can be the effects that virtual objects display during dynamic changes.
- the dynamic effect may be any one or a combination of the following: rotation, movement, change between transparent and opaque, color change, and the like.
- the virtual object rotates as the audio data changes, or rotates while moving.
- the dynamic effect type may include a rotation type, a movement type, a change type between transparent and non-transparent, and a color change type.
- the terminal when the terminal recognizes the music type corresponding to the audio data, it acquires a dynamic effect corresponding to the music type, and determines the visual state presented by the virtual object according to the acquired dynamic effect and attribute information. For example, when the acquired audio data is a type of rock music, the dynamic effect may be a more dynamic effect.
- the terminal recognizes the text features of the keywords in the audio data and selects the corresponding dynamic effect according to the recognized text features. For example, when the key word in the audio is recognized as an accent, the dynamic effect of the virtual object is switched to the dynamic effect corresponding to the accent.
- the corresponding dynamic effect is determined by the recognition result of the audio data, so that the virtual object presents different dynamic effects as the audio data changes, and the interaction between the user and the virtual object is improved.
- an embodiment of the present application provides an image processing method.
- a virtual object (such as a virtual object changing with music) can be dynamically adjusted according to music changes, so that the virtual object follows the volume value of the music. Changes in color, size, or rotation angle occur as the frequency value changes.
- the image processing method includes the following steps:
- the terminal may obtain audio data in the following ways: one is to collect audio data through the microphone of the terminal, and the other is to read audio data from the corresponding background music played by the terminal.
- the way to collect audio data through the microphone of the terminal is to collect audio data from the outside world, such as using the microphone function commonly used in mobile phones to collect the user's voice.
- the way to read audio data from the corresponding background music played by the terminal is that the terminal decodes the audio format file of the played background music to obtain audio data.
- one of the audio data obtained by the above two methods may be used as an input source, and the audio data obtained by the above two methods may also be used as an input source.
- the terminal encodes the acquired audio data into binary encoded audio data in a PCM coding and modulation mode. Audio data may also be referred to as audio signals, and no distinction is made in the embodiments of the present application.
- PCM is a common encoding method, which samples the simulated audio data at a preset time interval, discretizes the simulated audio data, and then quantizes the sampled values, and encodes the quantized sampled values. Obtains the amplitude of the sampling pulse in binary code.
- the terminal obtains the encoded audio data after PCM encoding, and parses out the sound-related attribute information from the encoded audio data.
- the attribute information may include a volume value, a frequency value, a tone color, and the like.
- the volume value can be expressed by the loudness of the audio data or the amplitude of the corresponding waveform, which characterizes the volume of the audio data over a period of time.
- the calculation formula is as follows:
- v i represents After the resultant PCM-encoded audio data encoded in the amplitude of one sample point, N represents the number of sampling points, the present embodiment can take the values N 1024, may be other values. For audio data with a sampling rate of 48k, 48 volume values can be calculated in one second.
- the frequency value can be the number of times the audio data vibrates up and down the corresponding waveform in a unit time, and the unit is Hz.
- Audio data can be decomposed into superposition of sine waves with different frequency values and different amplitudes.
- FFT Fast Fourier Transform
- the encoded audio data obtained after PCM encoding can be converted into frequency domain audio data.
- Frequency and volume values ie, amplitude values
- a 512-point FFT can be used, and the frequency band from 0 to the cutoff frequency (for a 48kHz sampling rate, the cutoff frequency is 24kHz) can be equally divided into 256 frequency bands. And calculate the amplitude of each frequency band to get the volume value of the audio data. In addition, the frequency band with the largest amplitude is acquired, and the frequency value corresponding to the frequency data is determined according to the sub-frequency domain audio data in the frequency band with the highest amplitude.
- volume value calculation or FFT calculation real-time calculation can be realized at the terminal.
- the terminal obtains frequency domain audio data in each period. Taking audio data with a sampling rate of 48kHz as an input source, for example, it will be calculated 48 times in one second.
- the terminal designs different mapping formulas according to different requirements.
- the mapping formula is the mapping relationship described in the embodiment of the present application.
- the input variable of the mapping type is a volume value or a frequency value
- the output is attribute information of the virtual object, such as color, scaling, and rotation angle. Taking the volume value of the audio data as the input and the scale of the virtual object as the output as an example, the following segmented mapping formula can be designed:
- x is the volume value of the audio data, and the range can be from 0 to 120 db, and f (x) is the scaling ratio of the virtual object.
- the mapping formula can be configured in 3 dimensions: 1) Configure the mapping formula according to the input type of the mapping type, such as the volume type or frequency value; 2) Configure the mapping formula according to the output type of the mapping type, and the output type can be Attribute information of different dimensions such as the scale, color, rotation angle, and offset of the virtual object; 3) The mapping formula is configured according to the correspondence between the input and output of the function.
- the virtual object when the decibel value is less than 50db, the virtual object maintains a default size of 1.0.
- the decibel value is greater than 50db, the scale of the virtual object increases as the decibel value increases.
- the decibel value is 120db, the scale is 4.0.
- FIG. 5 (a) is the default size, that is, the zoom ratio is 1.0; and FIG. 5 (b) is the effect when the zoom size is about 2.0.
- the terminal can collect real-time scene images through a camera.
- the real scene image may be a frame image in a video captured by the camera in real time.
- the terminal performs feature detection on the target object in the real scene image, such as face feature detection.
- the detection method may be: using the open source opencv or dlib face registration point SDK, or using the facial feature point detection SDK provided by Youtu, Shangtang, etc. for feature detection.
- the virtual object after the terminal adjusts the attribute information is fused to a fixed area of the target object in the real scene image (using a certain facial feature point of the target object as the anchor point), so that the virtual object can follow the human face in real time, and follow the audio data. Attribute information changes.
- the effect is that the virtual object can follow the human face in real time, and the size of the virtual object will change in real time with the volume value collected by the microphone or the volume value of the background music.
- S1316 Output a real scene image including a virtual object.
- the playability of the selfie / short video APP can be greatly increased.
- the virtual object will change in size, color, or rotation angle as the parameter value of the audio data changes, and the virtual object is added.
- the variety of changes has improved the interaction between users and virtual objects.
- FIG. 2 is a schematic flowchart of an image processing method according to an embodiment. It should be understood that although the steps in the flowchart of FIG. 2 are sequentially displayed in accordance with the directions of the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless explicitly stated in this document, the execution of these steps is not strictly limited, and these steps can be performed in other orders. Moreover, at least a part of the steps in FIG. 2 may include multiple sub-steps or stages. These sub-steps or stages are not necessarily performed at the same time, but may be performed at different times. The execution of these sub-steps or stages The order is not necessarily performed sequentially, but may be performed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
- an image processing device specifically includes: an audio data acquisition module 1402, an attribute information determination module 1404, a target object determination module 1406, and a fusion position determination module 1408.
- fusion module 1410 of which:
- An audio data acquisition module 1402 configured to acquire audio data corresponding to a real-life scene image collected in real time
- the attribute information determining module 1404 is configured to dynamically determine attribute information of the virtual object according to the audio data, and the attribute information is used to determine the visual state of the virtual object;
- a target object determining module 1406, configured to determine a target object from a real scene image
- a fusion position determination module 1408, configured to determine, according to the target object, the fusion position of the virtual object determined by the attribute information in the real scene image;
- a fusion module 1410 is configured to fuse a virtual object determined according to the attribute information to a real scene image according to a fusion position; the virtual object presents different visual states when the attribute information is different.
- the audio data acquisition module 1402 is further configured to collect real-time audio data corresponding to the real scene image from the current environment when the real scene image is collected in real time; or from the background audio of the real scene image collected in real time To read the audio data corresponding to the timestamp corresponding to the real scene image.
- the target object determination module 1406 is further configured to identify a biological feature from a real scene image; when the biological feature meets a preset condition, a biological object corresponding to the biological feature in the real scene image is determined as the target object.
- audio data corresponding to a real-time scene image collected in real time is acquired, and the attributes of the virtual object are dynamically determined through the audio data, thereby implementing control of the attributes of the virtual object according to the audio data.
- the fusion position of the virtual object in the real scene image is determined, and the virtual object determined based on the attribute information is fused into the real scene image according to the fusion position. Since the attribute information of the virtual object is controlled by audio data, When the audio data changes, the attribute information of the virtual object integrated into the real scene image also changes, which improves the interactivity.
- the attribute information determination module 1404 is further configured to obtain parameter values of audio data; determine a preset mapping relationship between the parameter values and the attribute information of the virtual object; and map the parameter values to virtual according to the preset mapping relationship.
- Object attribute information is further configured to obtain parameter values of audio data; determine a preset mapping relationship between the parameter values and the attribute information of the virtual object; and map the parameter values to virtual according to the preset mapping relationship.
- the attribute information determination module 1404 samples the audio data; quantizes and encodes the sampled result to obtain encoded audio data; and determines parameter values of the audio data according to the obtained encoded audio data.
- a preset mapping relationship between a parameter value and attribute information of a virtual object is determined.
- the attribute information of the corresponding virtual object can be obtained through the preset mapping relationship, thereby The adjustment of the virtual object is realized, so that the virtual object presents different visual states, and the diversity of the virtual object is improved.
- the parameter value includes a frequency value
- the attribute information determination module 1404 is further configured to convert the time-domain encoded audio data into frequency-domain audio data; segment the frequency-domain audio data to obtain multiple sub-frequency-domain audio data; Determine the amplitude of each sub-frequency domain audio data; select the sub-frequency domain audio data with the largest amplitude from each sub-frequency domain audio data; determine the corresponding frequency value of the audio data according to the selected sub-frequency domain audio data.
- the frequency domain audio data is segmented, and the frequency value can be determined from the sub-frequency domain audio data obtained after the segmentation.
- the frequency value can be used to adjust the virtual object; on the other hand, the segmentation
- the subsequent frequency domain audio data can reduce the calculation amount and speed up the calculation rate during the calculation process.
- the parameter value includes a volume value
- the attribute information determination module 1404 is further configured to determine the volume value according to the obtained encoded audio data; or, convert the time-domain encoded audio data into frequency-domain audio data; according to the frequency domain Audio data determines the volume value.
- the volume value is determined according to the obtained encoded audio data or the frequency domain audio data converted from the time-domain encoded audio data, and attribute information for determining the visual state of the virtual object is obtained.
- the volume value can be adjusted for virtual objects.
- the attribute information includes an attribute adjustment amount; the virtual object determined according to the attribute information is a virtual object obtained by adjusting corresponding attributes according to the attribute adjustment amount; the fusion module 1410 is further configured to determine the virtual object and its attributes. The attribute corresponding to the adjustment amount; adjust the attribute of the virtual object according to the attribute adjustment amount to obtain the virtual object after adjusting the attribute; and merge the virtual object after adjusting the attribute into the real scene image according to the fusion position.
- the attributes of the virtual object are adjusted by the attribute adjustment amount, and the virtual object after the adjustment of the attribute is fused to the real scene image according to the fusion position to obtain the virtual object that changes with the parameters of the audio data, so as to adjust the virtual object according to the audio data.
- the diversity of virtual objects has been improved and the user experience has been enhanced.
- the attribute information includes a target value of the attribute change; the virtual object determined according to the attribute information is a virtual object after the corresponding attribute is changed to the target value of the attribute change; the fusion module 1410 is further configured to determine a virtual object having, And the attribute corresponding to the attribute change target value; change the attribute of the virtual object to the attribute change target value to obtain the virtual object after the attribute change; and merge the virtual object after the attribute change into the real scene image according to the fusion position.
- the attributes of the virtual object are adjusted by the target value of the attribute change, and the virtual object after adjusting the attribute is merged into the real scene image according to the fusion position to obtain the virtual object that changes with the parameters of the audio data, so as to adjust the virtual object according to the audio data , Enhances the diversity of virtual objects, and enhances the user experience.
- the fusion position determination module 1408 is further configured to detect the features of the target object; find the features that match the virtual object with attributes from the detected features; and determine the virtual object with attributes based on the matched features. Fusion position in a realistic scene image.
- the fusion position of the virtual object in the real scene image is determined by the feature points of the target object, so that the virtual object is fused in the real scene image according to the fusion position, and the virtual object with a changed visual state is obtained, which improves The diversity of virtual objects has changed.
- FIG. 15 shows an internal structure diagram of a computer device in one embodiment.
- the computer device may specifically be a terminal in FIG. 1.
- the computer device includes the computer device including a processor, a memory, a network interface, an input device, and a display screen connected through a system bus.
- the memory includes a non-volatile storage medium and an internal memory.
- the non-volatile storage medium of the computer device stores an operating system, and may also store a computer program.
- the processor can implement an image processing method.
- a computer program may also be stored in the internal memory, and when the computer program is executed by the processor, the processor may execute the image processing method.
- the display screen of a computer device can be a liquid crystal display or an electronic ink display screen.
- the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball or a touchpad provided on the computer device casing. It can be an external keyboard, trackpad, or mouse.
- FIG. 15 is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer equipment to which the solution of the present application is applied.
- the specific computer equipment may be Include more or fewer parts than shown in the figure, or combine certain parts, or have a different arrangement of parts.
- the image processing apparatus provided in this application may be implemented in the form of a computer program, and the computer program may be run on a computer device as shown in FIG. 15.
- the memory of the computer device may store various program modules constituting the 14 device, such as the audio data acquisition module 1402, the attribute information determination module 1404, the target object determination module 1406, the fusion position determination module 1408, and the fusion module 1410 shown in FIG. .
- the computer program constituted by each program module causes the processor to execute the steps in the image processing method of each embodiment of the present application described in this specification.
- the computer device shown in FIG. 15 may execute S202 through the audio data acquisition module 1402 in the image processing apparatus shown in FIG. 14.
- the computer device may execute S204 through the attribute information determination module 1404.
- the computer device may execute S206 through the target object determination module 1406.
- the computer device may execute S208 through the fusion position determination module 1408.
- the computer device may execute S210 through the fusion module 1410.
- a computer device which includes a memory and a processor.
- the memory stores a computer program.
- the processor causes the processor to perform the following steps: Audio data; dynamically determining the attribute information of the virtual object according to the audio data, the attribute information is used to determine the visual state of the virtual object; determining the target object from the real scene image; and determining the virtual object determined by the attribute information in the real scene according to the target object.
- the fusion position in the image; the virtual object determined by the attribute information is fused to the real scene image according to the fusion position; the virtual object presents different visual states when the attribute information is different.
- the processor when the computer program is executed by the processor to obtain the audio data corresponding to the real-life scene image acquired in real time, the processor is caused to specifically perform the following steps: in real-time acquisition of the real scene image, it is collected in real time from the current environment Audio data corresponding to the real scene image; or, from background audio of the real scene image collected in real time, audio data corresponding to the time stamp corresponding to the real scene image is read.
- the processor when the computer program is executed by the processor to dynamically determine the attribute information of the virtual object based on the audio data, the processor is caused to specifically perform the following steps: obtaining parameter values of the audio data; determining parameter values and attribute information of the virtual object A preset mapping relationship between the two; according to the preset mapping relationship, parameter values are mapped to attribute information of the virtual object.
- the processor when the computer program is executed by the processor to obtain the parameter values of the audio data, the processor is caused to specifically perform the following steps: sampling the audio data; quantizing and encoding the result of the sampling to obtain encoded audio data; Parameter values of audio data are determined based on the obtained encoded audio data.
- the parameter value includes a frequency value.
- the processor is caused to specifically perform the following steps: encode the audio data in the time domain. Convert to frequency domain audio data; segment the frequency domain audio data to obtain multiple sub-frequency domain audio data; determine the amplitude of each sub-frequency domain audio data; select the sub-frequency domain audio data with the largest amplitude from each sub-frequency domain audio data ; Determine the frequency value corresponding to the audio data according to the selected sub-frequency domain audio data.
- the parameter value includes a volume value; when the computer program is executed by the processor to determine the parameter value of the audio data according to the obtained encoded audio data, the processor is caused to specifically perform the following steps: according to the obtained encoded audio data Determine the volume value; or convert the time-domain encoded audio data into frequency-domain audio data; determine the volume value according to the frequency-domain audio data.
- the attribute information includes an attribute adjustment amount; the virtual object determined by the attribute information is a virtual object adjusted by the corresponding attribute according to the attribute adjustment amount; the computer program is executed by the processor to fuse the virtual object determined by the attribute information according to the fusion.
- the processor is caused to perform the following steps: determining the attributes of the virtual object and corresponding to the attribute adjustment amount; adjusting the attributes of the virtual object according to the attribute adjustment amount, and obtaining the virtual Object; the virtual object after adjusting the properties is fused to the real scene image according to the fusion position.
- the attribute information includes the attribute change target value; the virtual object determined according to the attribute information is a virtual object that changes the corresponding attribute to the attribute change target value; the computer program is executed by the processor and the virtual object is determined according to the attribute information.
- the processor is caused to perform the following steps: determining the attributes of the virtual object and corresponding to the attribute change target value; changing the attributes of the virtual object to the attribute change target value, The virtual object after changing the attributes is obtained; the virtual object after changing the attributes is fused to the real scene image according to the fusion position.
- the processor when the computer program is executed by the processor to determine the target object from the real scene image, the processor causes the processor to specifically perform the following steps: identify the biometric feature from the real scene image; when the biometric feature meets a preset condition, Then, the biological object corresponding to the biological feature in the real scene image is determined as the target object.
- the processor when the computer program is executed by the processor to determine the fusion position of the virtual object with attributes in the real scene image according to the target object, the processor is caused to specifically perform the following steps: detecting the characteristics of the target object; Among the detected features, find features that match the virtual object with attributes; based on the matched features, determine the fusion position of the virtual object with attributes in the real scene image.
- a computer-readable storage medium storing a computer program.
- the processor causes the processor to perform the following steps: fetch audio data corresponding to a real-time scene image acquired in real time;
- the audio data dynamically determines the attribute information of the virtual object, and the attribute information is used to determine the visual state of the virtual object;
- the target object is determined from the real scene image; according to the target object, the fusion of the virtual object determined by the attribute information in the real scene image is determined Position; the virtual object determined by the attribute information is fused into the real scene image according to the fusion position; the virtual object presents different visual states when the attribute information is different.
- the processor when the computer program is executed by the processor to obtain the audio data corresponding to the real-life scene image acquired in real time, the processor is caused to specifically perform the following steps: in real-time acquisition of the real scene image, it is collected in real time from the current environment Audio data corresponding to the real scene image; or, from background audio of the real scene image collected in real time, audio data corresponding to the time stamp corresponding to the real scene image is read.
- the processor when the computer program is executed by the processor to dynamically determine the attribute information of the virtual object based on the audio data, the processor is caused to specifically perform the following steps: obtaining parameter values of the audio data; determining parameter values and attribute information of the virtual object A preset mapping relationship between the two; according to the preset mapping relationship, parameter values are mapped to attribute information of the virtual object.
- the processor when the computer program is executed by the processor to obtain the parameter values of the audio data, the processor is caused to specifically perform the following steps: sampling the audio data; quantizing and encoding the result of the sampling to obtain encoded audio data; Parameter values of audio data are determined based on the obtained encoded audio data.
- the parameter value includes a frequency value.
- the processor is caused to specifically perform the following steps: encode the audio data in the time domain. Convert to frequency domain audio data; segment the frequency domain audio data to obtain multiple sub-frequency domain audio data; determine the amplitude of each sub-frequency domain audio data; select the sub-frequency domain audio data with the largest amplitude from each sub-frequency domain audio data ; Determine the frequency value corresponding to the audio data according to the selected sub-frequency domain audio data.
- the parameter value includes a volume value; when the computer program is executed by the processor to determine the parameter value of the audio data according to the obtained encoded audio data, the processor is caused to specifically perform the following steps: according to the obtained encoded audio data Determine the volume value; or convert the time-domain encoded audio data into frequency-domain audio data; determine the volume value according to the frequency-domain audio data.
- the attribute information includes an attribute adjustment amount; the virtual object determined by the attribute information is a virtual object adjusted by the corresponding attribute according to the attribute adjustment amount; the computer program is executed by the processor to fuse the virtual object determined by the attribute information according to the fusion.
- the processor is caused to perform the following steps: determining the attributes of the virtual object and corresponding to the attribute adjustment amount; adjusting the attributes of the virtual object according to the attribute adjustment amount, and obtaining the virtual Object; the virtual object after adjusting the properties is fused to the real scene image according to the fusion position.
- the attribute information includes the attribute change target value; the virtual object determined according to the attribute information is a virtual object that changes the corresponding attribute to the attribute change target value; the computer program is executed by the processor and the virtual object is determined according to the attribute information.
- the processor is caused to perform the following steps: determining the attributes of the virtual object and corresponding to the attribute change target value; changing the attributes of the virtual object to the attribute change target value, The virtual object after changing the attributes is obtained; the virtual object after changing the attributes is fused to the real scene image according to the fusion position.
- the processor when the computer program is executed by the processor to determine the target object from the real scene image, the processor causes the processor to specifically perform the following steps: identify the biometric feature from the real scene image; when the biometric feature meets a preset condition, Then, the biological object corresponding to the biological feature in the real scene image is determined as the target object.
- the processor when the computer program is executed by the processor to determine the fusion position of the virtual object with attributes in the real scene image according to the target object, the processor is caused to specifically perform the following steps: detecting the characteristics of the target object; Among the detected features, find features that match the virtual object with attributes; based on the matched features, determine the fusion position of the virtual object with attributes in the real scene image.
- the processor when the computer program is executed by the processor, the processor causes the processor to specifically perform the following steps: extract audio characteristics of the audio data; when the audio characteristics meet the first trigger condition, perform at least one of the following: adding a virtual object ; Switch virtual objects; switch the type of visual state.
- the processor when the computer program is executed by the processor, the processor is caused to specifically perform the following steps: identify according to the audio data to obtain the recognition result; determine the type of dynamic effect that matches the recognition result; determine according to the type of dynamic effect and attribute information The visual state presented by the virtual object; the visual state matches the type of dynamic effect.
- Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory.
- Volatile memory can include random access memory (RAM) or external cache memory.
- RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous chain (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
- SRAM static RAM
- DRAM dynamic RAM
- SDRAM synchronous DRAM
- DDRSDRAM dual data rate SDRAM
- ESDRAM enhanced SDRAM
- SLDRAM synchronous chain (Synchlink) DRAM
- SLDRAM synchronous chain (Synchlink) DRAM
- Rambus direct RAM
- DRAM direct memory bus dynamic RAM
- RDRAM memory bus dynamic RAM
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Multimedia (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- General Health & Medical Sciences (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Databases & Information Systems (AREA)
- Library & Information Science (AREA)
- Oral & Maxillofacial Surgery (AREA)
- Evolutionary Computation (AREA)
- General Engineering & Computer Science (AREA)
- Software Systems (AREA)
- Data Mining & Analysis (AREA)
- Artificial Intelligence (AREA)
- Medical Informatics (AREA)
- Computing Systems (AREA)
- Human Computer Interaction (AREA)
- Computer Graphics (AREA)
- Computer Hardware Design (AREA)
- Life Sciences & Earth Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Evolutionary Biology (AREA)
- Signal Processing (AREA)
- Processing Or Creating Images (AREA)
Abstract
本申请涉及一种图像处理方法、装置、存储介质和计算机设备,方法包括:获取与实时采集的现实场景图像对应的音频数据;根据音频数据动态确定虚拟对象的属性信息,属性信息用于确定虚拟对象的视觉状态;从现实场景图像中确定目标对象;根据目标对象,确定按属性信息确定的虚拟对象在现实场景图像中的融合位置;将按属性信息确定的虚拟对象按融合位置融合到现实场景图像;虚拟对象在属性信息不同时呈现不同的视觉状态。本申请提供的方案可以提升用户与虚拟对象之间的交互性。
Description
本申请要求于2018年7月4日提交、申请号为201810723144.8、发明名称为“图像处理方法、装置、存储介质和计算机设备”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本申请涉及图像处理技术领域,特别是涉及一种图像处理方法、装置、存储介质和计算机设备。
随着图像处理技术和计算机技术的快速发展,出现了各式各样的用于录制视频的客户端,比如用户可通过客户端录制虚拟和现实相结合的视频。
目前在录制视频时,用户可在客户端的录制界面中自由选取虚拟对象,将这些虚拟对象添加到视频对应图像帧的相应位置,使得虚拟对象可跟随图像帧中运动目标的移动而移动。
然而,通过上述方式所录制的视频,虚拟对象仅能够跟随运动目标的移动而移动,交互性差。
发明内容
基于此,提供一种图像处理方法、装置、存储介质和计算机设备,能够解决虚拟对象仅能够跟随运动目标的移动而移动,使交互性差的技术问题。
一种图像处理方法,其特征在于,应用于图像处理系统,所述方法包括:
获取与实时采集的现实场景图像对应的音频数据;
根据所述音频数据动态确定虚拟对象的属性信息,所述属性信息用于确定所述虚拟对象的视觉状态;
从所述现实场景图像中确定目标对象;
根据所述目标对象,确定按所述属性信息确定的虚拟对象在所述现实场景图像中的融合 位置;
将按所述属性信息确定的虚拟对象按所述融合位置融合到所述现实场景图像;所述虚拟对象在属性信息不同时呈现不同的视觉状态。
一种图像处理装置,其特征在于,包括:
音频数据获取模块,用于获取与实时采集的现实场景图像对应的音频数据;
属性信息确定模块,用于根据所述音频数据动态确定虚拟对象的属性信息,所述属性信息用于确定所述虚拟对象的视觉状态;
目标对象确定模块,用于从所述现实场景图像中确定目标对象;
融合位置确定模块,用于根据所述目标对象,确定按所述属性信息确定的虚拟对象在所述现实场景图像中的融合位置;
融合模块,用于将按所述属性信息确定的虚拟对象按所述融合位置融合到所述现实场景图像;所述虚拟对象在属性信息不同时呈现不同的视觉状态。
一种存储介质,其特征在于,存储有计算机程序,所述计算机程序被处理器执行时,使得所述处理器执行上述的图像处理方法的步骤。
一种计算机设备,其特征在于,包括存储器和处理器,所述存储器存储有计算机程序,所述计算机程序被所述处理器执行时,使得所述处理器执行如下操作:
获取与实时采集的现实场景图像对应的音频数据;
根据所述音频数据动态确定虚拟对象的属性信息,所述属性信息用于确定所述虚拟对象的视觉状态;
从所述现实场景图像中确定目标对象;
根据所述目标对象,确定按所述属性信息确定的虚拟对象在所述现实场景图像中的融合位置;
将按所述属性信息确定的虚拟对象按所述融合位置融合到所述现实场景图像;所述虚拟对象在属性信息不同时呈现不同的视觉状态。
上述的图像处理方法、装置、存储介质和计算机设备,获取与实时采集的现实场景图像对应的音频数据,通过音频数据动态确定虚拟对象的属性,从而实现了根据音频数据对虚拟对象属性的控制。通过现实场景图像中的目标对象,确定虚拟对象在现实场景图像中的融合 位置,按照融合位置将根据属性信息确定的虚拟对象融合到现实场景图像中,由于虚拟对象的属性信息由音频数据控制,当音频数据发生变化时,融合到现实场景图像中的虚拟对象的属性信息也随之产生变化,提高了交互性。
图1为一个实施例中应用图像处理方法的系统结构图;
图2为一个实施例中图像处理方法的流程示意图;
图3为一个实施例中将虚拟对象融合到现实场景图像的示意图;
图4为一个实施例中将虚拟对象融合到现实场景图像的示意图;
图5为一个实施例中根据音频数据确定属性信息,将具有该属性信息的虚拟对象融合到现实场景图像的示意图;
图6为一个实施例中根据音频数据的参数值确定虚拟对象的属性信息的步骤的流程示意图;
图7为一个实施例中对音频数据进行抽样、量化和编码的示意图;
图8为一个实施例中根据编码音频数据确定频率值的步骤的流程示意图;
图9为一个实施例中按照属性调整量调整虚拟对象,并将调整后的虚拟对象融合到现实场景图像的步骤的流程示意图;
图10为一个实施例中按照属性变化目标值调整虚拟对象,并将调整后的虚拟对象融合到现实场景图像的步骤的流程示意图;
图11为一个实施例中根据目标对象的特征确定虚拟对象在现实场景图像的融合位置的步骤的流程示意图;
图12为一个实施例中目标对象脸部特征点的示意图;
图13为另一个实施例中图像处理方法的流程示意图;
图14为一个实施例中图像处理装置的结构框图;
图15为一个实施例中计算机设备的结构框图。
为了使本申请的目的、技术方案及优点更加清楚明白,以下结合附图及实施例,对本申请进行进一步详细说明。应当理解,此处所描述的具体实施例仅仅用以解释本申请,并不用于限定本申请。
图1为一个实施例中图像处理方法的应用环境图。参照图1,该图像处理方法应用于图像处理系统。该图像处理系统可以是一个终端或多个终端的组合,该终端可以是智能手机、电脑或其它可支持AR(Augmented Reality,增强现实)技术的设备。如图1所示,图像处理系统可以包括:摄像头、场景产生器、图像合成器和显示器;其中:
摄像头,用于获取目标对象对应环境的现实场景图像,将获取到的现实场景图像发送至图像合成器,以与增强现实模型的虚拟对象进行合成操作。
场景产生器,用于根据现实场景图像中目标对象的位置信息确定虚拟对象的融合位置,如通过分析目标对象的特征确定虚拟对象的融合位置,然后将虚拟对象发送至图像合成器。
图像合成器,用于将关于目标对象的现实场景图像和虚拟对象,按照融合位置进行融合,将融合的结果输出至显示屏。
显示器,用于将图像合成器输出的融合图像进行显示,形成目标对象和用于增强现实模型的虚拟对象共同显示的效果。
如图2所示,在一个实施例中,提供了一种图像处理方法。本实施例以该方法应用于上述图1中的终端来举例说明。参照图2,该图像处理方法具体包括如下步骤:
S202,获取与实时采集的现实场景图像对应的音频数据。
其中,现实场景可以指用户通过某种媒介所观看到的现实画面,在该现实画面中至少包括以下之一:人物、自然风景、人文风景和人类的智慧作品等。人类的智慧作品指的是人类通过劳动和智慧所创造的作品。
举例来说,现实场景为用户通过裸眼所观看到的人与自然的画面,或者用户通过3D眼镜所观看到的立体电影的画面。现实场景图像可以是通过终端采集到的关于现实场景的图像。例如,现实场景图像为通过图1中的摄像头实时采集到的现实场景的图像。终端采集多个现实场景图像后,将多个现实场景图像按照采集时间进行组合,可得到视频。
音频数据为时域的音频信号,音频数据中携带有语音、音乐和音效的有规律的声波的频 率、幅度变化信息等。根据声波的特征,可把音频数据分类为规则音频和不规则音频。其中规则音频又可以分为语音、音乐和音效。音频数据的来源可以是终端从外界采集所得,也可以是从与现实场景图像的背景音频中读取所得。
在一个实施例中,S202具体可以包括:在实时采集现实场景图像时,终端从当前环境中实时采集与现实场景图像对应的音频数据;或者,终端从实时采集的现实场景图像的背景音频中,读取与现实场景图像所对应时间戳对应的音频数据。
例如,在实时采集现实场景图像时,终端通过麦克风采集当前环境中的音乐或说话的声音。或者,由于在开发过程中,开发人员设置在实时采集的现实场景图像时自动播放背景音乐,终端实时采集现实场景图像时会生成现实场景图像对应的时间戳,在自动播放的背景音乐中读取与现实场景图像所对应时间戳对应的音频数据。或者,在实时采集的现实场景图像时,音乐客户端同时在播放音乐,终端则实时采集现实场景图像时会生成现实场景图像对应的时间戳,在播放的音乐中读取与现实场景图像所对应时间戳对应的音频数据。
S204,根据音频数据动态确定虚拟对象的属性信息,该属性信息用于确定虚拟对象的视觉状态。
其中,虚拟对象可以包括图像素材,该图像素材如静态贴纸或动态贴纸等。此外,虚拟对象还可以包括用于增强目标对象显示效果的虚拟道具。例如,虚拟道具可以是用于装扮目标对象的各种虚拟挂件和虚拟背景。
虚拟对象可以包括一个或多个属性,每个属性具有一个属性信息,该属性信息用于确定虚拟对象的视觉状态。
虚拟对象的属性信息可以包括以下至少一种:属性调整量和属性变化目标值。属性调整量包括虚拟对象的缩放比例、旋转角度和偏移量等。属性变化目标值包括虚拟对象的颜色RGB值等。属性调整量用于表示虚拟对象的属性的调整幅度,根据属性调整量能够对虚拟对象相应的属性进行调整,确定调整后的属性。属性变化目标值用于表示虚拟对象的属性的目标属性值,根据属性变化目标值能够对虚拟对象相应的属性进行调整,以使调整后的属性为该属性变化目标值。
需要说明的是,所确定虚拟对象的属性信息为实时采集现实场景图像当下的属性信息,则对于不同的现实场景图像,所确定虚拟对象的属性信息可能会不同。例如,假设用户拍摄 一个具有AR效果的视频,若视频中的第i帧现实场景图像为当前的现实场景图像,第i帧现实场景图像对应时刻的虚拟对象的属性信息为当前的属性信息。视频中的第i-1帧现实场景图像为上一时刻的现实场景图像,第i-1帧现实场景图像对应时刻的虚拟对象的属性信息为上一时刻的属性信息。其中,i为大于等于1的正整数。
在一个实施例中,终端根据音频数据动态确定虚拟对象的属性信息,根据确定的属性信息调整虚拟对象,获得按属性信息确定的虚拟对象。
在一个实施例中,音频数据具有音量值、频率值和音色等参数值。终端根据音频数据动态确定虚拟对象的属性信息。当终端将虚拟对象的原始属性信息更新为所确定的属性信息时,便可实现对虚拟对象的调整。例如,终端根据音频数据的频率值来调整虚拟对象的缩放比例,或调整虚拟对象的颜色RGB值。
S206,从现实场景图像中确定目标对象。
其中,目标对象可以是自然界中具有生命的物体,如人类、动物、植物等。
在一个实施例中,S206具体可以包括:从现实场景图像中识别生物特征;当生物特征满足预设条件时,则将现实场景图像中与生物特征对应的生物对象确定为目标对象。
其中,生物特征可以是生物的轮廓特征,也可以是生物的细节特征,如人的脸部特征等。
在一个实施例中,预设条件可以包括预设生物特征或生物特征的预设完整度阈值。终端从现实场景图像中识别生物特征;当生物特征满足预设生物特征,和/或生物特征的完整度达到预设完整度阈值时,则将现实场景图像中与生物特征对应的对象确定为目标对象。
即,当生物特征满足预设生物特征时,则将现实场景图像中与生物特征对应的对象确定为目标对象。或者,当生物特征的完整度达到预设完整度阈值时,则将现实场景图像中与生物特征对应的对象确定为目标对象。或者,当生物特征满足预设生物特征,且生物特征的完整度达到预设完整度阈值时,则将现实场景图像中与生物特征对应的对象确定为目标对象。
S208,根据目标对象,确定按属性信息确定的虚拟对象在现实场景图像中的融合位置。
其中,融合位置可以指:在融合到现实场景图像时,虚拟对象的中心点或关键位置点在现实场景图像中所处的位置。融合位置可以是一个位置点,也可以是多个位置点。针对不同的虚拟对象或同一虚拟对象的不同部位,可以确定不同的融合位置,例如不同的虚拟对象可以融合在目标对象的不同部位,同一个虚拟对象的不同部位可以融合在目标对象的不同部位。
在一个实施例中,终端检测目标对象的特征,从目标对象的特征中选取与虚拟对象匹配的特征,根据所选取的特征确定按属性信息确定的虚拟对象在现实场景图像中的融合位置。
其中,虚拟对象匹配的特征可以根据虚拟对象的类型确定,不同类型的虚拟对象可以匹配相同的特征或不同的特征。作为一个示例,可以设置多种类型的虚拟对象,并为每种类型的虚拟对象设置匹配的特征,从而建立虚拟对象与特征之间的匹配关系,后续即可根据该匹配关系确定虚拟对象匹配的特征。例如AR潜水眼镜匹配的特征为眼部特征,AR兔子装扮的虚拟道具匹配的特征为嘴部特征。
根据所选取的特征确定按属性信息确定的虚拟对象在现实场景图像中的融合位置,可以包括:将所选取的特征所在的位置确定为虚拟对象在现实场景图像中的融合位置。
作为一个示例,如图3所示,当虚拟对象为AR潜水眼镜的虚拟道具时,终端检测目标对象的特征,从检测到的特征中确定眼部特征,根据眼部特征确定虚拟对象的融合位置为用户的眼部位置。
作为另一个示例,如图4所示,当虚拟对象为AR兔子装扮的虚拟道具时,对于AR兔子装扮中的兔牙虚拟道具,终端检测目标对象的特征,从检测到的特征中确定嘴部特征,根据嘴部特征确定虚拟对象的融合位置为用户的嘴部位置。对于AR兔子装扮中的耳朵虚拟道具,终端检测目标对象的特征,从检测到的特征中确定头部特征(如头发),根据头部特征确定虚拟对象的融合位置为用户的头部位置。
或者,对于同一虚拟对象来说,该虚拟对象中可以包括多个部位,则不同的部位可以匹配不同的特征。作为一个示例,当虚拟对象为雨衣虚拟道具时,雨衣虚拟道具中的头部匹配的特征为头部特征,雨衣虚拟道具中的胳膊匹配的特征为胳膊特征。
在一个实施例中,终端按属性信息调整虚拟对象,从而使虚拟对象进行尺寸的缩放、或更换颜色RGB值、或改变旋转角度等。需要说明的是,终端按属性信息调整的虚拟对象,是以上一帧融合后的现实场景图像中的虚拟对象为参考,从而实现对当前虚拟对象进行缩放、或更换颜色RGB值、或改变旋转角度等操作。
S210,将按属性信息确定的虚拟对象按融合位置融合到现实场景图像;虚拟对象在属性信息不同时呈现不同的视觉状态。
在一个实施例中,终端确定虚拟对象的中心点或关键位置点,将虚拟对象的中心点或关 键位置点融合至与融合位置匹配的区域,从而将虚拟对象融合到现实场景图像中。其中,中心点或关键位置点用于确定虚拟对象所要融合的区域,用于与对应的融合位置匹配,即在融合时,将虚拟对象的中心点置于对应的融合位置,或者将虚拟对象的关键位置点置于对应的融合位置。
如图3所示,p表示虚拟对象的关键位置点,该关键位置点为AR潜水眼镜的虚拟道具的三个位置。在将AR潜水眼镜的虚拟道具融合到现实场景图像中时,根据图中3中的p点将AR潜水眼镜的虚拟道具对准目标对象的两眼和鼻子,从而实现AR潜水眼镜的虚拟道具的准确融合,具体效果如图3(b)所示。
如图4所示,q表示虚拟对象的关键位置点,分别位于AR兔子装扮的兔牙虚拟道具中的上部,以及,两个AR兔子装扮的耳朵虚拟道具的下部。在将AR兔子装扮的虚拟道具融合到现实场景图像时,根据关键位置点将AR兔子装扮的虚拟道具融合到目标对象的嘴唇部位和头部,从而实现AR兔子装扮的虚拟道具的准确融合,具体效果如图4(b)所示。
虚拟对象在属性信息不同时呈现不同的视觉状态,举例来说,假设用户拍摄一个具有AR效果的视频,若视频中的第i帧现实场景图像为当前的现实场景图像,第i帧现实场景图像对应时刻的虚拟对象的属性信息为当前的属性信息。视频中的第i-1帧现实场景图像为上一时刻的现实场景图像,第i-1帧现实场景图像对应时刻的虚拟对象的属性信息为上一时刻的属性信息。其中,i为大于等于1的正整数。当各帧现实场景图像和对应属性信息的虚拟对象按照时间组合,那么可获得一个具有AR效果的视频,在该视频中,虚拟对象的属性信息随着音频数据的变化而发生变化。如音频数据的音量值或频率值发生变化,虚拟对象的尺寸、或颜色RGB值、或方位随之发生变化。
作为一个示例,如图5所示,假设图5(a)为上一时刻融合后的现实场景图像,m为原始尺寸的虚拟对象。当所获取的音频数据的参数值发生变化,如音量大小发生变化,或频率值发生变化,那么,对应的虚拟对象的尺寸也发生变化,变化后的虚拟对象如图5(b)中的n所示,n为尺寸放大了的虚拟对象。
上述实施例中,获取与实时采集的现实场景图像对应的音频数据,通过音频数据动态确定虚拟对象的属性,从而实现了根据音频数据对虚拟对象属性的控制。通过现实场景图像中的目标对象,确定虚拟对象在现实场景图像中的融合位置,按照融合位置将根据属性信息确 定的虚拟对象融合到现实场景图像中,由于虚拟对象的属性信息由音频数据控制,当音频数据发生变化时,融合到现实场景图像中的虚拟对象的属性信息也随之产生变化,提高了交互性。
在一个实施例中,如图6所示,S204具体可以包括:
S602,获取音频数据的参数值。
其中,音频数据的参数值包括音频数据的音量值、频率值和音色等。音量值可以是以下任一种:平均频率值、最大音量值或最小音量值。频率值可以是以下任一种:平均频率值、最大频率值或最小频率值。
具体地,终端通过分析音频数据,获得音频数据中的音量值、频率值和音色等参数值。
在一个实施例中,S602具体可以包括:对音频数据进行抽样;将抽样所得的结果量化并编码,获得编码音频数据;根据所获得的编码音频数据确定音频数据的参数值。
具体地,终端采用PCM((Pulse Code Modulation,脉冲编码调制)的方式,对所采集的时间上连续的时域音频数据进行抽样、量化和编码,获得二进制的编码音频数据。终端根据所获得的编码音频数据确定音频数据的参数值,如确定音频数据的音量值。
作为一个示例,如图7所示,对音频数据U(t)进行抽样,使时间上连续的时域音频数据离散化。将离散化的音频数据进行量化,获得M进制的量化音频数据,其中,M为大于2的正整数。将量化后的音频数据进行编码,得到二进制的编码音频数据。
S604,确定参数值与虚拟对象的属性信息之间的预设映射关系。
在一个实施例中,不同类型的参数值对应不同的预设映射关系,终端根据参数值的类型,确定参数值与虚拟对象的属性信息之间的预设映射关系。其中,类型包括音量类型、频率类型和音色类型。与类型对应的,参数值包括音量值、频率值和音色。
例如,以参数值为音量值,以属性信息为缩放比例为例,终端确定音量与虚拟对象的缩放比例之间的预设映射关系,该预设映射关系可以是函数关系式,如下所示:
其中,x为音频数据的音量值,范围可以是0~120分贝(db),f(x)为虚拟对象的缩放比例。
S606,根据预设映射关系,将参数值映射为虚拟对象的属性信息。
在一个实施例中,终端确定预设映射关系时,将参数值作为变量输入该预设映射关系,得到虚拟对象的对应属性信息。例如,当音量值为40db时,将音量值40输入函数f(x),得到的属性信息为1的缩放比例,即对原始的虚拟对象不做任何放大或缩小处理。又例如,当音量值为120db时,将音量值120输入函数f(x),得到的属性信息为4的缩放比例,即将原始的虚拟对象放大到原来的4倍。可以看出,虚拟对象的属性信息随着音量数据的参数值变化而变化,从而根据音量数据的参数值实现对虚拟对象的控制,从而使虚拟对象呈现不同的视觉状态。
上述实施例中,确定参数值与虚拟对象的属性信息之间的预设映射关系,当获得对应的参数值的大小时,便可以通过该预设映射关系获得对应的虚拟对象的属性信息,从而实现对虚拟对象的调整,使虚拟对象呈现不同的视觉状态,提升了虚拟对象的多样化变化。
在一个实施例中,虚拟对象的属性信息可以通过音频数据的参数值确定,该参数值可以是频率值或音量值,当参数值为频率值时;如图8所示,上述根据所获得的编码音频数据确定音频数据的参数值的步骤,具体可以包括:
S802,将时域的编码音频数据转换为频域音频数据。
在一个实施例中,编码音频数据为离散化的音频信号。终端根据离散傅里叶变换,将编码音频数据转换为频域音频数据。其中,频域音频数据包含有音频数据的幅值(即音量值)、频率值和相位。
S804,将频域音频数据分段,获得多个子频域音频数据。
在一个实施例中,终端按照设定的步长,对频域音频数据进行分段,获得多个子频域音频数据。例如,使用512点傅里叶变换时,最多可以将0到截止频率(若抽样率48kHz时,截止频率为24kHz)的频段等分为256个频段,然后执行S806,即确定各频段内的振幅。
在一个实施例中,终端按照不等分的方式,将频域音频数据划分为不等长的多个频段,获得多个子频域音频数据。
S806,确定各子频域音频数据的振幅。
在一个实施例中,各子频域音频数据中包含了幅值、频率值和相位,终端确定各子频域音频数据中的振幅,从而获取到各子频域音频数据中的音量值。
S808,从各子频域音频数据中选取振幅最大的子频域音频数据。
其中,音频数据的振幅大表示音频数据的功率较大,对于所获取的音频数据而言,当音频数据的功率大,对应的有用信号多,当音频数据的功率小,对应的有用信号少。例如,终端通过麦克风采集音频数据,当所采集的音频数据的功率小,说明当前所采集的音频数据可能为噪声信号,因此,可以选取振幅最大的子频域音频数据。
具体地,当确定各子频域音频数据的振幅时,将各子频域音频数据之间的振幅进行比较,获取振幅最大的子频域音频数据。
在一个实施例中,终端对各子频域音频数据按照振幅的大小进行排列,从排列的各子频域音频数据中选取振幅最大的子频域音频数据。
S810,按照所选取的子频域音频数据确定音频数据对应的频率值。
上述实施例中,对频域音频数据进行分段,可以通过分段后所得子频域音频数据确定频率值,一方面,通过该频率值可以实现对虚拟对象的调整,另一方面,分段后的频域音频数据,在计算过程中可降低计算量,加快计算速率。
在一个实施例中,虚拟对象的属性信息可以通过音频数据的参数值确定,该参数值可以是频率值或音量值,当参数值为音量值时,根据所获得的编码音频数据确定音频数据的参数值的步骤,具体可以包括:根据所获得的编码音频数据确定音量值;或者,将时域的编码音频数据转换为频域音频数据;根据频域音频数据确定音量值。
时域的编码音频数据的幅值可以表示音频数据的音量值。在一个实施例中,终端将时域的编码音频数据的幅值,确定为音频数据的音量值。
在一个实施例中,编码音频数据为离散化的音频信号。终端根据离散傅里叶变换,将编码音频数据转换为频域音频数据。其中,频域音频数据包含有音频数据对应波形的幅值(即音量值)、频率值和相位。终端将频域音频数据中的幅值确定为音频数据的音量值。
在一个实施例中,终端将编码音频数据转换为频域音频数据后,按照设定的步长对频域音频数据进行分段,获得多个子频域音频数据。终端根据各子频域音频数据确定对应的幅值,将最大的幅值确定为音频数据的音量值,或者,将平均幅值确定为音频数据的音量值。
上述实施例中,根据所获得的编码音频数据,或者根据由时域的编码音频数据转换的频域音频数据这两种方式确定音量值,得到用于确定虚拟对象视觉状态的属性信息,通过该音 量值可以实现对虚拟对象的调整。
在一个实施例中,虚拟对象的属性信息可以包括以下至少一种:属性调整量和属性变化目标值。属性调整量可以包括虚拟对象的缩放比例、旋转角度和偏移量等,属性变化目标值可以包括虚拟对象的颜色RGB值。当属性信息为属性调整量时,上述的按属性信息确定的虚拟对象,是按属性调整量调整相应属性后的虚拟对象,如图9所示,S210具体可以包括:
S902,确定虚拟对象所具有的、且与属性调整量对应的属性。
其中,属性指的是虚拟对象所拥有的属性,包括缩放、颜色、旋转和偏移等。与属性调整量对应的属性包括缩放、旋转和偏移等。对应的,属性调整量为属性对应的具体值。
在一个实施例中,终端根据音频数据的参数值,确定虚拟对象所具有的、且与属性调整量对应的属性。
S904,按照属性调整量调整虚拟对象的属性,得到调整属性后的虚拟对象。
在一个实施例中,终端根据映射关系,确定与音频数据的参数值对应的属性调整量后,按照属性调整量调整虚拟对象,使虚拟对象的属性发生变化,得到调整属性后的虚拟对象。
例如,若属性调整量为缩放比例,根据缩放比例调整虚拟对象的尺寸大小,获得调整尺寸后的虚拟对象。
S906,将调整属性后的虚拟对象按融合位置融合到现实场景图像。
上述实施例中,通过属性调整量调整虚拟对象的属性,将调整属性后的虚拟对象按融合位置融合到现实场景图像,获得随音频数据的参数变化的虚拟对象,实现根据音频数据调整虚拟对象,提升了虚拟对象的多样化变化,增强用户的体验。
在一个实施例中,虚拟对象的属性信息可以包括以下至少一种:属性调整量和属性变化目标值。属性调整量可以包括虚拟对象的缩放比例、旋转角度和偏移量等,属性变化目标值可以包括虚拟对象的颜色。当属性信息为属性变化目标值时,上述的按属性信息确定的虚拟对象,是将相应属性变化至属性变化目标值后的虚拟对象,如图10所示,S210具体可以包括:
S1002,确定虚拟对象所具有的、且与属性变化目标值对应的属性。
其中,与属性变化目标值对应的属性包括虚拟对象的颜色。对应的,属性变化目标值为属性对应的具体值,如颜色RGB值。
在一个实施例中,终端根据音频数据的参数值,确定虚拟对象所具有的、且与属性变化目标值对应的属性。
S1004,将虚拟对象的属性变化至属性变化目标值,得到属性变化后的虚拟对象。
在一个实施例中,终端根据映射关系,确定与音频数据的参数值对应的属性变化目标值后,按照属性变化目标值调整虚拟对象,使虚拟对象的属性发生变化,得到调整属性后的虚拟对象。
例如,若属性变化目标值为目标颜色RGB值,根据目标颜色RGB值调整虚拟对象的显示颜色,使虚拟对象的原始显示颜色调整为目标颜色RGB值所对应的颜色,如原始为红色,通过目标颜色RGB值的调整后获得蓝色的虚拟对象。
S1006,将属性变化后的虚拟对象按融合位置融合到现实场景图像。
上述实施例中,通过属性变化目标值调整虚拟对象的属性,将调整属性后的虚拟对象按融合位置融合到现实场景图像,获得随音频数据的参数变化的虚拟对象,实现根据音频数据调整虚拟对象,提升了虚拟对象的多样化变化,增强用户的体验。
在一个实施例中,当属性信息为第一属性的属性调整量和第二属性的属性变化目标值时,上述的按属性信息确定的虚拟对象,是按属性调整量调整第一属性,且将相应的第二属性变化至属性变化目标值后的虚拟对象。S210具体可以包括:确定虚拟对象所具有的、且与属性调整量对应的第一属性,确定虚拟对象所具有的、且与属性变化目标值对应的第二属性,按照属性调整量调整虚拟对象的第一属性,将虚拟对象的第二属性变化至属性变化目标值,得到属性变化后的虚拟对象,将属性变化后的虚拟对象按融合位置融合到现实场景图像。
在一个实施例中,如图11所示,S208具体可以包括:
S1102,检测目标对象的特征。
具体地,终端通过特征点检测方式,如级联回归CNN(Convolutional Neural Networks,卷积神经网络)、或Dlib、Libfacedetect、或Seetaface等特征点检测方式检测目标对象的特征。
作为一个示例,如图12所示,为目标对象的脸部特征点的检测结果,为了描述方便,采用数字标记识别得到的各个脸部特征点,例如图12中所示的1~17表示目标对象的脸部边缘特征点,18~22表示目标对象的左眉部特征点,23~27表示目标对象的右眉部特征点,28~36表示目标对象的鼻子特征点,37~42表示目标对象的左眼特征点,43~48表示目标对象的右眼 特征点,49~68表示目标对象的嘴唇特征点。需要指出的是,以上仅为示例,在可选实施例中可以在以上脸部特征点中仅识别部分或更多的特征点,或采用其他方式标记各个特征点,均属于本申请实施例的范畴。
S1104,在所检测到的特征中查找与具有属性的虚拟对象匹配的特征。
不同的虚拟对象所对应现实场景图像中的融合位置也不同。如图3所示,AR潜水眼镜的虚拟道具对应现实场景图像中的融合位置,应该是目标对象的眼部位置。如图4所示,AR兔子装扮的虚拟道具的耳朵对应现实场景图像中的融合位置,应该是目标对象的头部位置;而AR兔子装扮的虚拟道具的兔牙对应现实场景图像中的融合位置,应该是目标对象的牙齿部位(或嘴唇部位)。如图5所示,AR小猫装扮对应现实场景图像中的融合位置,应该是目标对象的两边脸部位置。
在一个实施例中,终端确定虚拟对象的功能或用途,根据功能或用途确定虚拟对象所要挂载在目标对象的部位,进而确定虚拟对象与目标对象的特征之间的关系。终端在所检测到的特征中,根据所确定的关系获得与具有属性的虚拟对象匹配的特征。
S1106,根据匹配的特征,确定具有属性的虚拟对象在现实场景图像中的融合位置。
上述实施例中,通过目标对象的特征点,确定虚拟对象在现实场景图像中的融合位置,以便于根据该融合位置将虚拟对象融合在现实场景图像中,获得视觉状态发生变化的虚拟对象,提升了虚拟对象的多样性变化。
在一个实施例中,该方法还可以包括:提取音频数据的音频特征;当音频特征符合第一触发条件时,则执行以下至少一种:新增虚拟对象;切换虚拟对象;切换视觉状态的类型。
其中,音频特征可以包括至少以下之一:音频数据的音量值、频率值、音色、对数功率谱和梅尔频率倒谱系数等。对数功率谱和梅尔频率倒谱系数可以反映出音频数据的功率值以及说话人的风格特征和语音表现力等特征。语音表现力可以是语音的声调、轻重和节奏等特征。对应的第一触发条件可以包括音量值达到预设音量阈值,或频率值达到预设频率阈值,或音色满足音色条件,或功率值达到功率阈值,或说话人的风格特征满足风格特征条件,或说话人的语音表现力满足表现力条件等。
视觉状态的类型可以是虚拟对象的显示尺寸、显示颜色和运动轨迹等。
在一个实施例中,终端对时域的音频数据进行分帧和加窗处理,获得各帧的音频数据。 终端对各帧的音频数据进行傅里叶变换,获得对应的频谱。终端根据各帧的频谱计算出功率谱,然后对该功率谱进行对数运算,获得对数功率谱。终端可以将该对数功率谱确定为语音特征,或者,将对数功率谱经过离散余弦变换所得的结果确定为语音特征。
例如,假设所采集的语音的信号表达式为x(n),分帧和加窗后的语音为x'(n)=x(n)×h(n),对加窗后的语音x'(n)=x(n)×h(n)进行离散傅里叶变换,得到对应的频谱信号为:
其中,N表示离散傅里叶变换的点数。
获得各帧语音的频谱时,终端计算出对应的功率谱,并求出功率谱的对数值得到对数功率谱,从而得到对应的语音特征。
或者,获得对数功率谱后,终端将对数功率谱输入梅尔尺度的三角滤波器,经离散余弦变换后得到梅尔频率倒谱系数,所得的梅尔频率倒谱系数为:
其中,L阶指的是梅尔频率倒谱系数阶数,可以取值取12-16。M指的是三角滤波器个数。
上述实施例中,通过提取音频数据的音频特征,在音频特征满足对应的触发条件时,可在原来虚拟对象的基础上新增虚拟对象,或将原来的虚拟对象切换为其它虚拟对象,或将原来所呈现的视觉状态进行切换,使融合在现实场景图像中的虚拟对象多样化,以及所呈现的视觉状态多样化,进而提高了用户与虚拟对象的交互性。
在一个实施例中,该方法还可以包括:根据音频数据进行识别,获得识别结果;确定与识别结果匹配的动态效果类型;按照动态效果类型和属性信息确定虚拟对象所呈现的视觉状态;视觉状态与动态效果类型相匹配。
其中,识别结果可以指音频类型和音频数据中关键字的文本特征。音频类型可以指音乐的类型,如轻音乐、摇滚音乐和爵士音乐等音乐类型。文本特征可以指关键字的轻音和重音等。重音可通过增加音强或音高来表示。
动态效果可以是虚拟对象在动态变化过程中所显示出来的效果。具体地,动态效果可以是以下中的任一种或多种的组合:旋转、移动、透明与非透明之间变化和颜色变化等。例如, 虚拟对象随音频数据的变化而发生旋转,或边移动边旋转等。对应的,动态效果类型可以包括旋转类型、移动类型、透明与非透明之间变化类型和颜色变化类型。
在一个实施例中,终端识别出音频数据所对应的音乐类型时,获取与音乐类型对应的动态效果,按照获取的动态效果和属性信息确定虚拟对象所呈现的视觉状态。例如,当所获取的音频数据为摇滚音乐类型时,动态效果可以是比较动感的效果。
在一个实施例中,终端识别出音频数据中关键字的文本特征,根据识别的文本特征选择对应的动态效果。例如,当识别出音频中的关键字为重音时,将虚拟对象的动态效果切换为重音所对应的动态效果。
上述实施例中,通过音频数据的识别结果确定对应的动态效果,使虚拟对象随音频数据的变化而呈现不同的动态效果,提升了用户与虚拟对象之间的交互性。
对于传统的图像处理方案中,大部分相机/短视频类的应用程序都有动态展示虚拟对象的能力,即虚拟对象会跟随人脸的运动而运动;也都有播放背景音乐或者接收麦克风的能力,即录制视频的时候可以带有背景音乐或麦克风声音。但是目前还没有一款应用程序可以实时获取声音,进而分析声音的属性来实时调节虚拟对象的属性信息。
为了解决上述问题,本申请实施例提供了一种图像处理方法,通过该图像处理方法,可实现根据音乐变化动态的调整虚拟对象(如虚拟对象随音乐变化),使虚拟对象跟随音乐的音量值或频率值变化而发生颜色,或尺寸,或旋转角度的变化,如图13所示,该图像处理方法包括以下步骤:
S1302,获取音频数据。
终端获取音频数据的方式可以是:一是通过终端的麦克风采集音频数据的方式,另一是从终端播放的相应背景音乐读取音频数据的方式。通过终端的麦克风采集音频数据的方式是采集外界的音频数据,如利用手机常用的话筒功能采集用户发出的语音。从终端播放的相应背景音乐读取音频数据的方式,是终端解码所播放背景音乐的音频格式文件,从而获得音频数据。需要说明的是,可以将上述两种方式所获得的音频数据中的一种作为输入源,也可以将上述两种方式所获得的音频数据之间的混合作为输入源。终端通过PCM编码调制方式,将所获取的音频数据编码成二进制的编码音频数据。音频数据也可称为音频信号,在本申请实 施例中不做区分。
其中,PCM是一种常见的编码方式,是将模拟的音频数据按照预设的时间间隔进行抽样,使模拟的音频数据离散化,然后对抽样值进行量化,将量化后的抽样值进行编码,获得按二进制码表示抽样脉冲的幅值。
S1304,解析音频数据,获得对应的参数值,参数值如频率值和音量值。
终端获得经过PCM编码后的编码音频数据,从该编码音频数据中解析出与声音相关的属性信息,该属性信息可以包括:音量值、频率值和音色等。
音量值可以通过音频数据的响度或对应波形的幅值表示,表征一段时间内音频数据的音量大小,计算公式如下:
其中,v
i表示经过PCM编码之后所得的编码音频数据中一个抽样点的振幅,N表示抽样点的个数,本实施例中N可以取值1024,也可以为其它数值。对于抽样率为48k的音频数据,1秒钟可计算48次音量值。
频率值可以是音频数据在单位时间内对应波形上下震动的次数,单位为Hz。音频数据可以被分解为不同频率值、不同幅值的正弦波的叠加,利用FFT(Fast Fourier Transformation,快速傅里叶变换)算法可以将PCM编码后所得的编码音频数据转化为频域音频数据,通过该频域音频数据可获得频率值和音量值(即幅值)。
将PCM编码后所得的编码音频数据转化为频域音频数据时,可使用512点FFT,最多可将0到截止频率(对于48kHz抽样率,截止频率为24kHz)的频段等分为256个频段,并计算出每个频段的幅值,从而得到音频数据的音量值。此外,获取振幅最大的频段,根据振幅最大频段中的子频域音频数据确定频数据对应的频率值。
无论是音量值的计算还是FFT的计算,在终端都可以实现实时计算。
S1306,选取对应的映射式,将获得的参数值输入映射式。
终端获得每个时段内的频域音频数据。以48kHz抽样率的音频数据作为输入源为例,1秒钟内会计算48次。终端根据不同的需求,设计不同的映射式。其中,该映射式即为本申请实施例中所述的映射关系。该映射式的输入变量为音量值或频率值,输出为虚拟对象的属性信息,属性信息如颜色、缩放比例、旋转角度等。以音频数据的音量值为输入,以虚拟对象 的缩放比例为输出为例,可以设计下述分段的映射式:
其中,x为音频数据的音量值,范围可以是0~120db,f(x)为虚拟对象的缩放比例。
根据实际需求,可以配置各种不同的映射式。其中,映射式在3个维度内可配:1)根据映射式的输入类型配置映射式,如输入类型为音量值或频率值;2)根据映射式的输出类型配置映射式,输出类型可以是虚拟对象的缩放比例、颜色、旋转角度、偏移等各个不同维度的属性信息;3)根据函数的输入与输出之间的对应关系配置映射式。
S1308,输出虚拟对象的属性信息。
根据上述映射式,当分贝值小于50db时,虚拟对象保持默认大小1.0。当分贝值大于50db时,虚拟对象的缩放比例随分贝值的增大而增大,当分贝值为120db时,缩放比例为4.0。
如图5所示,图5(a)为默认大小,即缩放比例为1.0;图5(b)为缩放大小约为2.0时的效果。
S1310,采集现实场景图像。
终端可通过摄像头实时采集现实场景图像。其中,现实场景图像可以是摄像头实时所采集视频中的一帧图像。
S1312,检测现实场景图像中对象的特征。
终端对现实场景图像中的目标对象做特征检测,如人脸特征检测。
其中,检测的方式可以是:采用开源的opencv或dlib的人脸配准点SDK,或使用优图、商汤等提供的人脸特征点检测SDK进行特征检测。
S1314,将改变属性信息的虚拟对象与现实场景图像融合。
终端调整属性信息后的虚拟对象融合到现实场景图像中目标对象的固定区域(以目标对象的某个脸部特征点位锚点),即可实现虚拟对象实时跟随人脸,并随音频数据的属性信息变化而发生变化。
以音频数据的音量值控制虚拟对象的缩放比例为例,效果是虚拟对象可以实时跟随人脸,并且虚拟对象的尺寸会随着麦克风采集到的音量值,或者背景音乐的音量值实时发生变化。
S1316,输出包括有虚拟对象的现实场景图像。
通过上述实施例,可以很大程度上增加自拍/短视频类APP的可玩性,虚拟对象会随着音频数据的参数值变化而发生大小、或颜色、或旋转角度等变化,增加了虚拟对象的多样化变化,提高了用户与虚拟对象之间的交互性。
图2为一个实施例中图像处理方法的流程示意图。应该理解的是,虽然图2的流程图中的各个步骤按照箭头的指示依次显示,但是这些步骤并不是必然按照箭头指示的顺序依次执行。除非本文中有明确的说明,这些步骤的执行并没有严格的顺序限制,这些步骤可以以其它的顺序执行。而且,图2中的至少一部分步骤可以包括多个子步骤或者多个阶段,这些子步骤或者阶段并不必然是在同一时刻执行完成,而是可以在不同的时刻执行,这些子步骤或者阶段的执行顺序也不必然是依次进行,而是可以与其它步骤或者其它步骤的子步骤或者阶段的至少一部分轮流或者交替地执行。
如图14所示,在一个实施例中,提供了一种图像处理装置,该图像处理装置具体包括:音频数据获取模块1402、属性信息确定模块1404、目标对象确定模块1406、融合位置确定模块1408和融合模块1410;其中:
音频数据获取模块1402,用于获取与实时采集的现实场景图像对应的音频数据;
属性信息确定模块1404,用于根据音频数据动态确定虚拟对象的属性信息,该属性信息用于确定虚拟对象的视觉状态;
目标对象确定模块1406,用于从现实场景图像中确定目标对象;
融合位置确定模块1408,用于根据目标对象,确定按属性信息确定的虚拟对象在现实场景图像中的融合位置;
融合模块1410,用于将按属性信息确定的虚拟对象按融合位置融合到现实场景图像;虚拟对象在属性信息不同时呈现不同的视觉状态。
在一个实施例中,音频数据获取模块1402还用于在实时采集现实场景图像时,从当前环境中实时采集与现实场景图像对应的音频数据;或者,从实时采集的现实场景图像的背景音频中,读取与现实场景图像所对应时间戳对应的音频数据。
在一个实施例中,目标对象确定模块1406还用于从现实场景图像中识别生物特征;当生物特征满足预设条件时,则将现实场景图像中与生物特征对应的生物对象确定为目标对象。
上述实施例中,获取与实时采集的现实场景图像对应的音频数据,通过音频数据动态确定虚拟对象的属性,从而实现了根据音频数据对虚拟对象属性的控制。通过现实场景图像中的目标对象,确定虚拟对象在现实场景图像中的融合位置,按照融合位置将根据属性信息确定的虚拟对象融合到现实场景图像中,由于虚拟对象的属性信息由音频数据控制,当音频数据发生变化时,融合到现实场景图像中的虚拟对象的属性信息也随之产生变化,提高了交互性。
在一个实施例中,属性信息确定模块1404还用于获取音频数据的参数值;确定参数值与虚拟对象的属性信息之间的预设映射关系;根据预设映射关系,将参数值映射为虚拟对象的属性信息。
在一个实施例中,属性信息确定模块1404对音频数据进行抽样;将抽样所得的结果量化并编码,获得编码音频数据;根据所获得的编码音频数据确定音频数据的参数值。
上述实施例中,确定参数值与虚拟对象的属性信息之间的预设映射关系,当获得对应的参数值的大小时,便可以通过该预设映射关系获得对应的虚拟对象的属性信息,从而实现对虚拟对象的调整,使虚拟对象呈现不同的视觉状态,提升了虚拟对象的多样化变化。
在一个实施例中,参数值包括频率值;属性信息确定模块1404还用于将时域的编码音频数据转换为频域音频数据;将频域音频数据分段,获得多个子频域音频数据;确定各子频域音频数据的振幅;从各子频域音频数据中选取振幅最大的子频域音频数据;按照所选取的子频域音频数据确定音频数据对应的频率值。
上述实施例中,对频域音频数据进行分段,可以通过分段后所得子频域音频数据确定频率值,一方面,通过该频率值可以实现对虚拟对象的调整,另一方面,分段后的频域音频数据,在计算过程中可降低计算量,加快计算速率。
在一个实施例中,参数值包括音量值;属性信息确定模块1404还用于根据所获得的编码音频数据确定音量值;或者,将时域的编码音频数据转换为频域音频数据;根据频域音频数据确定音量值。
上述实施例中,根据所获得的编码音频数据,或者根据由时域的编码音频数据转换的频 域音频数据这两种方式确定音量值,得到用于确定虚拟对象视觉状态的属性信息,通过该音量值可以实现对虚拟对象的调整。
在一个实施例中,属性信息包括属性调整量;按属性信息确定的虚拟对象,是按属性调整量调整相应属性后的虚拟对象;融合模块1410还用于确定虚拟对象所具有的、且与属性调整量对应的属性;按照属性调整量调整虚拟对象的属性,得到调整属性后的虚拟对象;将调整属性后的虚拟对象按融合位置融合到现实场景图像。
上述实施例中,通过属性调整量调整虚拟对象的属性,将调整属性后的虚拟对象按融合位置融合到现实场景图像,获得随音频数据的参数变化的虚拟对象,实现根据音频数据调整虚拟对象,提升了虚拟对象的多样化变化,增强用户的体验。
在一个实施例中,属性信息包括属性变化目标值;按属性信息确定的虚拟对象,是将相应属性变化至属性变化目标值后的虚拟对象;融合模块1410还用于确定虚拟对象所具有的、且与属性变化目标值对应的属性;将虚拟对象的属性变化至属性变化目标值,得到属性变化后的虚拟对象;将属性变化后的虚拟对象按融合位置融合到现实场景图像。
上述实施例中,通过属性变化目标值调整虚拟对象的属性,将调整属性后的虚拟对象按融合位置融合到现实场景图像,获得随音频数据的参数变化的虚拟对象,实现根据音频数据调整虚拟对象,提升了虚拟对象的多样化变化,增强用户的体验。
在一个实施例中,融合位置确定模块1408还用于检测目标对象的特征;在所检测到的特征中查找与具有属性的虚拟对象匹配的特征;根据匹配的特征,确定具有属性的虚拟对象在现实场景图像中的融合位置。
上述实施例中,通过目标对象的特征点,确定虚拟对象在现实场景图像中的融合位置,以便于根据该融合位置将虚拟对象融合在现实场景图像中,获得视觉状态发生变化的虚拟对象,提升了虚拟对象的多样性变化。
图15示出了一个实施例中计算机设备的内部结构图。该计算机设备具体可以是图1中的终端。如图15所示,该计算机设备包括该计算机设备包括通过系统总线连接的处理器、存储器、网络接口、输入装置和显示屏。其中,存储器包括非易失性存储介质和内存储器。该计算机设备的非易失性存储介质存储有操作系统,还可存储有计算机程序,该计算机程序被处理器执行时,可使得处理器实现图像处理方法。该内存储器中也可储存有计算机程序,该 计算机程序被处理器执行时,可使得处理器执行图像处理方法。计算机设备的显示屏可以是液晶显示屏或者电子墨水显示屏,计算机设备的输入装置可以是显示屏上覆盖的触摸层,也可以是计算机设备外壳上设置的按键、轨迹球或触控板,还可以是外接的键盘、触控板或鼠标等。
本领域技术人员可以理解,图15中示出的结构,仅仅是与本申请方案相关的部分结构的框图,并不构成对本申请方案所应用于其上的计算机设备的限定,具体的计算机设备可以包括比图中所示更多或更少的部件,或者组合某些部件,或者具有不同的部件布置。
在一个实施例中,本申请提供的图像处理装置可以实现为一种计算机程序的形式,计算机程序可在如图15所示的计算机设备上运行。计算机设备的存储器中可存储组成该14装置的各个程序模块,比如,图14所示的音频数据获取模块1402、属性信息确定模块1404、目标对象确定模块1406、融合位置确定模块1408和融合模块1410。各个程序模块构成的计算机程序使得处理器执行本说明书中描述的本申请各个实施例的图像处理方法中的步骤。
例如,图15所示的计算机设备可以通过如图14所示的图像处理装置中的音频数据获取模块1402执行S202。计算机设备可通过属性信息确定模块1404执行S204。计算机设备可通过目标对象确定模块1406执行S206。计算机设备可通过融合位置确定模块1408执行S208。计算机设备可通过融合模块1410执行S210。
在一个实施例中,提供了一种计算机设备,包括存储器和处理器,存储器存储有计算机程序,计算机程序被处理器执行时,使得处理器执行以下步骤:取与实时采集的现实场景图像对应的音频数据;根据音频数据动态确定虚拟对象的属性信息,该属性信息用于确定虚拟对象的视觉状态;从现实场景图像中确定目标对象;根据目标对象,确定按属性信息确定的虚拟对象在现实场景图像中的融合位置;将按属性信息确定的虚拟对象按融合位置融合到现实场景图像;虚拟对象在属性信息不同时呈现不同的视觉状态。
在一个实施例中,计算机程序被处理器执行获取与实时采集的现实场景图像对应的音频数据的步骤时,使得处理器具体执行以下步骤:在实时采集现实场景图像时,从当前环境中实时采集与现实场景图像对应的音频数据;或者,从实时采集的现实场景图像的背景音频中,读取与现实场景图像所对应时间戳对应的音频数据。
在一个实施例中,计算机程序被处理器执行根据音频数据动态确定虚拟对象的属性信息 的步骤时,使得处理器具体执行以下步骤:获取音频数据的参数值;确定参数值与虚拟对象的属性信息之间的预设映射关系;根据预设映射关系,将参数值映射为虚拟对象的属性信息。
在一个实施例中,计算机程序被处理器执行获取音频数据的参数值的步骤时,使得处理器具体执行以下步骤:对音频数据进行抽样;将抽样所得的结果量化并编码,获得编码音频数据;根据所获得的编码音频数据确定音频数据的参数值。
在一个实施例中,参数值包括频率值,计算机程序被处理器执行根据所获得的编码音频数据确定音频数据的参数值的步骤时,使得处理器具体执行以下步骤:将时域的编码音频数据转换为频域音频数据;将频域音频数据分段,获得多个子频域音频数据;确定各子频域音频数据的振幅;从各子频域音频数据中选取振幅最大的子频域音频数据;按照所选取的子频域音频数据确定音频数据对应的频率值。
在一个实施例中,参数值包括音量值;计算机程序被处理器执行根据所获得的编码音频数据确定音频数据的参数值的步骤时,使得处理器具体执行以下步骤:根据所获得的编码音频数据确定音量值;或者,将时域的编码音频数据转换为频域音频数据;根据频域音频数据确定音量值。
在一个实施例中,属性信息包括属性调整量;按属性信息确定的虚拟对象,是按属性调整量调整相应属性后的虚拟对象;计算机程序被处理器执行将按属性信息确定的虚拟对象按融合位置融合到现实场景图像的步骤时,使得处理器具体执行以下步骤:确定虚拟对象所具有的、且与属性调整量对应的属性;按照属性调整量调整虚拟对象的属性,得到调整属性后的虚拟对象;将调整属性后的虚拟对象按融合位置融合到现实场景图像。
在一个实施例中,属性信息包括属性变化目标值;按属性信息确定的虚拟对象,是将相应属性变化至属性变化目标值后的虚拟对象;计算机程序被处理器执行将按属性信息确定的虚拟对象按融合位置融合到现实场景图像的步骤时,使得处理器具体执行以下步骤:确定虚拟对象所具有的、且与属性变化目标值对应的属性;将虚拟对象的属性变化至属性变化目标值,得到属性变化后的虚拟对象;将属性变化后的虚拟对象按融合位置融合到现实场景图像。
在一个实施例中,计算机程序被处理器执行从现实场景图像中确定目标对象的步骤时,使得处理器具体执行以下步骤:从现实场景图像中识别生物特征;当生物特征满足预设条件时,则将现实场景图像中与生物特征对应的生物对象确定为目标对象。
在一个实施例中,计算机程序被处理器执行根据目标对象,确定具有属性的虚拟对象在现实场景图像中的融合位置的步骤时,使得处理器具体执行以下步骤:检测目标对象的特征;在所检测到的特征中查找与具有属性的虚拟对象匹配的特征;根据匹配的特征,确定具有属性的虚拟对象在现实场景图像中的融合位置。
在一个实施例中,提供了一种计算机可读存储介质,存储有计算机程序,计算机程序被处理器执行时,使得处理器执行以下步骤:取与实时采集的现实场景图像对应的音频数据;根据音频数据动态确定虚拟对象的属性信息,该属性信息用于确定虚拟对象的视觉状态;从现实场景图像中确定目标对象;根据目标对象,确定按属性信息确定的虚拟对象在现实场景图像中的融合位置;将按属性信息确定的虚拟对象按融合位置融合到现实场景图像;虚拟对象在属性信息不同时呈现不同的视觉状态。
在一个实施例中,计算机程序被处理器执行获取与实时采集的现实场景图像对应的音频数据的步骤时,使得处理器具体执行以下步骤:在实时采集现实场景图像时,从当前环境中实时采集与现实场景图像对应的音频数据;或者,从实时采集的现实场景图像的背景音频中,读取与现实场景图像所对应时间戳对应的音频数据。
在一个实施例中,计算机程序被处理器执行根据音频数据动态确定虚拟对象的属性信息的步骤时,使得处理器具体执行以下步骤:获取音频数据的参数值;确定参数值与虚拟对象的属性信息之间的预设映射关系;根据预设映射关系,将参数值映射为虚拟对象的属性信息。
在一个实施例中,计算机程序被处理器执行获取音频数据的参数值的步骤时,使得处理器具体执行以下步骤:对音频数据进行抽样;将抽样所得的结果量化并编码,获得编码音频数据;根据所获得的编码音频数据确定音频数据的参数值。
在一个实施例中,参数值包括频率值,计算机程序被处理器执行根据所获得的编码音频数据确定音频数据的参数值的步骤时,使得处理器具体执行以下步骤:将时域的编码音频数据转换为频域音频数据;将频域音频数据分段,获得多个子频域音频数据;确定各子频域音频数据的振幅;从各子频域音频数据中选取振幅最大的子频域音频数据;按照所选取的子频域音频数据确定音频数据对应的频率值。
在一个实施例中,参数值包括音量值;计算机程序被处理器执行根据所获得的编码音频数据确定音频数据的参数值的步骤时,使得处理器具体执行以下步骤:根据所获得的编码音 频数据确定音量值;或者,将时域的编码音频数据转换为频域音频数据;根据频域音频数据确定音量值。
在一个实施例中,属性信息包括属性调整量;按属性信息确定的虚拟对象,是按属性调整量调整相应属性后的虚拟对象;计算机程序被处理器执行将按属性信息确定的虚拟对象按融合位置融合到现实场景图像的步骤时,使得处理器具体执行以下步骤:确定虚拟对象所具有的、且与属性调整量对应的属性;按照属性调整量调整虚拟对象的属性,得到调整属性后的虚拟对象;将调整属性后的虚拟对象按融合位置融合到现实场景图像。
在一个实施例中,属性信息包括属性变化目标值;按属性信息确定的虚拟对象,是将相应属性变化至属性变化目标值后的虚拟对象;计算机程序被处理器执行将按属性信息确定的虚拟对象按融合位置融合到现实场景图像的步骤时,使得处理器具体执行以下步骤:确定虚拟对象所具有的、且与属性变化目标值对应的属性;将虚拟对象的属性变化至属性变化目标值,得到属性变化后的虚拟对象;将属性变化后的虚拟对象按融合位置融合到现实场景图像。
在一个实施例中,计算机程序被处理器执行从现实场景图像中确定目标对象的步骤时,使得处理器具体执行以下步骤:从现实场景图像中识别生物特征;当生物特征满足预设条件时,则将现实场景图像中与生物特征对应的生物对象确定为目标对象。
在一个实施例中,计算机程序被处理器执行根据目标对象,确定具有属性的虚拟对象在现实场景图像中的融合位置的步骤时,使得处理器具体执行以下步骤:检测目标对象的特征;在所检测到的特征中查找与具有属性的虚拟对象匹配的特征;根据匹配的特征,确定具有属性的虚拟对象在现实场景图像中的融合位置。
在一个实施例中,计算机程序被处理器执行时,使得处理器具体执行以下步骤:提取音频数据的音频特征;当音频特征符合第一触发条件时,则执行以下至少一种:新增虚拟对象;切换虚拟对象;切换视觉状态的类型。
在一个实施例中,计算机程序被处理器执行时,使得处理器具体执行以下步骤:根据音频数据进行识别,获得识别结果;确定与识别结果匹配的动态效果类型;按照动态效果类型和属性信息确定虚拟对象所呈现的视觉状态;视觉状态与动态效果类型相匹配。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机程序来指令相关的硬件来完成,所述的程序可存储于一非易失性计算机可读取存储介质 中,该程序在执行时,可包括如上述各方法的实施例的流程。其中,本申请所提供的各实施例中所使用的对存储器、存储、数据库或其它介质的任何引用,均可包括非易失性和/或易失性存储器。非易失性存储器可包括只读存储器(ROM)、可编程ROM(PROM)、电可编程ROM(EPROM)、电可擦除可编程ROM(EEPROM)或闪存。易失性存储器可包括随机存取存储器(RAM)或者外部高速缓冲存储器。作为说明而非局限,RAM以多种形式可得,诸如静态RAM(SRAM)、动态RAM(DRAM)、同步DRAM(SDRAM)、双数据率SDRAM(DDRSDRAM)、增强型SDRAM(ESDRAM)、同步链路(Synchlink)DRAM(SLDRAM)、存储器总线(Rambus)直接RAM(RDRAM)、直接存储器总线动态RAM(DRDRAM)、以及存储器总线动态RAM(RDRAM)等。
以上实施例的各技术特征可以进行任意的组合,为使描述简洁,未对上述实施例中的各个技术特征所有可能的组合都进行描述,然而,只要这些技术特征的组合不存在矛盾,都应当认为是本说明书记载的范围。
以上所述实施例仅表达了本申请的几种实施方式,其描述较为具体和详细,但并不能因此而理解为对本申请专利范围的限制。应当指出的是,对于本领域的普通技术人员来说,在不脱离本申请构思的前提下,还可以做出若干变形和改进,这些都属于本申请的保护范围。因此,本申请专利的保护范围应以所附权利要求为准。
Claims (28)
- 一种图像处理方法,其特征在于,应用于图像处理系统,所述方法包括:获取与实时采集的现实场景图像对应的音频数据;根据所述音频数据动态确定虚拟对象的属性信息,所述属性信息用于确定所述虚拟对象的视觉状态;从所述现实场景图像中确定目标对象;根据所述目标对象,确定按所述属性信息确定的虚拟对象在所述现实场景图像中的融合位置;将按所述属性信息确定的虚拟对象按所述融合位置融合到所述现实场景图像;所述虚拟对象在属性信息不同时呈现不同的视觉状态。
- 根据权利要求1所述的方法,其特征在于,所述获取与实时采集的现实场景图像对应的音频数据包括:在实时采集现实场景图像时,从当前环境中实时采集与所述现实场景图像对应的音频数据;或者,从实时采集的现实场景图像的背景音频中,读取与所述现实场景图像所对应时间戳对应的音频数据。
- 根据权利要求1所述的方法,其特征在于,所述根据所述音频数据动态确定虚拟对象的属性信息包括:获取所述音频数据的参数值;确定参数值与虚拟对象的属性信息之间的预设映射关系;根据所述预设映射关系,将所述参数值映射为所述虚拟对象的属性信息。
- 根据权利要求3所述的方法,其特征在于,所述获取所述音频数据的参数值包括:对所述音频数据进行抽样;将抽样所得的结果量化并编码,获得编码音频数据;根据所获得的编码音频数据确定所述音频数据的参数值。
- 根据权利要求3所述的方法,其特征在于,所述参数值包括频率值;所述根据所获 得的编码音频数据确定所述音频数据的参数值包括:将时域的所述编码音频数据转换为频域音频数据;将所述频域音频数据分段,获得多个子频域音频数据;确定各子频域音频数据的振幅;从各子频域音频数据中选取振幅最大的子频域音频数据;按照所选取的子频域音频数据确定所述音频数据对应的频率值。
- 根据权利要求3所述的方法,其特征在于,所述参数值包括音量值;所述根据所获得的编码音频数据确定所述音频数据的参数值包括:根据所获得的编码音频数据确定音量值;或者,将时域的所述编码音频数据转换为频域音频数据;根据所述频域音频数据确定音量值。
- 根据权利要求1所述的方法,其特征在于,所述属性信息包括属性调整量;所述按所述属性信息确定的虚拟对象,是按所述属性调整量调整相应属性后的虚拟对象;所述将按所述属性信息确定的虚拟对象按所述融合位置融合到所述现实场景图像包括:确定所述虚拟对象所具有的、且与所述属性调整量对应的属性;按照所述属性调整量调整所述虚拟对象的所述属性,得到调整属性后的虚拟对象;将调整属性后的虚拟对象按所述融合位置融合到所述现实场景图像。
- 根据权利要求1所述的方法,其特征在于,所述属性信息包括属性变化目标值;所述按所述属性信息确定的虚拟对象,是将相应属性变化至所述属性变化目标值后的虚拟对象;所述将按所述属性信息确定的虚拟对象按所述融合位置融合到所述现实场景图像包括:确定所述虚拟对象所具有的、且与所述属性变化目标值对应的属性;将所述虚拟对象的所述属性变化至所述属性变化目标值,得到属性变化后的虚拟对象;将属性变化后的虚拟对象按所述融合位置融合到所述现实场景图像。
- 根据权利要求1至8任一项所述的方法,其特征在于,所述从所述现实场景图像中确定目标对象包括:从所述现实场景图像中识别生物特征;当所述生物特征满足预设条件时,则将所述现实场景图像中与所述生物特征对应的生物对象确定为目标对象。
- 根据权利要求1至8任一项所述的方法,其特征在于,所述根据所述目标对象,确定具有所述属性的虚拟对象在所述现实场景图像中的融合位置包括:检测所述目标对象的特征;在所检测到的特征中查找与具有所述属性的虚拟对象匹配的特征;根据所述匹配的特征,确定具有所述属性的虚拟对象在所述现实场景图像中的融合位置。
- 根据权利要求1至8任一项所述的方法,其特征在于,还包括:提取所述音频数据的音频特征;当所述音频特征符合第一触发条件时,则执行以下至少一种:新增虚拟对象;切换虚拟对象;切换所述视觉状态的类型。
- 根据权利要求1至8任一项所述的方法,其特征在于,还包括:根据所述音频数据进行识别,获得识别结果;确定与所述识别结果匹配的动态效果类型;按照所述动态效果类型和所述属性信息确定所述虚拟对象所呈现的视觉状态;所述视觉状态与所述动态效果类型相匹配。
- 一种图像处理装置,其特征在于,包括:音频数据获取模块,用于获取与实时采集的现实场景图像对应的音频数据;属性信息确定模块,用于根据所述音频数据动态确定虚拟对象的属性信息,所述属性信息用于确定所述虚拟对象的视觉状态;目标对象确定模块,用于从所述现实场景图像中确定目标对象;融合位置确定模块,用于根据所述目标对象,确定按所述属性信息确定的虚拟对象在所述现实场景图像中的融合位置;融合模块,用于将按所述属性信息确定的虚拟对象按所述融合位置融合到所述现实场景图像;所述虚拟对象在属性信息不同时呈现不同的视觉状态。
- 根据权利要求13所述的装置,其特征在于,所述音频数据获取模块还用于在实时 采集现实场景图像时,从当前环境中实时采集与所述现实场景图像对应的音频数据;或者,从实时采集的现实场景图像的背景音频中,读取与所述现实场景图像所对应时间戳对应的音频数据。
- 根据权利要求14所述的装置,其特征在于,属性信息确定模块还用于获取所述音频数据的参数值;确定参数值与虚拟对象的属性信息之间的预设映射关系;根据所述预设映射关系,将所述参数值映射为虚拟对象的属性信息。
- 一种存储介质,其特征在于,存储有计算机程序,所述计算机程序被处理器执行时,使得所述处理器执行如权利要求1至12中任一项所述方法的步骤。
- 一种计算机设备,其特征在于,包括存储器和处理器,所述存储器存储有计算机程序,所述计算机程序被所述处理器执行时,使得所述处理器执行以下步骤:获取与实时采集的现实场景图像对应的音频数据;根据所述音频数据动态确定虚拟对象的属性信息,所述属性信息用于确定所述虚拟对象的视觉状态;从所述现实场景图像中确定目标对象;根据所述目标对象,确定按所述属性信息确定的虚拟对象在所述现实场景图像中的融合位置;将按所述属性信息确定的虚拟对象按所述融合位置融合到所述现实场景图像;所述虚拟对象在属性信息不同时呈现不同的视觉状态。
- 根据权利要求17所述的计算机设备,其特征在于,所述计算机程序被所述处理器执行获取与实时采集的现实场景图像对应的音频数据的步骤时,使得所述处理器执行以下步骤:在实时采集现实场景图像时,从当前环境中实时采集与所述现实场景图像对应的音频数据;或者,从实时采集的现实场景图像的背景音频中,读取与所述现实场景图像所对应时间戳对应的音频数据。
- 根据权利要求17所述的计算机设备,其特征在于,所述计算机程序被所述处理器执行根据音频数据动态确定虚拟对象的属性信息的步骤时,使得所述处理器执行以下步骤:获取所述音频数据的参数值;确定参数值与虚拟对象的属性信息之间的预设映射关系;根据所述预设映射关系,将所述参数值映射为所述虚拟对象的属性信息。
- 根据权利要求19所述的计算机设备,其特征在于,所述计算机程序被所述处理器执行获取所述音频数据的参数值的步骤时,使得所述处理器执行以下步骤:对所述音频数据进行抽样;将抽样所得的结果量化并编码,获得编码音频数据;根据所获得的编码音频数据确定所述音频数据的参数值。
- 根据权利要求19所述的计算机设备,其特征在于,所述参数值包括频率值;所述计算机程序被所述处理器执行根据所获得的编码音频数据确定所述音频数据的参数值的步骤时,使得所述处理器执行以下步骤:将时域的所述编码音频数据转换为频域音频数据;将所述频域音频数据分段,获得多个子频域音频数据;确定各子频域音频数据的振幅;从各子频域音频数据中选取振幅最大的子频域音频数据;按照所选取的子频域音频数据确定所述音频数据对应的频率值。
- 根据权利要求19所述的计算机设备,其特征在于,所述参数值包括音量值;所述计算机程序被所述处理器执行根据所获得的编码音频数据确定所述音频数据的参数值的步骤时,使得所述处理器执行以下步骤:根据所获得的编码音频数据确定音量值;或者,将时域的所述编码音频数据转换为频域音频数据;根据所述频域音频数据确定音量值。
- 根据权利要求17所述的计算机设备,其特征在于,所述属性信息包括属性调整量;所述按所述属性信息确定的虚拟对象,是按所述属性调整量调整相应属性后的虚拟对象;所述计算机程序被所述处理器执行将按所述属性信息确定的虚拟对象按所述融合位置融合到所述现实场景图像的步骤时,使得所述处理器执行以下步骤:确定所述虚拟对象所具有的、且与所述属性调整量对应的属性;按照所述属性调整量调整所述虚拟对象的所述属性,得到调整属性后的虚拟对象;将调整属性后的虚拟对象按所述融合位置融合到所述现实场景图像。
- 根据权利要求17所述的计算机设备,其特征在于,所述属性信息包括属性变化目标值;所述按所述属性信息确定的虚拟对象,是将相应属性变化至所述属性变化目标值后的虚拟对象;所述计算机程序被所述处理器执行将按所述属性信息确定的虚拟对象按所述融合位置融合到所述现实场景图像的步骤时,使得所述处理器执行以下步骤:确定所述虚拟对象所具有的、且与所述属性变化目标值对应的属性;将所述虚拟对象的所述属性变化至所述属性变化目标值,得到属性变化后的虚拟对象;将属性变化后的虚拟对象按所述融合位置融合到所述现实场景图像。
- 根据权利要求17至24任一项所述的计算机设备,其特征在于,所述计算机程序被所述处理器执行从所述现实场景图像中确定目标对象的步骤时,使得所述处理器执行以下步骤:从所述现实场景图像中识别生物特征;当所述生物特征满足预设条件时,则将所述现实场景图像中与所述生物特征对应的生物对象确定为目标对象。
- 根据权利要求17至24任一项所述的计算机设备,其特征在于,所述计算机程序被所述处理器执行根据所述目标对象,确定具有所述属性的虚拟对象在所述现实场景图像中的融合位置的步骤时,使得所述处理器执行以下步骤:检测所述目标对象的特征;在所检测到的特征中查找与具有所述属性的虚拟对象匹配的特征;根据所述匹配的特征,确定具有所述属性的虚拟对象在所述现实场景图像中的融合位置。
- 根据权利要求17至24任一项所述的计算机设备,其特征在于,所述计算机程序被所述处理器执行时,使得所述处理器执行以下步骤:提取所述音频数据的音频特征;当所述音频特征符合第一触发条件时,则执行以下至少一种:新增虚拟对象;切换虚拟对象;切换所述视觉状态的类型。
- 根据权利要求17至24任一项所述的计算机设备,其特征在于,所述计算机程序被所述处理器执行时,使得所述处理器执行以下步骤:根据所述音频数据进行识别,获得识别结果;确定与所述识别结果匹配的动态效果类型;按照所述动态效果类型和所述属性信息确定所述虚拟对象所呈现的视觉状态;所述视觉状态与所述动态效果类型相匹配。
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US16/999,047 US11392642B2 (en) | 2018-07-04 | 2020-08-20 | Image processing method, storage medium, and computer device |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201810723144.8 | 2018-07-04 | ||
| CN201810723144.8A CN108769535B (zh) | 2018-07-04 | 2018-07-04 | 图像处理方法、装置、存储介质和计算机设备 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US16/999,047 Continuation US11392642B2 (en) | 2018-07-04 | 2020-08-20 | Image processing method, storage medium, and computer device |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020007185A1 true WO2020007185A1 (zh) | 2020-01-09 |
Family
ID=63975934
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2019/091359 Ceased WO2020007185A1 (zh) | 2018-07-04 | 2019-06-14 | 图像处理方法、装置、存储介质和计算机设备 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US11392642B2 (zh) |
| CN (1) | CN108769535B (zh) |
| TW (1) | TWI793344B (zh) |
| WO (1) | WO2020007185A1 (zh) |
Cited By (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111833460A (zh) * | 2020-07-10 | 2020-10-27 | 北京字节跳动网络技术有限公司 | 增强现实的图像处理方法、装置、电子设备及存储介质 |
| US20210321157A1 (en) * | 2020-06-28 | 2021-10-14 | Baidu Online Network Technology (Beijing) Co., Ltd. | Special effect processing method and apparatus for live broadcasting, and server |
| CN113948105A (zh) * | 2021-09-30 | 2022-01-18 | 深圳追一科技有限公司 | 基于语音的图像生成方法、装置、设备及介质 |
| CN114007091A (zh) * | 2021-10-27 | 2022-02-01 | 北京市商汤科技开发有限公司 | 一种视频处理方法、装置、电子设备及存储介质 |
| CN115239621A (zh) * | 2022-06-10 | 2022-10-25 | 国网安徽省电力有限公司超高压分公司 | 一种基于虚实匹配的变电站视觉智能巡检方法 |
| CN117014675A (zh) * | 2022-09-16 | 2023-11-07 | 腾讯科技(深圳)有限公司 | 虚拟对象的视频生成方法、装置和计算机可读存储介质 |
Families Citing this family (29)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108769535B (zh) * | 2018-07-04 | 2021-08-10 | 腾讯科技(深圳)有限公司 | 图像处理方法、装置、存储介质和计算机设备 |
| CN110070896B (zh) * | 2018-10-19 | 2020-09-01 | 北京微播视界科技有限公司 | 图像处理方法、装置、硬件装置 |
| CN109597481B (zh) * | 2018-11-16 | 2021-05-04 | Oppo广东移动通信有限公司 | Ar虚拟人物绘制方法、装置、移动终端及存储介质 |
| CN110072047B (zh) | 2019-01-25 | 2020-10-09 | 北京字节跳动网络技术有限公司 | 图像形变的控制方法、装置和硬件装置 |
| CN109935318B (zh) * | 2019-03-06 | 2021-07-30 | 智美康民(珠海)健康科技有限公司 | 三维脉波的显示方法、装置、计算机设备及存储介质 |
| CN113906370B (zh) * | 2019-05-29 | 2024-08-02 | 苹果公司 | 为物理元素生成内容 |
| CN112738634B (zh) * | 2019-10-14 | 2022-08-02 | 北京字节跳动网络技术有限公司 | 视频文件的生成方法、装置、终端及存储介质 |
| CN110866509B (zh) * | 2019-11-20 | 2023-04-28 | 腾讯科技(深圳)有限公司 | 动作识别方法、装置、计算机存储介质和计算机设备 |
| CN113225613B (zh) * | 2020-01-21 | 2022-07-08 | 北京达佳互联信息技术有限公司 | 图像识别、视频直播方法和装置 |
| CN111638781B (zh) * | 2020-05-15 | 2024-03-19 | 广东小天才科技有限公司 | 基于ar的发音引导方法、装置、电子设备及存储介质 |
| CN111627117B (zh) * | 2020-06-01 | 2024-04-16 | 上海商汤智能科技有限公司 | 画像展示特效的调整方法、装置、电子设备及存储介质 |
| CN111815782A (zh) * | 2020-06-30 | 2020-10-23 | 北京市商汤科技开发有限公司 | Ar场景内容的显示方法、装置、设备及计算机存储介质 |
| CN111899350A (zh) * | 2020-07-31 | 2020-11-06 | 北京市商汤科技开发有限公司 | 增强现实ar图像的呈现方法及装置、电子设备、存储介质 |
| CN111915744B (zh) * | 2020-08-31 | 2024-09-27 | 深圳传音控股股份有限公司 | 增强现实图像的交互方法、终端和存储介质 |
| CN112367426B (zh) | 2020-11-09 | 2021-06-04 | Oppo广东移动通信有限公司 | 虚拟对象显示方法及装置、存储介质和电子设备 |
| CN112767484B (zh) * | 2021-01-25 | 2023-09-05 | 脸萌有限公司 | 定位模型的融合方法、定位方法、电子装置 |
| CN113256815B (zh) * | 2021-02-24 | 2024-03-22 | 北京华清易通科技有限公司 | 虚拟现实场景融合及播放方法和虚拟现实设备 |
| CN112990283B (zh) * | 2021-03-03 | 2024-07-26 | 网易(杭州)网络有限公司 | 图像生成方法、装置和电子设备 |
| US11521341B1 (en) | 2021-06-21 | 2022-12-06 | Lemon Inc. | Animation effect attachment based on audio characteristics |
| US11769289B2 (en) | 2021-06-21 | 2023-09-26 | Lemon Inc. | Rendering virtual articles of clothing based on audio characteristics |
| CN113905189B (zh) * | 2021-09-28 | 2024-06-14 | 安徽尚趣玩网络科技有限公司 | 一种视频内容动态拼接方法及装置 |
| CN113867530A (zh) * | 2021-09-28 | 2021-12-31 | 深圳市慧鲤科技有限公司 | 虚拟物体控制方法、装置、设备及存储介质 |
| CN113870438A (zh) * | 2021-09-28 | 2021-12-31 | 深圳市慧鲤科技有限公司 | 虚拟物体控制方法、装置、设备及存储介质 |
| US20240354764A1 (en) * | 2022-01-13 | 2024-10-24 | Garteh John Llewelyn | Digital Gateway Manager for Physical Luxury Items and Collectibles |
| CN114419213B (zh) * | 2022-01-24 | 2025-12-12 | 北京字跳网络技术有限公司 | 图像处理方法、装置、设备和存储介质 |
| US12548225B2 (en) * | 2022-06-17 | 2026-02-10 | Lemon Inc. | Audio or visual input interacting with video creation |
| CN115359220B (zh) * | 2022-08-16 | 2024-05-07 | 支付宝(杭州)信息技术有限公司 | 虚拟世界的虚拟形象更新方法及装置 |
| CN115512003B (zh) * | 2022-11-16 | 2023-04-28 | 之江实验室 | 一种独立关系检测的场景图生成方法和系统 |
| US12511849B2 (en) * | 2022-12-30 | 2025-12-30 | Htc Corporation | Method for improving visual quality of reality service content, host, and computer readable storage medium |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20140064517A1 (en) * | 2012-09-05 | 2014-03-06 | Acer Incorporated | Multimedia processing system and audio signal processing method |
| CN104036789A (zh) * | 2014-01-03 | 2014-09-10 | 北京智谷睿拓技术服务有限公司 | 多媒体处理方法及多媒体装置 |
| CN107329980A (zh) * | 2017-05-31 | 2017-11-07 | 福建星网视易信息系统有限公司 | 一种基于音频的实时联动显示方法及存储设备 |
| CN107343211A (zh) * | 2016-08-19 | 2017-11-10 | 北京市商汤科技开发有限公司 | 视频图像处理方法、装置和终端设备 |
| CN108769535A (zh) * | 2018-07-04 | 2018-11-06 | 腾讯科技(深圳)有限公司 | 图像处理方法、装置、存储介质和计算机设备 |
Family Cites Families (19)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6917370B2 (en) * | 2002-05-13 | 2005-07-12 | Charles Benton | Interacting augmented reality and virtual reality |
| US20110191674A1 (en) * | 2004-08-06 | 2011-08-04 | Sensable Technologies, Inc. | Virtual musical interface in a haptic virtual environment |
| TWI423113B (zh) * | 2010-11-30 | 2014-01-11 | Compal Communication Inc | 互動式人機介面及其所適用之可攜式通訊裝置 |
| US9563265B2 (en) * | 2012-01-12 | 2017-02-07 | Qualcomm Incorporated | Augmented reality with sound and geometric analysis |
| TW201404133A (zh) * | 2012-07-09 | 2014-01-16 | Wistron Corp | 自動拍照裝置及方法 |
| TWI579731B (zh) * | 2013-08-22 | 2017-04-21 | Chunghwa Telecom Co Ltd | Combined with the reality of the scene and virtual components of the interactive system and methods |
| CN103685946B (zh) * | 2013-11-30 | 2017-08-08 | 北京智谷睿拓技术服务有限公司 | 一种拍照方法和控制设备 |
| WO2015185110A1 (en) * | 2014-06-03 | 2015-12-10 | Metaio Gmbh | Method and system for presenting a digital information related to a real object |
| US9361447B1 (en) * | 2014-09-04 | 2016-06-07 | Emc Corporation | Authentication based on user-selected image overlay effects |
| CN104580888B (zh) * | 2014-12-17 | 2018-09-04 | 广东欧珀移动通信有限公司 | 一种图像处理方法及终端 |
| KR20160133328A (ko) * | 2015-05-12 | 2016-11-22 | 삼성전자주식회사 | 웨어러블 디바이스를 이용한 원격 제어 방법 및 장치 |
| US10462524B2 (en) * | 2015-06-23 | 2019-10-29 | Facebook, Inc. | Streaming media presentation system |
| US11024088B2 (en) * | 2016-05-27 | 2021-06-01 | HoloBuilder, Inc. | Augmented and virtual reality |
| US10445936B1 (en) * | 2016-08-01 | 2019-10-15 | Snap Inc. | Audio responsive augmented reality |
| US10277834B2 (en) * | 2017-01-10 | 2019-04-30 | International Business Machines Corporation | Suggestion of visual effects based on detected sound patterns |
| CN107105168A (zh) * | 2017-06-02 | 2017-08-29 | 哈尔滨市舍科技有限公司 | 可虚拟拍照的共享观景系统 |
| EP3633587B1 (en) * | 2018-10-04 | 2024-07-03 | Nokia Technologies Oy | An apparatus and associated methods for presentation of comments |
| US11017231B2 (en) * | 2019-07-10 | 2021-05-25 | Microsoft Technology Licensing, Llc | Semantically tagged virtual and physical objects |
| US11089427B1 (en) * | 2020-03-31 | 2021-08-10 | Snap Inc. | Immersive augmented reality experiences using spatial audio |
-
2018
- 2018-07-04 CN CN201810723144.8A patent/CN108769535B/zh active Active
-
2019
- 2019-06-14 WO PCT/CN2019/091359 patent/WO2020007185A1/zh not_active Ceased
- 2019-07-04 TW TW108123650A patent/TWI793344B/zh active
-
2020
- 2020-08-20 US US16/999,047 patent/US11392642B2/en active Active
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20140064517A1 (en) * | 2012-09-05 | 2014-03-06 | Acer Incorporated | Multimedia processing system and audio signal processing method |
| CN104036789A (zh) * | 2014-01-03 | 2014-09-10 | 北京智谷睿拓技术服务有限公司 | 多媒体处理方法及多媒体装置 |
| CN107343211A (zh) * | 2016-08-19 | 2017-11-10 | 北京市商汤科技开发有限公司 | 视频图像处理方法、装置和终端设备 |
| CN107329980A (zh) * | 2017-05-31 | 2017-11-07 | 福建星网视易信息系统有限公司 | 一种基于音频的实时联动显示方法及存储设备 |
| CN108769535A (zh) * | 2018-07-04 | 2018-11-06 | 腾讯科技(深圳)有限公司 | 图像处理方法、装置、存储介质和计算机设备 |
Cited By (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20210321157A1 (en) * | 2020-06-28 | 2021-10-14 | Baidu Online Network Technology (Beijing) Co., Ltd. | Special effect processing method and apparatus for live broadcasting, and server |
| US11722727B2 (en) * | 2020-06-28 | 2023-08-08 | Baidu Online Network Technology (Beijing) Co., Ltd. | Special effect processing method and apparatus for live broadcasting, and server |
| CN111833460A (zh) * | 2020-07-10 | 2020-10-27 | 北京字节跳动网络技术有限公司 | 增强现实的图像处理方法、装置、电子设备及存储介质 |
| JP2023533295A (ja) * | 2020-07-10 | 2023-08-02 | 北京字節跳動網絡技術有限公司 | 拡張現実の画像処理方法、装置、電子機器及び記憶媒体 |
| US11756276B2 (en) | 2020-07-10 | 2023-09-12 | Beijing Bytedance Network Technology Co., Ltd. | Image processing method and apparatus for augmented reality, electronic device, and storage medium |
| EP4167192A4 (en) * | 2020-07-10 | 2023-12-13 | Beijing Bytedance Network Technology Co., Ltd. | IMAGE PROCESSING METHOD AND APPARATUS FOR AUGMENTED REALITY, ELECTRONIC DEVICE AND STORAGE MEDIUM |
| CN111833460B (zh) * | 2020-07-10 | 2024-07-26 | 北京字节跳动网络技术有限公司 | 增强现实的图像处理方法、装置、电子设备及存储介质 |
| JP7674462B2 (ja) | 2020-07-10 | 2025-05-09 | 北京字節跳動網絡技術有限公司 | 拡張現実の画像処理方法、装置、電子機器及び記憶媒体 |
| CN113948105A (zh) * | 2021-09-30 | 2022-01-18 | 深圳追一科技有限公司 | 基于语音的图像生成方法、装置、设备及介质 |
| CN114007091A (zh) * | 2021-10-27 | 2022-02-01 | 北京市商汤科技开发有限公司 | 一种视频处理方法、装置、电子设备及存储介质 |
| CN115239621A (zh) * | 2022-06-10 | 2022-10-25 | 国网安徽省电力有限公司超高压分公司 | 一种基于虚实匹配的变电站视觉智能巡检方法 |
| CN117014675A (zh) * | 2022-09-16 | 2023-11-07 | 腾讯科技(深圳)有限公司 | 虚拟对象的视频生成方法、装置和计算机可读存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| TW202022683A (zh) | 2020-06-16 |
| TWI793344B (zh) | 2023-02-21 |
| CN108769535B (zh) | 2021-08-10 |
| US11392642B2 (en) | 2022-07-19 |
| US20200380031A1 (en) | 2020-12-03 |
| CN108769535A (zh) | 2018-11-06 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| TWI793344B (zh) | 影像處理方法、裝置、儲存介質和電腦設備 | |
| Czyzewski et al. | An audio-visual corpus for multimodal automatic speech recognition | |
| CN111489424A (zh) | 虚拟角色表情生成方法、控制方法、装置和终端设备 | |
| US12027165B2 (en) | Computer program, server, terminal, and speech signal processing method | |
| US8447065B2 (en) | Method of facial image reproduction and related device | |
| US20230290382A1 (en) | Method and apparatus for matching music with video, computer device, and storage medium | |
| CN112512649B (zh) | 用于提供音频和视频效果的技术 | |
| CN109509483A (zh) | 产生频率增强音频信号的译码器和产生编码信号的编码器 | |
| CN113113040B (zh) | 音频处理方法及装置、终端及存储介质 | |
| CN115439614B (zh) | 虚拟形象的生成方法、装置、电子设备和存储介质 | |
| CN117854535B (zh) | 基于交叉注意力的视听语音增强方法及其模型搭建方法 | |
| CN113724683A (zh) | 音频生成方法、计算机设备及计算机可读存储介质 | |
| CA3053032A1 (fr) | Methode et appareil de modification dynamique du timbre de la voix par decalage en frequence des formants d'une enveloppe spectrale | |
| CN114220414A (zh) | 语音合成方法以及相关装置、设备 | |
| WO2023160515A1 (zh) | 视频处理方法、装置、设备及介质 | |
| CN120673773A (zh) | 基于噪声感知的语音增强方法、装置、设备及介质 | |
| CN113301372A (zh) | 直播方法、装置、终端及存储介质 | |
| CN112885318A (zh) | 多媒体数据生成方法、装置、电子设备及计算机存储介质 | |
| CN117877504A (zh) | 一种联合语音增强方法及其模型搭建方法 | |
| CN110619886A (zh) | 一种针对低资源土家语的端到端语音增强方法 | |
| WO2022041192A1 (zh) | 语音消息处理方法、设备及即时通信客户端 | |
| CN112466306A (zh) | 会议纪要生成方法、装置、计算机设备及存储介质 | |
| CN119360830A (zh) | 基于大模型的语音风格识别系统 | |
| JP5427622B2 (ja) | 音声変更装置、音声変更方法、プログラム及び記録媒体 | |
| CN113362432A (zh) | 一种面部动画生成方法及装置 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19831229 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 19831229 Country of ref document: EP Kind code of ref document: A1 |


