WO2024217237A1 - 数据生成方法、装置、电子设备、计算机可读介质 - Google Patents
数据生成方法、装置、电子设备、计算机可读介质 Download PDFInfo
- Publication number
- WO2024217237A1 WO2024217237A1 PCT/CN2024/084022 CN2024084022W WO2024217237A1 WO 2024217237 A1 WO2024217237 A1 WO 2024217237A1 CN 2024084022 W CN2024084022 W CN 2024084022W WO 2024217237 A1 WO2024217237 A1 WO 2024217237A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- data
- target
- processed
- background
- adjusted
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T11/00—Two-dimensional [2D] image generation
- G06T11/60—Creating or editing images; Combining images with text
Definitions
- the present disclosure relates to a data generation method, device, electronic device, and computer-readable medium.
- VR virtual reality
- VR image data usually needs to be shot in a room where VR shooting equipment is pre-arranged.
- the background involved in the VR image data is also relatively simple or single, which makes the VR image data unable to meet the background requirements of some application scenarios, thus affecting the development potential of VR image data.
- the present disclosure provides a data generation method, device, electronic device, and computer-readable medium, which can overcome the adverse effects caused by the inability of VR image data to meet the background requirements of some application scenarios.
- the present disclosure provides a data generation method, the method comprising:
- target background representation data corresponding to the data to be processed; the target background described by the target background representation data is different from the original background;
- the target object in the data to be processed is subjected to brightness adjustment processing to obtain adjusted data; the brightness state of the target object in the adjusted data is different from the target environment map. Describing the brightness state of the target object in the data to be processed;
- performing brightness adjustment processing on the target object in the data to be processed according to the target environment map to obtain adjusted data includes:
- the adjusted data is determined according to the normal estimation result, the data to be processed, and the light map.
- determining the adjusted data according to the normal estimation result, the data to be processed, and the light map includes:
- brightness adjustment processing is performed on the data to be processed to obtain a first adjustment result
- the adjusted data is determined according to the first adjustment result.
- the method before determining the adjusted data according to the first adjustment result, the method further includes:
- the determining the adjusted data according to the first adjustment result includes:
- the adjusted data is determined according to the object position representation data and the first adjustment result, so that the adjusted brightness corresponding to the target object is recorded in the adjusted data.
- the adjusted data is virtual reality (VR) image data
- the data to be processed is image data converted from the original VR image data
- the step of determining the adjusted data according to the object position representation data and the first adjustment result includes:
- the adjusted data According to the first VR image data and the second VR image data, the adjusted data.
- the method before determining the adjusted data according to the normal estimation result, the data to be processed, and the light map, the method further includes:
- the step of determining the adjusted data according to the normal estimation result, the data to be processed, and the light map includes:
- the adjusted data is determined according to the object position representation data, the normal estimation result, the data to be processed, and the light map.
- determining the adjusted data according to the normal estimation result, the data to be processed, and the light map includes:
- the adjusted data is determined according to the object normal result, the object representation data, and the light map.
- determining the adjusted data according to the object normal result, the object representation data, and the light map includes:
- the second light brightness representation information performing brightness adjustment processing on the object representation data to obtain a second adjustment result
- the adjusted data is determined according to the second adjustment result.
- the data to be processed, the object position representation data, and the normal estimation result are all image sequences
- the method further comprises:
- Adjacent frames in the normal estimation result are smoothed using a preset smoothing algorithm.
- the data to be processed is data converted from original VR image data
- the process of obtaining the data to be processed includes:
- the original VR image data is subjected to object region interception processing to obtain region representation data
- Projection dedistortion processing is performed on the region representation data to obtain the data to be processed.
- the target background representation data, the adjusted data, and the target data are all VR image data
- the target background representation data and the adjusted data are fused to obtain the target data.
- the data to be processed is image data converted from original VR image data
- the original VR image data is monocular panoramic image data, binocular panoramic image data, monocular semi-panoramic image data or binocular semi-panoramic image data.
- the target environment map is a high dynamic range imaging HDR environment map.
- obtaining a target environment map and data to be processed includes:
- the background replacement request carries the target environment map and the original VR image data
- the original VR image data is transformed to obtain the data to be processed.
- the present disclosure provides a data generation device, comprising:
- An information acquisition unit used to acquire a target environment map and data to be processed; the data to be processed includes at least one image data; the data to be processed is used to describe the state of the target object in the original background;
- a background generating unit configured to generate target background representation data corresponding to the data to be processed by using the target environment map; the target background described by the target background representation data is different from the original background;
- An object relighting unit is used to perform brightness adjustment processing on the target object in the data to be processed according to the target environment map to obtain adjusted data; the brightness state of the target object in the adjusted data is different from the brightness state of the target object in the data to be processed;
- a data generating unit is configured to generate a target background characterization data and the adjusted data according to the target background characterization data and the adjusted data.
- the target data is used to describe the state of the target object under the target background; the brightness state of the target object in the target data is consistent with the brightness state of the target object in the adjusted data.
- the present disclosure provides an electronic device, the device comprising: a processor and a memory;
- the memory is used to store instructions or computer programs
- the processor is used to execute the instructions or computer programs in the memory so that the electronic device executes the data generation method provided by the present disclosure.
- the present disclosure provides a computer-readable medium, in which instructions or computer programs are stored.
- the device executes the data generating method provided by the present disclosure.
- the present disclosure provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, wherein the computer program contains program codes for executing the data generation method provided by the present disclosure.
- FIG1 is a flow chart of a data generation method provided by an embodiment of the present disclosure.
- FIG2 is a schematic diagram of a process of background replacement processing and object relighting processing provided by an embodiment of the present disclosure
- FIG3 is a schematic diagram of the structure of a data generating device provided by an embodiment of the present disclosure.
- FIG. 4 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure.
- the data generation method provided by the present disclosure includes the following S1-S4. Among them, Figure 1 is a flow chart of a data generation method provided by an embodiment of the present disclosure.
- S1 Obtain a target environment map and data to be processed; the data to be processed includes at least one image data; the data to be processed is used to describe the state of the target object in the original background.
- the data to be processed refers to image data that needs background replacement processing (for example, one frame of image data or an image sequence including multiple frames of image data, etc.). That is, the data to be processed can refer to one piece of image data or an image sequence (for example, a video, etc.).
- the present disclosure does not limit the implementation methods of the above data to be processed.
- the data to be processed may be a frame of image data.
- the data to be processed may be video data (e.g., the data to be processed may be a normal video of a portrait as shown in FIG. 2 ).
- the data to be processed can be used to describe the state of a target object under the original background (for example, movement state, brightness state, etc.).
- the original background refers to the scenery that appears in the data to be processed and is used to set off the target object; and the present disclosure does not limit the original background.
- the original background may refer to other scenery in the live broadcast room except the anchor.
- the target object refers to the object described by the data to be processed (for example, a person, an animal, an object, etc.); and the present disclosure does not limit the target object.
- the target object may refer to the anchor.
- the present disclosure does not limit the implementation methods of the above-mentioned data to be processed.
- the data to be processed may include at least one ordinary image.
- ordinary images refer to image data that conforms to the real reality style.
- the real reality style refers to the image style of image data collected by an ordinary camera from the real world.
- the ordinary image can be used to define a plane (for example, the plane refers to a plane composed of the right side of the screen of the ordinary camera as the positive direction of the X axis, the top of the screen as the positive direction of the Y axis, and the direction from the screen to the photographer as the positive direction of the Z axis).
- the above data to be processed can be implemented using ordinary images, or the data to be processed can also be implemented using ordinary videos (for example, ordinary portrait videos shown in FIG. 2 ).
- the ordinary video refers to an image sequence composed of multiple ordinary images.
- the present disclosure does not limit the method for obtaining the above-mentioned data to be processed.
- the process of obtaining the data to be processed can be specifically as follows: the video data captured by an ordinary camera (for example, an ordinary live video recorded by an ordinary camera for a host in a live broadcast room, etc.) is directly determined as the data to be processed.
- the data to be processed may refer to data converted from the original VR image data, and the process of obtaining the data to be processed may specifically include steps 11 to 13 below.
- Step 11 After acquiring the original VR image data, perform object detection processing on the original VR image data to obtain an object detection result.
- the original VR image data refers to VR image data that needs to undergo background replacement processing (e.g., one frame of VR image data or a VR image sequence including multiple frames of VR image data, similar to VR video data, etc.).
- the original VR image data may refer to a VR live video recorded by a VR shooting device for a host in a live broadcast room.
- the present disclosure does not limit the implementation method of the above original VR image data.
- the original VR image data may be a frame of VR image.
- the original VR image data may be a VR video (for example, the original VR image data may be a VR live video shown in FIG. 2).
- VR images or VR videos are collected by means of pre-arranged VR shooting equipment.
- the present disclosure does not limit the implementation method of the above-mentioned original VR image data.
- the original VR image data can be implemented using monocular panoramic image data, binocular panoramic image data, monocular semi-panoramic image data or binocular semi-panoramic image data (for example, binocular semi-panoramic image data with a resolution of 7680 ⁇ 8404).
- Object detection processing is used to detect objects (e.g., humans, animals, etc.) present in VR image data (e.g., a frame of VR image data or a VR image sequence, etc.); and the present disclosure does not limit the object detection processing.
- the object detection processing may be portrait detection processing.
- the object detection result is used to describe the area occupied by the target object in the original VR image data.
- the present disclosure does not limit the representation method of the object detection result.
- the object detection result can represent the area occupied by the target object in the original VR image data with the help of a rectangular box.
- step 12 Based on the relevant content of step 12 above, it can be known that in some application scenarios (for example, scenarios such as background replacement processing for VR live video), if it is necessary to perform background replacement processing on a certain VR live video (for example, the VR live video shown in Figure 2, etc.), then in order to ensure that the video data after the background replacement is more natural, it is necessary not only to replace the background image of the VR live video, but also to perform re-lighting processing on the portrait (for example, the anchor portrait, etc.) in the VR live video, so as to ensure that the light brightness of the portrait in the video data after the background replacement is coordinated with the light brightness of its background, thereby ensuring that the brightness distribution in the video data after the background replacement is more natural (that is, not abrupt).
- scenarios such as background replacement processing for VR live video
- Step 12 According to the above object detection results, the original VR image data is processed by object region interception to obtain region representation data.
- the region representation data is used to represent the region occupied by the target object in the original VR image data; and the present disclosure does not limit the implementation method of the region representation data.
- the region representation data is a frame of VR image
- the region representation data is an image sequence (for example, video data).
- the present disclosure does not limit the information carried by the above region representation data.
- the region representation data may include an object and part of the background, so that the result of subsequent normal estimation based on the region representation data is more accurate.
- step 12 Based on the relevant content of step 12 above, it can be known that in some application scenarios (for example, scenarios such as background replacement processing for VR live video, etc.), if it is necessary to perform background replacement processing on a certain VR live video (for example, the VR live video shown in Figure 2, etc.), then after obtaining the object detection result (for example, the portrait detection result) for the VR live video, the object area interception processing can be performed on the original VR image data according to the object detection result to obtain area representation data, so that the area representation data can use at least one image data to represent the area occupied by the above target object in the original VR image data, so that the VR live video can be converted into an ordinary video (for example, the ordinary portrait video shown in Figure 2) with the help of the area representation data, so that the re-lighting processing can be completed with the help of the ordinary video.
- an ordinary video for example, the ordinary portrait video shown in Figure 2
- the re-lighting processing can be completed with the help of the ordinary video.
- Step 13 Perform projection dedistortion processing on the above region representation data to obtain the data to be processed.
- the present disclosure does not limit the implementation method of the "projection dedistortion processing" in step 13.
- it can be implemented by any method that can perform projection dedistortion processing on a VR image, and the embodiments of the present disclosure are not limited to this.
- the present disclosure does not limit the implementation method of the "data to be processed” in step 13.
- the data to be processed can be implemented using ordinary image data with a resolution of 1280 ⁇ 2176.
- step 13 Based on the relevant content of step 13 above, it can be known that in some application scenarios (for example, scenarios such as background replacement processing for VR live video, etc.), if it is necessary to perform background replacement processing on a VR live video (for example, the VR live video shown in Figure 2, etc.), then after cutting out the regional representation data corresponding to the portrait from the VR live video, projection de-distortion processing is performed on the regional representation data to obtain an ordinary video (for example, the ordinary video of the portrait shown in Figure 2). In this way, the purpose of converting a VR video into an ordinary video can be achieved, so that the re-lighting processing can be completed with the help of the ordinary video later.
- a VR live video for example, the VR live video shown in Figure 2, etc.
- the data to be processed above is an ordinary image
- the data to be processed refers to an ordinary image converted from the original VR image data.
- the data to be processed is an ordinary video
- the data to be processed refers to an ordinary video converted from the original VR image data.
- the data to be processed may include a Or multiple ordinary images, and the data to be processed can represent the state of the target object under the original background, so that the data to be processed can be used to complete the re-lighting processing of the target object later, so that the brightness state of the target object after the re-lighting processing conforms to the light distribution law recorded in the above target environment map, which is conducive to improving the background replacement effect.
- the target environment map refers to the environment map that is required to be referenced when performing background replacement processing on the above-mentioned data to be processed; and the scene state represented by the target environment map (for example, scene type, scene posture, scene brightness, etc.) is different from the scene state represented by the above-mentioned original background.
- the present disclosure does not limit the implementation method of the environment map.
- it can be implemented using a panoramic map, for example, a spherical environment map, or a cubic environment map.
- the present disclosure does not limit the implementation method of the above-mentioned target environment map.
- it can be implemented using a high dynamic range (High Dynamic Range Imaging, HDR) environment map.
- HDR High Dynamic Range Imaging
- S2 Generate target background representation data corresponding to the data to be processed using the target environment map; the target background described by the target background representation data is different from the original background.
- the target background characterization data is used to characterize the background required for background replacement processing for the above data to be processed; and the present disclosure does not limit the implementation method of the target background characterization data.
- the target background characterization data is also a frame of image data.
- the target background characterization data is also an image sequence.
- the target background characterization data belongs to VR image data (for example, the VR virtual background shown in Figure 2).
- target background refers to the background represented by the target background representation data above; and the target background The target background is determined based on the target environment map described above, so that the scene state represented by the target background is consistent with the scene state represented by the target environment map, thereby making the target background different from the original background.
- the present disclosure does not limit the implementation method of S2 above.
- it can be implemented by any method of generating a VR virtual background using an environment map.
- it since the left and right eyes of the human visual system have parallax, it is necessary to generate a virtual background with parallax to simulate the three-dimensional sense of the VR scene.
- the virtual background corresponding to the left and right eyes can be adjusted horizontally according to the size of the HDR environment map (that is, the "target environment map" above), and the present disclosure does not make specific limitations on this.
- a VR live video for example, the VR live video shown in Figure 2, etc.
- a virtual background of the VR live video for example, the VR virtual background shown in Figure 2
- the VR virtual background shown in Figure 2 can be generated according to the HDR environment map, so that the virtual background can be used to replace the existing background in the VR live video later.
- S3 According to the target environment map, brightness adjustment is performed on the target object in the data to be processed to obtain adjusted data; the brightness state of the target object in the adjusted data is different from the brightness state of the target object in the data to be processed.
- the adjusted data refers to the brightness adjustment processing result of the target object in the processed data, so that the brightness state of the target object in the adjusted data conforms to the light distribution law recorded in the above target environment map, so that the brightness state of the target object in the adjusted data is better coordinated with the brightness state of the above "target background representation data", so that when the target object is merged with the target background representation data, a more natural light distribution state can be presented.
- the present disclosure does not limit the implementation of the above “adjusted data”.
- the adjusted data may include at least one ordinary image.
- the adjusted data may include at least one VR image.
- the VR image for example, may be a panoramic image or a semi-panoramic image.
- the above adjusted data is also a frame of image data; however, if the data to be processed is an image sequence (e.g., video data), then the adjusted data is also an image sequence (e.g., video data).
- the data to be processed is an image sequence (e.g., video data)
- the adjusted data is also an image sequence (e.g., video data).
- the adjusted data above can be implemented using VR images or VR videos.
- the embodiment of the present disclosure does not limit the determination process of the above “adjusted data”, for example, it may include the following steps 21 to 23.
- Step 21 Convert the light information recorded in the target environment map above into an irradiance map.
- the light map is used to describe the light information recorded in the target environment map above.
- the light map is essentially a panoramic image, so that the light map records the sum of all light energies in each direction; and the sum can be determined using the following formula (1).
- ⁇ represents the longitude value
- ⁇ represents the latitude value
- L( ⁇ , ⁇ ) represents the light brightness recorded in the above target environment map at the point where the longitude value is ⁇ and the latitude value is ⁇
- c represents a preset value.
- the present disclosure does not limit the implementation method of the above step 21.
- it can be implemented by any method that can convert the light information recorded in an environment map into an irradiance map.
- step 21 Based on the relevant content of step 21 above, it can be known that in some application scenarios (for example, scenarios such as background replacement processing for VR live videos), if it is necessary to perform background replacement processing for a certain VR live video (for example, the VR live video shown in Figure 2, etc.) based on an HDR environment map (for example, the HDR environment map shown in Figure 2), then after obtaining the HDR environment map, the light information recorded in the HDR environment map can be directly converted into a light map, so that the light map can be used for subsequent re-lighting processing.
- scenarios such as background replacement processing for VR live videos
- an HDR environment map for example, the HDR environment map shown in Figure 2
- Step 22 Perform normal estimation processing on the above data to be processed to obtain the normal estimation result.
- the normal estimation result is used to represent the directional characteristics of each pixel point in the above-mentioned data to be processed; and the present disclosure does not limit the implementation method of the normal estimation result. For example, it can be implemented using a normal map.
- the present disclosure does not limit the data type of the above normal estimation result.
- the normal estimation result is also a frame of image data.
- the to-be-processed data is an image sequence (e.g., video data)
- the normal estimation result is also an image sequence (e.g., video data).
- the present disclosure does not limit the implementation of the "normal estimation processing" in step 22.
- it can be implemented using any method that can perform normal estimation processing on an image data (for example, based on a Pix2Pix network, etc.).
- the present disclosure also provides a possible implementation method of the above "normal estimation result" determination process, which may specifically include: first performing normal estimation processing on the data to be processed to obtain a normal estimation result; then using a preset smoothing algorithm to smooth adjacent frames in the normal estimation result.
- the preset smoothing algorithm is used to smooth a time-series image sequence; and the present disclosure does not limit the preset smoothing algorithm.
- it can be implemented using any method that can smooth an image sequence arranged in a time sequence (for example, an optical flow algorithm or a Raft algorithm).
- the optical flow algorithm is used as an example for explanation below.
- the process of determining the "normal estimation result" above may specifically include the following steps 31-32.
- Step 31 Perform normal estimation processing on the data to be processed to obtain a normal estimation result, so that the normal estimation result includes a normal map corresponding to N frames of ordinary images.
- the normal map corresponding to the nth frame of ordinary image is used to represent the directional characteristics of each pixel in the nth frame of ordinary image.
- n is a positive integer
- n ⁇ N is a positive integer.
- Step 32 Determine the optical flow between any adjacent frames in the data to be processed.
- the present disclosure does not limit the method for obtaining the “optical flow” in step 32 .
- it can be implemented by any method that can calculate the optical flow for two image data.
- Step 33 Using the optical flow between the n-1th frame of the ordinary image and the nth frame of the ordinary image, warping the normal map corresponding to the n-1th frame of the ordinary image to obtain the normal estimation result corresponding to the nth frame of the ordinary image.
- n is a positive integer
- 2 ⁇ n ⁇ N and N is a positive integer.
- the present disclosure does not limit the implementation method of the "warping processing" in step 33.
- it can be implemented by using the warping processing involved in any optical flow algorithm (for example, adding the optical flow between the n-1th frame ordinary image and the nth frame ordinary image, and the normal map corresponding to the n-1th frame ordinary image, etc.).
- Step 34 Perform weighted average processing on the normal calculation result corresponding to the n-th frame of the ordinary image and the normal map corresponding to the n-th frame of the ordinary image to obtain the final normal map corresponding to the n-th frame of the ordinary image, where n is a positive integer, 2 ⁇ n ⁇ N, and N is a positive integer.
- step 22 Based on the relevant content of step 22 above, it can be known that in some application scenarios (for example, scenarios such as background replacement processing for VR live video, etc.), if it is necessary to perform background replacement processing on a certain VR live video (for example, the VR live video shown in Figure 2, etc.), then after converting the VR live video into an ordinary video (for example, the ordinary portrait video shown in Figure 2), you can first perform normal estimation processing on the ordinary video to obtain a normal estimation result; then use the optical flow algorithm to smooth the adjacent frames in the normal estimation result to obtain the final normal estimation result corresponding to the data to be processed, so that the normal estimation result has better timing stability, so that the subsequent re-lighting processing can be achieved based on the normal estimation result.
- a certain VR live video for example, the VR live video shown in Figure 2, etc.
- an ordinary video for example, the ordinary portrait video shown in Figure 2
- Step 23 Determine the adjusted data based on the normal estimation result, the data to be processed, and the light map.
- step 23 does not limit the implementation of step 23.
- two possible implementations are described below.
- the above step 23 may specifically include the following steps 231 to 233.
- Step 231 Based on the normal estimation result, determine the first light brightness table from the above light map Solicit information.
- the first light brightness representation information is used to record the irradiance corresponding to each pixel point in the above to-be-processed data in the above target environment map.
- step 231 does not limit the implementation of step 231 .
- step 231 Based on the relevant content of step 231 above, it can be known that after obtaining the above light map, the normal recorded for each pixel in the above normal estimation result can be used to index the light map to obtain the irradiance corresponding to each pixel, so that the irradiance corresponding to each pixel can be used later to adjust the brightness of the corresponding pixel in the above data to be processed.
- Step 232 According to the first light brightness characterization information, brightness adjustment processing is performed on the data to be processed to obtain a first adjustment result.
- the first adjustment result is a brightness adjustment result of the pointer to the data to be processed, so that the first adjustment result is used to record the brightness after relighting corresponding to each pixel point in the data to be processed.
- the present disclosure does not limit the implementation method of the "brightness adjustment processing" in the above step 232.
- it can be specifically: multiplying the irradiance corresponding to a pixel point by the brightness of the pixel point in the above processed data to obtain the brightness after re-lighting corresponding to the pixel point.
- Step 233 Determine the adjusted data according to the first adjustment result above.
- step 233 it can be specifically: directly determining the first adjustment result as the adjusted data. Since the data to be processed is ordinary data (for example, ordinary images or ordinary videos), the first adjustment result obtained by adjusting the data to be processed also belongs to ordinary data, and thus the adjusted data also belongs to ordinary data.
- the data to be processed is ordinary data (for example, ordinary images or ordinary videos)
- the first adjustment result obtained by adjusting the data to be processed also belongs to ordinary data, and thus the adjusted data also belongs to ordinary data.
- the above data to be processed not only contains the target object, but also contains some background areas that are relatively close to the target object. Therefore, in order to avoid interference caused by the background areas, the present disclosure also provides a possible implementation method of the process of determining the above adjusted data, which may specifically include the following steps 41-42.
- Step 41 Perform object position determination processing on the above data to be processed to obtain object position representation data.
- the object position representation data is used to represent the position of the target object in the data to be processed; and the present disclosure does not limit the implementation method of the object position representation data.
- it can be The process is implemented by matting.
- the present disclosure does not limit the data type of the above object position representation data.
- the object position representation data is also a frame of image data.
- the to-be-processed data is an image sequence (e.g., video data)
- the object position representation data is also an image sequence (e.g., video data).
- the present disclosure does not limit the implementation method of the "object position determination processing" in step 41.
- it can be implemented using any method that can perform object matting estimation processing on an image data (for example, based on MobileNetV3+ASPP network, etc.).
- the present disclosure also provides a possible implementation method of the determination process of the above "object position representation data”, which may specifically include: firstly performing object position determination processing on the data to be processed to obtain object position representation data; and then using a preset smoothing algorithm to smooth adjacent frames in the object position representation data.
- object position representation data may specifically include: firstly performing object position determination processing on the data to be processed to obtain object position representation data; and then using a preset smoothing algorithm to smooth adjacent frames in the object position representation data.
- the relevant content of the preset smoothing algorithm can be found in the relevant content of step 22 above, and for the sake of brevity, it will not be repeated here.
- the object position determination processing can be performed on the data to be processed to obtain object position representation data; then the optical flow algorithm is used to smooth the adjacent frames in the object position representation data to obtain the final object position representation data corresponding to the data to be processed, so that the object position representation data has better timing stability, which is conducive to improving the re-lighting effect.
- step 41 Based on the relevant content of step 41 above, it can be known that in some application scenarios (for example, scenarios such as background replacement processing for VR live video, etc.), if it is necessary to perform background replacement processing on a certain VR live video (for example, the VR live video shown in Figure 2, etc.), then after converting the VR live video into an ordinary video (for example, the portrait ordinary video shown in Figure 2), the object position determination processing (for example, portrait matting estimation processing) can be performed on the ordinary video first to obtain object position representation data (for example, portrait matting results); then the optical flow algorithm is used to smooth the adjacent frames in the object position representation data to obtain the final object position representation data corresponding to the data to be processed, so that the object position representation data has better timing stability, so that the subsequent re-lighting processing can be achieved based on the object position representation data.
- the object position determination processing for example, portrait matting estimation processing
- object position representation data for example, portrait matting results
- the optical flow algorithm is used to smooth the adjacent frames in the object position representation data to obtain the final
- step 41 does not limit the execution time of step 41.
- it may be earlier than the execution time of step 41 below.
- the execution time of step 42 is sufficient.
- Step 42 determining adjusted data according to the above object position representation data and the above first adjustment result, so that the adjusted data records the adjusted brightness corresponding to the above target object.
- step 42 it can be specifically: according to the above object position characterization data, the adjusted data is extracted from the above first adjustment result, so that the adjusted data only records the brightness after relighting corresponding to the target object.
- the following is an example.
- the above step 42 may specifically include the following steps 421-423.
- Step 421 Convert the first adjustment result into first VR image data.
- the first VR image data refers to the first adjustment result expressed in the VR image data format; and the present disclosure does not limit the determination process of the first VR image data, for example, it can be specifically: the first adjustment result is subjected to reverse reprojection processing to obtain the first VR image data, so that the first VR image data conforms to the VR image data format (especially, conforms to the image data format of the above "original VR image data”).
- the "reverse reprojection processing” and the above “projection dedistortion” are two opposite processing processes; and the present disclosure does not limit the implementation method of the "reverse reprojection processing".
- the process of determining the first VR image data above may specifically be: according to a preset VR image data format, converting the first adjustment result above into the first VR image data, so that the first VR image data conforms to the preset VR image data format.
- the preset VR image data format is determined according to the format requirements of the target data below; and the preset VR image data format may be pre-set, for example, the preset VR image data format is the same as the image data format of the "original VR image data" above.
- step 421 Based on the relevant content of step 421 above, it can be known that in some application scenarios (for example, scenarios such as background replacement processing for VR live video), if it is necessary to perform background replacement processing on a VR live video (for example, the VR live video shown in FIG. 2), when the VR live video is converted into a normal video (for example, the portrait normal video shown in FIG. 2), and the first adjustment result corresponding to the normal video is obtained (for example, the output of the “normal estimation and relighting” module shown in FIG. 2 Data), the first adjustment result is reversely reprojected to obtain a VR video in a video format that conforms to the above “original VR video”, so that the VR video can better represent the first adjustment result in accordance with the VR video format.
- a VR live video for example, the VR live video shown in FIG. 2
- a normal video for example, the portrait normal video shown in FIG. 2
- the first adjustment result corresponding to the normal video for example, the output of the “normal estimation and relighting” module shown
- Step 422 Convert the above object position representation data into second VR image data.
- the second VR image data refers to the above object position representation data represented in the VR image data format; and the present disclosure does not limit the determination process of the second VR image data to be similar to the above "determination process of the first VR image data". For ease of understanding, it is explained below with examples.
- the above “determination process of the second VR image data” may specifically be: reverse reprojecting the object position representation data to obtain the second VR image data, so that the second VR image data conforms to the VR image data format (especially, conforms to the image data format of the above “original VR image data”).
- the “reverse reprojection process” and the above “projection dedistortion” are two opposite processes; and the present disclosure does not limit the implementation method of the “reverse reprojection process”.
- step 422 Based on the relevant content of step 422 above, it can be known that in some application scenarios (for example, scenarios such as background replacement processing for VR live video, etc.), if it is necessary to perform background replacement processing on a certain VR live video (for example, the VR live video shown in Figure 2, etc.), then when the VR live video is converted into an ordinary video (for example, the portrait ordinary video shown in Figure 2), and after obtaining the object position representation data corresponding to the ordinary video (for example, the output data of the "portrait cutout" module shown in Figure 2), the object position representation data is reversely reprojected to obtain a VR video in the video format that conforms to the above "original VR video", so that the VR video can better represent the object position representation data according to the VR video format.
- a certain VR live video for example, the VR live video shown in Figure 2, etc.
- Step 423 Determine adjusted data according to the first VR image data and the second VR image data.
- adjusted data can be determined according to the first VR image data and the second VR image data (for example, the first VR image data and the second VR image data are fused to obtain the adjusted data; or, according to the second VR image data, the first VR image data is subjected to a cutout process to obtain the adjusted data), so that the adjusted data can represent the re-lit target object according to the image data format of the above-mentioned "original VR image data", so that the brightness state of the target object in the adjusted data is different from that of the target object in the data to be processed.
- the present disclosure does not limit the implementation method of the "fusion processing" in step 423.
- it can be implemented by any method that can fuse two VR image data.
- the normal estimation result can be first used to index the corresponding irradiance from the above light map, and use the irradiance to perform brightness adjustment processing on the above data to be processed to obtain a first adjustment result, so that the adjusted data corresponding to the data to be processed can be determined based on the first adjustment result, so that the brightness state of the target object in the adjusted data is different from the brightness state of the target object in the data to be processed.
- the ordinary video can be firstly subjected to portrait matting estimation processing and normal estimation processing; then, using the normal estimation result, the corresponding irradiance is indexed from the light map corresponding to the HDR environment map, and the irradiance is multiplied by the brightness in the ordinary video to obtain the result after relighting; then, the portrait matting estimation results (that is, the above HDR environment map) are respectively converted into an ordinary video (for example, the portrait ordinary video shown in FIG. 2 ).
- the "object position representation data" and the re-lighting result are respectively inversely mapped and projected to the video format of the above VR live video to obtain the VR video corresponding to the portrait matting estimation result (that is, the "second VR video” above) and the VR video corresponding to the re-lighting result (that is, the "first VR video” above); finally, the VR video corresponding to the portrait matting estimation result and the VR video corresponding to the re-lighting result are fused to obtain the portrait re-lighting VR video corresponding to the VR live video (that is, the "adjusted data” above), so that the brightness state of the portrait in the portrait re-lighting VR video is different from the brightness state of the portrait in the VR live video, so that the purpose of re-lighting an object in a VR video can be achieved.
- the present disclosure also provides another possible implementation of the process of determining the above “adjusted data”, which may specifically include the following steps 51 to 54.
- Step 51 Convert the light information recorded in the target environment map above into an irradiance map.
- step 51 please refer to step 21 above.
- Step 52 Perform normal estimation processing on the above data to be processed to obtain a normal estimation result.
- step 51 please refer to step 22 above.
- Step 53 Perform object position determination processing on the above data to be processed to obtain object position representation data.
- step 53 please refer to the above step 41.
- Step 54 Determine adjusted data based on the above object position representation data, the above normal estimation result, the above data to be processed, and the above light map.
- step 54 does not limit the implementation of step 54.
- it may specifically include the following steps 541 to 543.
- Step 541 extract object representation data from the data to be processed according to the above object position representation data.
- the object characterization data is used to characterize the state of the target object in the data to be processed (for example, the position, brightness, etc.).
- step 541 can specifically be: fusing the object position representation data with the data to be processed to obtain the object representation data, so that the object representation data also belongs to the video data, thereby enabling the object representation data to represent the target object recorded in the data to be processed.
- Step 542 extracting an object normal result from the normal estimation result according to the object position characterization data, wherein the object normal result is used to characterize the normal state of the target object in the data to be processed.
- Step 543 Determine adjusted data according to the above object normal result, the above object representation data, and the above light map.
- step 543 does not limit the implementation of the above step 543.
- it may specifically include the following steps 5431 to 5433.
- Step 5431 According to the above object normal result, determine the second light brightness representation information from the above light map.
- the second light brightness representation information is used to record the irradiance corresponding to each pixel in the above object representation data (that is, each pixel in the area occupied by the target object in the above to-be-processed data) in the above target environment map.
- step 5431 does not limit the implementation of step 5431.
- step 5431 Based on the relevant content of step 5431 above, it can be known that after obtaining the above light map, the normal recorded for each pixel point in the area occupied by the target object in the above object normal result can be used to index the light map to obtain the irradiance corresponding to each pixel point, so that the irradiance corresponding to each pixel point can be used later to adjust the brightness of the corresponding pixel point in the above object representation data.
- Step 5432 According to the second light brightness characterization information, brightness adjustment processing is performed on the object characterization data to obtain a second adjustment result.
- the second adjustment result refers to the brightness adjustment processing result for the above object characterization data (ie, the target object), so that the second adjustment result is used to record the brightness after relighting corresponding to each pixel point in the area occupied by the target object.
- step 5432 does not limit the implementation of the "brightness adjustment process" in step 5432 above to be similar to the implementation of the "brightness adjustment process” in step 232 above.
- Step 5433 Determine the adjusted data based on the second adjustment result above.
- step 5433 it may specifically be: directly determining the second adjustment result as adjusted data, so that the adjusted data belongs to ordinary data (for example, ordinary image or ordinary video).
- step 5433 it may specifically be: performing reverse reprojection processing on the second adjustment result to obtain the adjusted data, so that the adjusted data belongs to VR image data (for example, VR image or VR video).
- the above step 5433 can specifically be: according to a preset VR image data format, converting the second adjustment result into adjusted data so that the adjusted data conforms to the preset VR image data format.
- steps 51 to 54 Based on the relevant contents of steps 51 to 54 above, it can be known that in some application scenarios (for example, scenarios such as background replacement processing for VR live video, etc.), if it is necessary to perform background replacement processing on a certain VR live video (for example, the VR live video shown in FIG. 2 ) based on the HDR environment map, then after converting the VR live video into an ordinary video (for example, the ordinary portrait video shown in FIG.
- scenarios such as background replacement processing for VR live video, etc.
- the ordinary video can be firstly subjected to portrait matting estimation processing and normal estimation processing; then, using the portrait matting estimation result (that is, the “object position representation data” above), the ordinary video and the normal estimation result are subjected to portrait extraction processing to obtain object representation data and object normal result; then, with the help of the object representation data, the object normal result, and the like, the portrait extraction processing and the normal estimation processing can be performed on the ordinary video and the normal estimation result.
- portrait matting estimation result that is, the “object position representation data” above
- the ordinary video and the normal estimation result are subjected to portrait extraction processing to obtain object representation data and object normal result; then, with the help of the object representation data, the object normal result, and the like, the portrait extraction processing and the normal estimation processing can be performed on the ordinary video and the normal estimation result.
- the portrait re-lighting VR video corresponding to the VR live video that is, the "adjusted data" above
- the brightness state of the portrait in the portrait re-lighting VR video is different from the brightness state of the portrait in the VR live video, thereby achieving the purpose of re-lighting an object in a VR video.
- S4 Generate target data according to the target background characterization data and the adjusted data; the target data is used to describe the state of the target object under the target background; the brightness state of the target object in the target data is consistent with the brightness state of the target object in the adjusted data.
- the target data refers to the background replacement result of the above data to be processed, so that the target data is used to describe the state of the above target object in the above target background.
- Difference point 1 The background used by the target data (that is, the target background mentioned above) is different from the background used by the data to be processed mentioned above (that is, the original background mentioned above).
- the number of images in the target data is the same as the number of images in the above-mentioned data to be processed, that is, if the data to be processed is a frame of image data, then the target data is also a frame of image data; if the data to be processed is video data (that is, an image data sequence), then the target data is Data also belongs to video data.
- the above target data not only replaces the background, but also adjusts the brightness of the target object, so that the brightness state of the target object in the target data is more coordinated with the brightness state of the background in the target data, which can effectively ensure that the target data presents a more natural light distribution, thereby effectively improving the background replacement effect.
- the present disclosure does not limit the implementation of the above target data, for example, it may include at least one VR image. It can be seen that in a possible implementation, the target data may be a frame of VR image data, or a VR image sequence (for example, VR video), and the present disclosure does not specifically limit this.
- the present disclosure does not limit the determination process of the above target data. For ease of understanding, two situations are described below.
- Case 1 In some application scenarios, when the above target background representation data, the above adjusted data, and the above target data all belong to VR image data, the process of determining the target data can be specifically: fusing the target background representation data with the adjusted data to obtain the target data.
- the target background representation data and the adjusted data both belong to VR image data, and the above target data also belongs to VR image data, then it can be determined that the target data can be generated with the help of the fusion process of two VR image data, so the target background representation data and the adjusted data can be directly fused to obtain the target data, so that the target data can comprehensively represent the background represented by the target background representation data and the re-lighting result of the target object represented by the adjusted data in accordance with the VR image data format.
- the target data determination process may specifically include the following steps 61-62.
- the third VR image data is used to describe the re-lighting result of the above target object according to the VR image data format.
- step 61 does not limit the implementation of the above step 61.
- it can adopt the above step Any implementation method of step 421 is implemented.
- Step 62 The target background representation data is fused with the third VR image data to obtain target data.
- the present disclosure does not limit the implementation method of the "fusion processing" in step 62.
- it can be implemented by any method that can fuse two VR image data.
- the target background characterization data or the adjusted data can be subjected to image data format conversion processing according to the image data format requirements of the above target data, so that the target data can be subsequently determined by means of a fusion method of two image data having the same image data format, which is conducive to better meeting the format requirements of the target data.
- the background replacement process for the data to be processed is specifically as follows: after obtaining the target environment map (for example, the HDR environment map), first use the target environment map to generate target background representation data corresponding to the data to be processed, so that the target background described by the target background representation data conforms to the scene presentation state (for example, the scene posture, the brightness state of the scene, etc.) in the target environment map, so that the target background described by the target background representation data is different from the target background of the data to be processed.
- the target environment map for example, the HDR environment map
- the target background representation data conforms to the scene presentation state (for example, the scene posture, the brightness state of the scene, etc.) in the target environment map, so that the target background described by the target background representation data is different from the target background of the data to be processed.
- the original background appearing in the processing data is processed, and according to the target environment map, the target object in the data to be processed is subjected to brightness adjustment processing to obtain adjusted data, so that the brightness state of the target object in the adjusted data conforms to the light distribution law recorded in the target environment map, so that the brightness state of the target object in the adjusted data is different from the brightness state of the target object in the data to be processed; and then the target data is generated according to the target background characterization data and the adjusted data, so that the brightness state of the target object in the target data is consistent with the brightness state of the target object in the adjusted data.
- the brightness state of the target object in the target data and the brightness state of the background in the target data both conform to the light distribution law recorded in the target environment map
- the brightness state of the target object in the target data and the brightness state of the background in the target data are relatively coordinated (for example, the brightness transition is relatively natural), so that the target data can more naturally describe the state of the target object under the target background (for example, the light brightness state), so that it can be achieved while ensuring that the object and the background light are consistent.
- the purpose of background replacement is to overcome the adverse effects caused by the inability of VR image data to meet the background requirements of some application scenarios.
- the present disclosure does not limit the execution subject of the data generation method provided in the embodiments of the present disclosure.
- the data generation method provided in the embodiments of the present disclosure can be applied to devices with data processing functions such as terminal devices or servers.
- the data generation method provided in the embodiments of the present disclosure can also be implemented with the help of data communication processes between different devices (for example, a terminal device and a server, two terminal devices, or two servers).
- the terminal device can be a smart phone, a computer, a personal digital assistant (PDA) or a tablet computer.
- PDA personal digital assistant
- the server can be an independent server, a cluster server or a cloud server.
- Scenario 1 The data generation method provided in the present disclosure can be applied to a VR live broadcast scenario, and the implementation process in this scenario can include the following steps 71 to 73.
- Step 71 receiving a background replacement request triggered by a target user for original VR image data; the background replacement request carries a target environment map and the original VR image data.
- the target user refers to the request triggerer; and the present disclosure does not limit the target user.
- the target user may refer to the live publisher (for example, the live anchor or other staff in the live studio) to meet the live publisher's need to change the live background.
- the target user may refer to the live viewer to meet the live viewer's need to change the live background.
- the background replacement request is used to request to replace the original background in the original VR image data with the target background described by the target environment map, wherein the original VR image data is used to describe the state of the target object under the original background.
- the present disclosure does not limit the triggering method of the above-mentioned background replacement request.
- it can be specifically: after the target user sets the original VR image data as the image data that needs to be processed for background replacement through some operations, and sets the target environment map as the target background to be used for background replacement processing through other operations, the target user can trigger the background replacement request by executing a preset operation (for example, clicking the "background replacement" button, etc.).
- Step 72 convert the original VR image data to obtain data to be processed.
- step 72 the process of obtaining the "data to be processed" in step 72 is described above.
- Step 73 Based on the data to be processed and the target environment map carried by the background replacement request, background replacement processing and object re-lighting processing are performed on the original VR image data described by the background replacement request to obtain target data, so that the target data can describe the state of the target object under the target background, thereby enabling the target data to meet the background replacement needs of the target user.
- step 73 the process of obtaining the "target data" in step 73 (that is, the background replacement process and the object re-lighting process) can be found in S1-S4 above, and for the sake of brevity, it will not be repeated here.
- Scenario 2 The data generation method provided in the present disclosure can be applied to a binocular scenario, and the implementation process in this scenario can include the following steps 81 to 83.
- Step 81 If the original VR image data is binocular data (for example, binocular panoramic image data or binocular semi-panoramic image data), then after acquiring the original VR image data, the original VR image data is converted into binocular data to be processed, so that the binocular data to be processed includes left-eye data to be processed and right-eye data to be processed (for example, left-eye ordinary image data + right-eye ordinary image data).
- binocular data for example, binocular panoramic image data or binocular semi-panoramic image data
- Step 82 After acquiring the target environment map, using the target environment map, generate binocular target background representation data corresponding to the data to be processed, so that the binocular target background representation data includes target background representation data corresponding to the left eye and target background representation data corresponding to the right eye.
- Step 83 After obtaining the target environment map, brightness adjustment is performed on the target object in the binocular data to be processed according to the target environment map to obtain binocular adjusted data, so that the binocular adjusted data includes the adjusted data corresponding to the left eye + the adjusted data corresponding to the right eye.
- Step 84 Generate binocular target data according to binocular target background representation data and binocular adjusted data (for example, generate target data corresponding to the left eye according to the target background representation data corresponding to the left eye and the adjusted data corresponding to the left eye, and generate target data corresponding to the right eye according to the target background representation data corresponding to the right eye and the adjusted data corresponding to the left eye).
- the adjusted data corresponding to the right eye is used to generate the target data corresponding to the right eye).
- the target data corresponding to the left eye is used to describe the state of the target object under the target background from the perspective of the left eye; the target data corresponding to the right eye is used to describe the state of the target object under the target background from the perspective of the right eye.
- the certain data corresponding to the left eye and the certain data corresponding to the right eye can be determined by the data processing process shown in the above S1 to S4 for the certain data. For the sake of brevity, they will not be repeated here.
- the VR image data of the left eye can be processed to obtain a series of data corresponding to the left eye (e.g., target background representation data, adjusted data, target data), and the VR image data of the right eye can be processed to obtain a series of data corresponding to the right eye (e.g., target background representation data, adjusted data, target data), so that the final generated binocular target data can be determined based on these data, so that the binocular target data can represent the state of the target object under the target background under the binocular perspective, thereby realizing background replacement processing and object re-lighting processing in the binocular scene.
- a series of data corresponding to the left eye e.g., target background representation data, adjusted data, target data
- the VR image data of the right eye can be processed to obtain a series of data corresponding to the right eye (e.g., target background representation data, adjusted data, target data)
- the binocular target data can represent the state of the target object under the target background under the binocular perspective, thereby realizing
- Figure 3 is a schematic diagram of the structure of a data generation device provided by the embodiment of the present disclosure. It should be noted that for the technical details of the data generation device provided by the embodiment of the present disclosure, please refer to the relevant content of the data generation method above.
- the data generating device 300 provided in the embodiment of the present disclosure includes:
- the information acquisition unit 301 is configured to acquire a target environment map and data to be processed; the data to be processed includes at least one image data; the data to be processed is used to describe the state of the target object in the original background;
- the background generation unit 302 is configured to generate target background representation data corresponding to the data to be processed using the target environment map; the target background described by the target background representation data is different from the original background;
- the object relighting unit 303 is configured to perform brightness adjustment processing on the target object in the to-be-processed data according to the target environment map to obtain adjusted data; the brightness state of the target object in the adjusted data is different from the brightness state of the target object in the to-be-processed data;
- the data generating unit 304 is configured to generate target data according to the target background characterization data and the adjusted data; the target data is used to describe the state of the target object under the target background; the brightness state of the target object in the target data is kept constant with the brightness state of the target object in the adjusted data. Consistent.
- the object relighting unit 303 includes:
- An information conversion subunit configured to convert light information recorded in a target environment map into a light map
- the normal estimation subunit is configured to perform normal estimation processing on the data to be processed to obtain a normal estimation result
- the first determination subunit is configured to determine the adjusted data according to the normal estimation result, the data to be processed and the light map.
- the first determining subunit includes:
- a first determination subunit is configured to determine first light brightness representation information from the light map according to the normal estimation result
- a first adjustment subunit is configured to perform brightness adjustment processing on the data to be processed according to the first light brightness representation information to obtain a first adjustment result
- the second determining subunit is configured to determine adjusted data according to the first adjustment result.
- the data generating device 300 further includes:
- a position determination unit configured to perform object position determination processing on the data to be processed to obtain object position representation data
- the second determining subunit is specifically configured to determine the adjusted data according to the object position representation data and the first adjustment result, so that the adjusted brightness corresponding to the target object is recorded in the adjusted data.
- the adjusted data is virtual reality VR image data
- the data to be processed is image data converted from the original VR image data
- the second determination subunit is specifically configured to: convert the first adjustment result into first VR image data; convert the object position representation data into second VR image data; and determine the adjusted data according to the first VR image data and the second VR image data.
- the data generating device 300 further includes:
- a position determination unit configured to perform object position determination processing on the data to be processed to obtain object position representation data
- the first determination subunit is specifically configured to determine the adjusted data according to the object position representation data, the normal estimation result, the data to be processed and the light map.
- the first determining subunit includes:
- a first extraction subunit is configured to extract object representation data from the data to be processed according to the object position representation data
- a second extraction subunit is configured to extract an object normal result from the normal estimation result according to the object position representation data
- the third determining subunit is configured to determine the adjusted data according to the object normal result, the object representation data and the light map.
- the third determination subunit is specifically configured to: determine second light brightness representation information from the light map based on the object normal result; perform brightness adjustment processing on the object representation data based on the second light brightness representation information to obtain a second adjustment result; and determine the adjusted data based on the second adjustment result.
- the data to be processed, the object position representation data, and the normal estimation result are all image sequences
- the data generating device 300 further includes:
- a first smoothing unit is configured to perform smoothing processing on adjacent frames in the object position representation data using a preset smoothing algorithm
- the second smoothing unit is configured to use a preset smoothing algorithm to smooth adjacent frames in the normal estimation result.
- the data to be processed is data converted from original VR image data
- the information acquisition unit 301 includes:
- the object detection subunit is configured to perform object detection processing on the original VR image data to obtain an object detection result
- the region interception subunit is configured to perform object region interception processing on the original VR image data according to the object detection result to obtain region representation data;
- the projection processing subunit is configured to perform projection de-distortion processing on the regional representation data to obtain data to be processed.
- the target background representation data, the adjusted data, and the target data are all VR image data
- the data generation unit 304 is specifically configured to: combine the target background representation data with the adjusted data Perform fusion processing to obtain target data.
- the data to be processed is image data converted from original VR image data;
- the original VR image data is monocular panoramic image data, binocular panoramic image data, monocular semi-panoramic image data or binocular semi-panoramic image data.
- the target environment map is a high dynamic range imaging HDR environment map.
- the information acquisition unit 301 is specifically configured to: receive a background replacement request triggered by a target user for original VR image data; the background replacement request carries a target environment map and original VR image data; and transform the original VR image data to obtain data to be processed.
- the target environment map is first used to generate target background representation data corresponding to the data to be processed, so that the target background described by the target background representation data conforms to the scene presentation state (for example, the scene posture, the brightness state of the scene, etc.) in the target environment map, so that the target background described by the target background representation data is different from the original background appearing in the data to be processed, and according to the target environment map, the brightness of the target object in the data to be processed is adjusted to obtain adjusted data, so that the brightness state of the target object in the adjusted data conforms to the light distribution law recorded in the target environment map, so that the brightness state of the target object in the adjusted data is different from the brightness state of the target object in the data to be processed; and then, according to the target background representation data
- the brightness state of the target object in the target data and the brightness state of the background in the target data both conform to the light distribution law recorded in the target environment map
- the brightness state of the target object in the target data and the brightness state of the background in the target data are relatively coordinated (for example, the brightness transition is relatively natural), so that the target data can more naturally describe the state of the target object under the target background (for example, the light brightness state), so as to achieve the purpose of background replacement under the premise of ensuring that the object and the background light are coordinated, so as to overcome the adverse effects caused by the inability of VR image data to meet the background requirements of some application scenarios.
- an embodiment of the present disclosure also provides an electronic device, which includes a processor and a memory: the memory is configured to store instructions or computer programs; the processor is configured to execute the instructions or computer programs in the memory, so that the electronic device executes any implementation of the data generation method provided in the embodiment of the present disclosure.
- the terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
- mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
- PDAs personal digital assistants
- PADs tablet computers
- PMPs portable multimedia players
- vehicle-mounted terminals such as vehicle-mounted navigation terminals
- fixed terminals such as digital TVs, desktop computers, etc.
- the electronic device shown in FIG. 4 is only an example and should not bring any limitation to the functions and scope of use
- the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 to a random access memory (RAM) 403.
- a processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404.
- An input/output (I/O) interface 405 is also connected to the bus 404.
- the following devices may be connected to the I/O interface 405: input devices 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 408 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 409.
- the communication device 409 may allow the electronic device 400 to communicate wirelessly or wired with other devices to exchange data.
- FIG. 4 shows an electronic device 400 with various devices, it should be understood that it is not required to implement or have all the devices shown. More or fewer devices may be implemented or have alternatively.
- an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart.
- the computer program can be downloaded and installed from the network through the communication device 409, or installed from the storage device 408, or installed from the ROM402.
- the processing device 401 the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
- the embodiments of the present disclosure also provide a computer-readable medium, in which instructions or computer programs are stored.
- the instructions or computer programs are executed on a device, the device executes any implementation of the data generation method provided in the embodiments of the present disclosure.
- the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two.
- the computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above.
- Computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
- a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device.
- a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried.
- This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above.
- the computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device.
- the program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
- the client and server may communicate using any currently known or future developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network).
- HTTP Hyper Text Transfer Protocol
- Examples of communication networks include a local area network ("LAN”), a wide area network ("WAN”), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
- the computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
- the computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device can execute the method.
- Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages.
- the program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server.
- the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
- LAN local area network
- WAN wide area network
- Internet service provider e.g., AT&T, MCI, Sprint, EarthLink, MSN, GTE, etc.
- each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function.
- the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved.
- each square box in the block diagram and/or flow chart, and the combination of the square boxes in the block diagram and/or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
- the units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit/module does not, in some cases, constitute a limitation on the unit itself.
- exemplary types of hardware logic components include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
- FPGAs field programmable gate arrays
- ASICs application specific integrated circuits
- ASSPs application specific standard products
- SOCs systems on chip
- CPLDs complex programmable logic devices
- a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment.
- a machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
- a machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing.
- a more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
- RAM random access memory
- ROM read-only memory
- EPROM or flash memory erasable programmable read-only memory
- CD-ROM portable compact disk read-only memory
- CD-ROM compact disk read-only memory
- magnetic storage device or any suitable combination of the foregoing.
- At least one (item) means one or more, and “plurality” means two or more.
- “And/or” is used to describe the association relationship of associated objects, indicating that three relationships may exist.
- a and/or B can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural.
- the character “/” generally indicates that the objects associated before and after are in an “or” relationship.
- At least one of the following” or similar expressions refers to any combination of these items, including any combination of single or plural items.
- At least one of a, b or c can mean: a, b, c, "a and b", “a and c", “b and c", or "a and b and c", where a, b, c can be single or multiple.
- the steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two.
- the software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
Landscapes
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Image Processing (AREA)
Abstract
本公开的实施例公开了一种数据生成方法及装置、电子设备和计算机可读介质,该数据生成方法包括:在获取到目标环境贴图之后,先利用该目标环境贴图,生成待处理数据对应的目标背景表征数据,并依据该目标环境贴图,对该待处理数据中目标对象进行亮度调整处理,得到调整后数据;再根据该目标背景表征数据与该调整后数据,生成目标数据,以使该目标数据中背景不同于该待处理数据中背景,并使得该目标数据中目标对象所处的亮度状态不同于该待处理数据中目标对象所处的亮度状态。本公开的实施例可以实现在确保对象与背景光线相协调的前提下进行背景更换的目的,从而能够更好地满足针对VR图像数据的背景需求。
Description
本公开要求于2023年4月21日递交的中国专利申请第202310436249.6号的优先权,在此全文引用上述中国专利申请公开的内容以作为本公开的一部分。
本公开涉及一种数据生成方法、装置、电子设备、计算机可读介质。
随着虚拟现实(Virtual Reality,VR)技术的普及,VR图像数据的应用场景(例如,直播场景等)越来越多。
目前,VR图像数据通常需要在预先布置好VR拍摄设备的房间内进行拍摄,但是因该房间的环境比较简单或者单一,导致VR图像数据所涉及的背景也比较简单或者单一,从而导致VR图像数据无法满足一些应用场景的背景需求,如此影响了VR图像数据的发展潜力。
发明内容
本公开提供了一种数据生成方法、装置、电子设备、计算机可读介质,能够克服因VR图像数据无法满足一些应用场景的背景需求而导致的不良影响。
为了实现上述目的,本公开提供的技术方案如下:
本公开提供一种数据生成方法,所述方法包括:
获取目标环境贴图和待处理数据;所述待处理数据包括至少一个图像数据;所述待处理数据用于描述目标对象在原始背景下所处状态;
利用所述目标环境贴图,生成所述待处理数据对应的目标背景表征数据;所述目标背景表征数据所描述的目标背景不同于所述原始背景;
依据所述目标环境贴图,对所述待处理数据中目标对象进行亮度调整处理,得到调整后数据;所述调整后数据中目标对象所处的亮度状态不同于所
述待处理数据中目标对象所处的亮度状态;
根据所述目标背景表征数据与所述调整后数据,生成目标数据;所述目标数据用于描述所述目标对象在所述目标背景下所处状态;所述目标数据中目标对象所处的亮度状态与所述调整后数据中目标对象所处的亮度状态保持一致。
在一种可能的实施方式下,所述依据所述目标环境贴图,对所述待处理数据中目标对象进行亮度调整处理,得到调整后数据,包括:
将所述目标环境贴图中记录的光线信息转化为光照贴图;
对所述待处理数据进行法向估计处理,得到法向估计结果;
根据所述法向估计结果、所述待处理数据以及所述光照贴图,确定所述调整后数据。
在一种可能的实施方式下,所述根据所述法向估计结果、所述待处理数据以及所述光照贴图,确定所述调整后数据,包括:
依据所述法向估计结果,从所述光照贴图中确定第一光线亮度表征信息;
依据所述第一光线亮度表征信息,对所述待处理数据进行亮度调整处理,得到第一调整结果;
根据所述第一调整结果确定所述调整后数据。
在一种可能的实施方式下,所述根据所述第一调整结果确定所述调整后数据之前,所述方法还包括:
对所述待处理数据进行对象位置确定处理,得到对象位置表征数据;
所述根据所述第一调整结果确定所述调整后数据,包括:
根据所述对象位置表征数据和所述第一调整结果,确定所述调整后数据,以使所述调整后数据中记录有所述目标对象对应的调整后亮度。
在一种可能的实施方式下,所述调整后数据属于虚拟现实VR图像数据;所述待处理数据为利用原始VR图像数据转化所得的图像数据;
所述根据所述对象位置表征数据和所述第一调整结果,确定所述调整后数据,包括:
将所述第一调整结果转化为第一VR图像数据;
将所述对象位置表征数据转化为第二VR图像数据;
根据所述第一VR图像数据和所述第二VR图像数据,确定所述调整后
数据。
在一种可能的实施方式下,所述根据所述法向估计结果、所述待处理数据以及所述光照贴图,确定所述调整后数据之前,所述方法还包括:
对所述待处理数据进行对象位置确定处理,得到对象位置表征数据;
所述根据所述法向估计结果、所述待处理数据以及所述光照贴图,确定所述调整后数据,包括:
根据所述对象位置表征数据、所述法向估计结果、所述待处理数据以及所述光照贴图,确定所述调整后数据。
在一种可能的实施方式下,所述根据所述法向估计结果、所述待处理数据以及所述光照贴图,确定所述调整后数据,包括:
按照所述对象位置表征数据,从所述待处理数据中提取对象表征数据;
按照所述对象位置表征数据,从所述法向估计结果中提取对象法向结果;
根据所述对象法向结果、所述对象表征数据以及所述光照贴图,确定所述调整后数据。
在一种可能的实施方式下,所述根据所述对象法向结果、所述对象表征数据以及所述光照贴图,确定所述调整后数据,包括:
依据所述对象法向结果,从所述光照贴图中确定第二光线亮度表征信息;
依据所述第二光线亮度表征信息,对所述对象表征数据进行亮度调整处理,得到第二调整结果;
根据所述第二调整结果确定所述调整后数据。
在一种可能的实施方式下,所述待处理数据、所述对象位置表征数据、以及所述法向估计结果均为图像序列;
所述方法还包括:
利用预设平滑算法,对所述对象位置表征数据中相邻帧进行平滑处理;
和/或,
利用预设平滑算法,对所述法向估计结果中相邻帧进行平滑处理。
在一种可能的实施方式下,所述待处理数据为利用原始VR图像数据转化所得的数据;
所述待处理数据的获取过程,包括:
对所述原始VR图像数据进行对象检测处理,得到对象检测结果;
按照所述对象检测结果,对所述原始VR图像数据进行对象区域截取处理,得到区域表征数据;
对所述区域表征数据进行投影去畸变处理,得到所述待处理数据。
在一种可能的实施方式下,所述目标背景表征数据、所述调整后数据、以及所述目标数据均属于VR图像数据;
所述根据所述目标背景表征数据与所述调整后数据,生成目标数据,包括:
将所述目标背景表征数据与所述调整后数据进行融合处理,得到所述目标数据。
在一种可能的实施方式下,所述待处理数据为利用原始VR图像数据转化所得的图像数据;
所述原始VR图像数据为单目全景图像数据、双目全景图像数据、单目半全景图像数据或者双目半全景图像数据。
在一种可能的实施方式下,所述目标环境贴图为高动态范围成像HDR环境贴图。
在一种可能的实施方式下,所述获取目标环境贴图和待处理数据,包括:
接收目标用户针对原始VR图像数据触发的背景更换请求;所述背景更换请求携带有所述目标环境贴图和所述原始VR图像数据;
对所述原始VR图像数据进行转化处理,得到所述待处理数据。
本公开提供了一种数据生成装置,包括:
信息获取单元,用于获取目标环境贴图和待处理数据;所述待处理数据包括至少一个图像数据;所述待处理数据用于描述目标对象在原始背景下所处状态;
背景生成单元,用于利用所述目标环境贴图,生成所述待处理数据对应的目标背景表征数据;所述目标背景表征数据所描述的目标背景不同于所述原始背景;
对象重打光单元,用于依据所述目标环境贴图,对所述待处理数据中目标对象进行亮度调整处理,得到调整后数据;所述调整后数据中目标对象所处的亮度状态不同于所述待处理数据中目标对象所处的亮度状态;
数据生成单元,用于根据所述目标背景表征数据与所述调整后数据,生
成目标数据;所述目标数据用于描述所述目标对象在所述目标背景下所处状态;所述目标数据中目标对象所处的亮度状态与所述调整后数据中目标对象所处的亮度状态保持一致。
本公开提供了一种电子设备,所述设备包括:处理器和存储器;
所述存储器,用于存储指令或计算机程序;
所述处理器,用于执行所述存储器中的所述指令或计算机程序,以使得所述电子设备执行本公开提供的数据生成方法。
本公开提供了一种计算机可读介质,所述计算机可读介质中存储有指令或计算机程序,当所述指令或计算机程序在设备上运行时,使得所述设备执行本公开提供的数据生成方法。
本公开提供了一种计算机程序产品,其包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行本公开提供的数据生成方法的程序代码。
为了更清楚地说明本公开实施例的技术方案,下面将对实施例描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本公开中记载的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其它的附图。
图1为本公开实施例提供的一种数据生成方法的流程图;
图2为本公开实施例提供的一种背景更换处理以及对象重打光处理的过程示意图;
图3为本公开实施例提供的一种数据生成装置的结构示意图;以及
图4为本公开实施例提供的一种电子设备的结构示意图。
下面将结合本公开的实施例中的附图,为了使本技术领域的人员更好地理解本公开方案,下面将结合本公开实施例中的附图,对本公开实施例中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅是本公开一部分实施例,而不是全部的实施例。基于本公开中的实施例,本领域普通技术人
员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本公开保护的范围。
为了更好地理解本公开所提供的技术方案,下面先结合一些附图对本公开提供的数据生成方法进行说明。如图1所示,本公开实施例提供的数据生成方法,包括下文S1-S4。其中,该图1为本公开实施例提供的一种数据生成方法的流程图。
S1:获取目标环境贴图和待处理数据;该待处理数据包括至少一个图像数据;该待处理数据用于描述目标对象在原始背景下所处状态。
其中,待处理数据是指需要进行背景更换处理的图像数据(比如,一帧图像数据或者包括多帧图像数据的图像序列等)。也就是,该待处理数据可以是指一张图像数据,也可以是指一个图像序列(例如,视频等)。
另外,本公开不限定上文待处理数据的实施方式,例如,在一些应用场景(例如,图像背景更换场景)下,该待处理数据可以为一帧图像数据。又如,在另一些应用场景(例如,视频背景更换场景)下,该待处理数据可以属于视频数据(比如,该待处理数据可以为图2所示的人像普通视频等)。
此外,对于上文待处理数据来说,该待处理数据可以用于描述一个目标对象在原始背景下所处状态(例如,移动状态、亮度状态等)。其中,该原始背景是指在该待处理数据中出现的、用于衬托该目标对象的景物;而且本公开不限定该原始背景,例如,在一些应用场景(例如,VR直播视频的视频背景更换场景)下,该原始背景可以是指直播间内除了主播以外的其他景物。该目标对象是指该待处理数据所描述的对象(例如,人、动物、物体等);而且本公开不限定该目标对象,例如,在一些应用场景(例如,VR直播视频的视频背景更换场景)下,该目标对象可以是指该主播。
还有,本公开不限定上文待处理数据的实施方式,例如,在一些应用场景下,该待处理数据可以包括至少一个普通图像。其中,普通图像是指符合真实现实风格的图像数据。该真实现实风格是指利用普通照相机针对真实世界采集所得的图像数据所具有的图像风格。另外,该普通图像可以被用于定义一个平面(例如,该平面是指以该普通照相机的屏幕右侧为X轴正向、以屏幕上方为Y轴正向、并以从屏幕指向拍摄者的方向为Z轴正向所组成的平面)。
可见,在一种可能的实施方式下,上文待处理数据可以采用普通图像进行实施,或者,该待处理数据也可以采用普通视频(例如,图2所示的人像普通视频等)进行实施。其中,该普通视频是指由多个普通图像所组成的图像序列。
再者,本公开不限定上文待处理数据的获取方式,例如,在一些应用场景(比如,将一个普通直播视频转换为背景发生变化的VR直播视频等场景)下,如果该待处理数据属于普通视频,则该待处理数据的获取过程具体可以为:将由普通照相机采集所得的视频数据(例如,由普通照相机针对某个直播间主播所录制的普通直播视频等),直接确定为该待处理数据。
又如,在另一些应用场景(比如,针对一个VR直播视频或者VR图像进行背景更换处理等场景)下,如果该待处理数据属于普通数据(例如,普通视频或者普通图像),则该待处理数据可以是指利用原始VR图像数据转化所得的数据,而且该待处理数据的获取过程具体可以包括下文步骤11-步骤13。
步骤11:在获取到原始VR图像数据之后,对该原始VR图像数据进行对象检测处理,得到对象检测结果。
其中,原始VR图像数据是指需要进行背景更换处理的VR图像数据(例如,一帧VR图像数据或者包括多帧VR图像数据的VR图像序列,类似于VR视频数据等)。例如,该原始VR图像数据可以是指利用VR拍摄设备针对某个直播间主播所录制的VR直播视频。
另外,本公开不限定上文原始VR图像数据的实施方式,例如,在一些应用场景(例如,针对VR图像进行背景更换处理等场景)下,该原始VR图像数据可以为一帧VR图像。又如,在另一些应用场景(例如,针对VR视频进行背景更换处理等场景)下,该原始VR图像数据可以为一个VR视频(比如,该原始VR图像数据可以为图2所示的VR直播视频等)。其中,VR图像或者VR视频均是借助预先布置好的VR拍摄设备所采集的。
此外,本公开也不限定上文原始VR图像数据的实施方式,例如,在一些应用场景下,该原始VR图像数据可以采用单目全景图像数据、双目全景图像数据、单目半全景图像数据或者双目半全景图像数据(比如,分辨率为7680×8404的双目半全景图像数据)进行实施。
还有,本公开不限定上文原始VR图像数据的获取方式,例如,其可以借助预先布置好的VR拍摄设备进行采集。
对象检测处理用于检测VR图像数据(例如,一帧VR图像数据或者一个VR图像序列等)中所存在的对象(例如,人、动物等);而且本公开不限定该对象检测处理,例如,在一些应用场景(比如,针对VR直播视频进行背景更换处理等场景)下,该对象检测处理可以是人像检测处理。
对象检测结果用于描述上文目标对象在原始VR图像数据中所占区域;而且本公开不限定该对象检测结果的表示方式,例如,该对象检测结果可以借助矩形框表示该目标对象在原始VR图像数据中所占区域。
基于上文步骤12的相关内容可知,在一些应用场景(例如,针对VR直播视频进行背景更换处理等场景)下,如果需要针对某个VR直播视频(例如,图2所示的VR直播视频等)进行背景更换处理,则为了确保背景更换后的视频数据更自然,不仅需要更换该VR直播视频的背景图像,还需要针对该VR直播视频中人像(例如,主播人像等)进行重打光处理,以确保背景更换后的视频数据中人像所具有的光线亮度与其背景所具有的光线亮度比较协调,从而确保背景更换后的视频数据中的亮度分布比较自然(也就是,不突兀),故为了更好地实现针对该人像进行重打光处理,需要先针对VR直播视频进行人像检测处理,以便后续能够基于检测结果,实现将该VR直播视频转化为普通视频(例如,图2所示的人像普通视频)的目的,以便后续能够借助该普通视频完成重打光处理。
步骤12:按照上文对象检测结果,对原始VR图像数据进行对象区域截取处理,得到区域表征数据。
其中,区域表征数据用于表征上文目标对象在原始VR图像数据中所占区域;而且本公开不限定该区域表征数据的实施方式,例如,当上文“原始VR图像数据”为一帧VR图像时,该区域表征数据为一帧图像数据。又如,当上文“原始VR图像数据”为VR图像序列(比如,VR视频)时,该区域表征数据为图像序列(比如,视频数据)。
另外,本公开不限定上文区域表征数据所携带的信息,比如,该区域表征数据可以包括对象以及部分背景,以便后续基于该区域表征数据进行法向估计时结果更准确。
基于上文步骤12的相关内容可知,在一些应用场景(例如,针对VR直播视频进行背景更换处理等场景)下,如果需要针对某个VR直播视频(例如,图2所示的VR直播视频等)进行背景更换处理,则在获取到针对该VR直播视频的对象检测结果(例如,人像检测结果)之后,可以按照该对象检测结果,对该原始VR图像数据进行对象区域截取处理,得到区域表征数据,以使该区域表征数据能够借助至少一个图像数据表示出上文目标对象在原始VR图像数据中所占区域,以便后续能够借助该区域表征数据,实现将该VR直播视频转化为普通视频(例如,图2所示的人像普通视频)的目的,以便后续能够借助该普通视频完成重打光处理。
步骤13:对上文区域表征数据进行投影去畸变处理,得到待处理数据。
需要说明的是,本公开不限定步骤13中“投影去畸变处理”的实施方式,例如,其可以采用任意一种能够针对一个VR图像进行投影去畸变处理的方法进行实施,本公开的实施例对此并不限制。
还需要说明的是,本公开不限定步骤13中“待处理数据”的实施方式,例如,当上文“原始VR图像数据”采用分辨率为7680×8404的双目半全景图像数据进行实施时,该待处理数据可以采用分辨率为1280×2176的普通图像数据进行实施。
基于上文步骤13的相关内容可知,在一些应用场景(例如,针对VR直播视频进行背景更换处理等场景)下,如果需要针对某个VR直播视频(例如,图2所示的VR直播视频等)进行背景更换处理,则从该VR直播视频中截取出人像所对应的区域表征数据之后,针对该区域表征数据进行投影去畸变处理,以得到一个普通视频(例如,图2所示的人像普通视频),如此能够实现将一个VR视频转换成普通视频的目的,以便后续能够借助该普通视频完成重打光处理。
基于上文步骤11至步骤13的相关内容可知,在一种可能的实施方式下,如果上文原始VR图像数据为VR图像,则上文待处理数据为普通图像,而且该待处理数据是指由该原始VR图像数据转换成的普通图像。在另一种可能的实施方式下,如果该原始VR图像数据为VR视频,则该待处理数据为普通视频,而且该待处理数据是指由该原始VR图像数据转换成的普通视频。
基于上文“待处理数据”的相关内容可知,该待处理数据可以包括一个
或者多个普通图像,而且该待处理数据能够表示出目标对象在原始背景下所处状态,以便后续能够借助该待处理数据,完成针对该目标对象的重打光处理,以使重打光处理后的目标对象所处的亮度状态符合上文目标环境贴图中所记录的光线分布规律,如此有利于提高背景更换效果。
目标环境贴图是指在针对上文待处理数据进行背景更换处理时所需参考的环境贴图;而且该目标环境贴图所表征的景物状态(例如,景物类型、景物姿态、景物所处亮度等)不同于上文原始背景所表征的景物状态。需要说明的是,本公开不限定环境贴图的实施方式,比如,其可以采用全景图进行实施,例如可以采用球面环境贴图,也可以采用立方体环境贴图进行实施贴图。
另外,本公开不限定上文目标环境贴图的实施方式,例如,如图2所示,其可以采用高动态范围(High Dynamic Range Imaging,HDR)环境贴图进行实施。
基于上文S1的相关内容可知,在一些应用场景(例如,针对VR直播视频进行背景更换处理等场景)下,如果需要针对某个VR直播视频(例如,图2所示的VR直播视频等)进行背景更换处理,则不仅需要将该VR直播视频转换成普通视频(例如,图2所示的人像普通视频),还需要获取目标环境贴图(例如,图2所示的HDR环境贴图),以使该目标环境贴图能够为该VR直播视频的背景更换过程提供新的背景以及光线信息。
S2:利用目标环境贴图,生成待处理数据对应的目标背景表征数据;该目标背景表征数据所描述的目标背景不同于原始背景。
其中,目标背景表征数据用于表征在针对上文待处理数据进行背景更换处理时所需使用的背景;而且本公开不限定该目标背景表征数据的实施方式,例如,如果上文“待处理数据”为一帧图像数据,则该目标背景表征数据也为一帧图像数据。又如,如果上文“待处理数据”为图像序列(比如,视频数据),则该目标背景表征数据也为图像序列。还如,如果上文“待处理数据”为利用原始VR图像数据(例如,图2所示的VR直播视频)转化所得的普通图像数据,则该目标背景表征数据属于VR图像数据(例如,图2所示的VR虚拟背景)。
上文“目标背景”是指上文目标背景表征数据所表征的背景;而且该目
标背景是基于上文目标环境贴图所确定的,以使该目标背景所表征的景物状态与该目标环境贴图所表征的景物状态保持一致,从而使得该目标背景不同于该原始背景。
另外,本公开不限定上文S2的实施方式,例如,其可以采用任意一种利用一个环境贴图生成VR虚拟背景的方法进行实施。又如,在一些应用场景(比如,针对双目全景图像数据或者双目半全景图像数据进行背景更换处理等场景)下,由于人类视觉系统左右眼有视差,故需要生成具有视差的虚拟背景,来模拟VR场景的立体感。基于此可知,可以根据一些视觉规律(例如,人类左右眼之间的距离大概为6厘米、在人类观察3厘米以外物体时左右眼之间构成的角度时差大概为1.15度等规律),将左右眼对应的虚拟背景根据HDR环境贴图(也就是,上文“目标环境贴图”)的尺寸进行水平平移调整处理即可,本公开对此不做具体限定。
基于上文S2的相关内容可知,在一些应用场景(例如,针对VR直播视频进行背景更换处理等场景)下,如果需要基于HDR环境贴图,针对某个VR直播视频(例如,图2所示的VR直播视频等)进行背景更换处理,则可以根据该HDR环境贴图,生成该VR直播视频的虚拟背景(例如,图2所示的VR虚拟背景),以便后续能够利用该虚拟背景替换该VR直播视频中已存在的背景。
S3:依据目标环境贴图,对待处理数据中目标对象进行亮度调整处理,得到调整后数据;该调整后数据中目标对象所处的亮度状态不同于该待处理数据中目标对象所处的亮度状态。
其中,调整后数据是指针对待处理数据中目标对象的亮度调整处理结果,以使该调整后数据中目标对象所处的亮度状态符合上文目标环境贴图中所记录的光线分布规律,从而使得该调整后数据中目标对象所处的亮度状态与上文“目标背景表征数据”所具有的亮度状态之间具有更好地协调性,以使在将该目标对象与该目标背景表征数据进行融合时能够呈现出更自然的光线分布状态。
另外,本公开不限定上文“调整后数据”的实施方式,例如,在一些应用场景下,该调整后数据可以包括至少一个普通图像。又如,在另一些应用场景下,该调整后数据可以包括至少一个VR图像。其中,本公开不限定该
VR图像,例如,其可以是全景图或者半全景图。
此外,在一种可能的实施方式下,如果上文待处理数据为一帧图像数据,则上文调整后数据也为一帧图像数据;但是,如果该待处理数据为图像序列(比如,视频数据),则该调整后数据也为图像序列(比如,视频数据)。
基于上述两段内容可知,在一些应用场景(例如,将一个普通图像转换为背景发生变化的VR图像、将一个普通视频转换为背景发生变化的VR视频、针对一个VR图像进行背景更换处理、或者针对一个VR视频进行背景更换处理等场景)下,上文调整后数据可以采用VR图像或者VR视频进行实施。
另外,本公开实施例不限定上文“调整后数据”的确定过程,例如,其可以包括下文步骤21-步骤23。
步骤21:将上文目标环境贴图中记录的光线信息转化为光照贴图(irradiance map)。
其中,光照贴图用于描述上文目标环境贴图中记录的光线信息。
另外,光照贴图本质上也是一个全景图像,以使该光照贴图中记录了每个方向上所有光线能量的总和;而且该总和可以借助下文公式(1)进行确定。
式中,φ表示经度值;θ表示纬度值;L(φ,θ)表示在上文目标环境贴图中记录的、在经度值为φ而且纬度值为θ这一点上所呈现的光线亮度;c表示一个预先设定的数值。
此外,本公开不限定上文步骤21的实施方式,例如,其可以采用任意一种能够将一个环境贴图中记录的光线信息转化为irradiance map的方法进行实施。
基于上文步骤21的相关内容可知,在一些应用场景(例如,针对VR直播视频进行背景更换处理等场景)下,如果需要基于HDR环境贴图(例如,图2所示的HDR环境贴图),针对某个VR直播视频(例如,图2所示的VR直播视频等)进行背景更换处理,则在获取到该HDR环境贴图之后,可以直接将该HDR环境贴图中记录的光线信息转化为光照贴图,以便后续能够借助该光照贴图实现重打光处理。
步骤22:对上文待处理数据进行法向估计处理,得到法向估计结果。
其中,法向估计结果用于表示上文待处理数据中每个像素点所具有的方向特点;而且本公开不限定该法向估计结果的实施方式,例如,其可以采用法向贴图进行实施。
另外,本公开不限定上文法向估计结果的数据类型,例如,当上文待处理数据为一帧图像数据时,该法向估计结果也为一帧图像数据。又如,当该待处理数据为图像序列(比如,视频数据)时,该法向估计结果也为图像序列(比如,视频数据)。
此外,本公开不限定步骤22中“法向估计处理”的实施方式,例如,其可以采用任意一种能够针对一个图像数据进行法向估计处理的方法(例如,基于Pix2Pix网络等)进行实施。
还有,当上文法向估计结果为图像序列(比如,视频数据)时,为了更好地提高法向估计结果的时序稳定性,本公开还提供了上文“法向估计结果”的确定过程的一种可能的实施方式,其具体可以包括为:先对待处理数据进行法向估计处理,得到法向估计结果;再利用预设平滑算法,对该法向估计结果中相邻帧进行平滑处理。其中,该预设平滑算法用于针对一个时序图像序列进行平滑处理;而且本公开不限定该预设平滑算法,例如,其可以采用任意一种能够针对一个按照时序所排列的图像序列进行平滑处理的方法(例如,光流算法或者Raft算法)进行实施。为了便于理解,下面以光流算法为例进行说明。
作为示例,当上文待处理数据包括按照预设顺序进行排列的N帧普通图像,上文预设平滑算法为光流算法(例如,Dense Inverse Search Optical Flow,DISOpticalFlow等)时,上文“法向估计结果”的确定过程具体可以包括下文步骤31-步骤32。
步骤31:对待处理数据进行法向估计处理,得到法向估计结果,以使该法向估计结果包括N帧普通图像对应的法向贴图。其中,第n帧普通图像对应的法向贴图用于表示该第n帧普通图像中每个像素点所具有的方向特点。其中,n为正整数,n≤N,N为正整数。
步骤32:确定待处理数据中任意相邻帧之间的光流。
需要说明的是,本公开不限定步骤32中“光流”的获取方式,例如,其可以采用任意一种能够针对两个图像数据进行光流计算的方法进行实施。
步骤33:利用第n-1帧普通图像与第n帧普通图像之间的光流,对该第n-1帧普通图像对应的法向贴图进行扭曲(warping)处理,得到该第n帧普通图像对应的法向推算结果。其中,n为正整数,2≤n≤N,N为正整数。
需要说明的是,本公开不限定步骤33中“warping处理”的实施方式,例如,其可以采用任意一种光流算法中所涉及的warping处理(例如,将第n-1帧普通图像与第n帧普通图像之间的光流、以及第n-1帧普通图像对应的法向贴图进行加和处理等处理方式)进行实施。
步骤34:将第n帧普通图像对应的法向推算结果与该第n帧普通图像对应的法向贴图进行加权平均处理,得到该第n帧普通图像对应的最终法向贴图。其中,n为正整数,2≤n≤N,N为正整数。
基于上文步骤31至步骤34的相关内容可知,在一种可能的实施方式下,当获取到待处理数据之后,可以先针对该待处理数据进行法向估计处理,得到法向估计结果;再利用光流算法针对该法向估计结果中相邻帧进行平滑处理,以得到该待处理数据对应的最终的法向估计结果,以使该法向估计结果具有更好的时序稳定性,如此有利于提高重打光效果。
基于上文步骤22的相关内容可知,在一些应用场景(例如,针对VR直播视频进行背景更换处理等场景)下,如果需要针对某个VR直播视频(例如,图2所示的VR直播视频等)进行背景更换处理,则在将该VR直播视频转换为普通视频(例如,图2所示的人像普通视频)之后,可以先针对该普通视频进行法向估计处理,得到法向估计结果;再利用光流算法针对该法向估计结果中相邻帧进行平滑处理,以得到该待处理数据对应的最终的法向估计结果,以使该法向估计结果具有更好的时序稳定性,以便后续能够基于该法向估计结果实现重打光处理。
步骤23:根据上文法向估计结果、上文待处理数据以及上文光照贴图,确定调整后数据。
需要说明的是,本公开不限定步骤23的实施方式,为了便于理解,下面结合两种可能的实施方式进行说明。
在一种可能的实施方式下,上文步骤23具体可以包括下文步骤231-步骤233。
步骤231:依据法向估计结果,从上文光照贴图中确定第一光线亮度表
征信息。
其中,第一光线亮度表征信息用于记录上文待处理数据中各个像素点在上文目标环境贴图中所对应的辐照度。
另外,本公开不限定步骤231的实施方式。
基于上文步骤231的相关内容可知,在获取到上文光照贴图之后,可以利用上文法向估计结果中针对各个像素点所记录的法向,对该光照贴图进行索引,以得到各个像素点对应的辐照度,以便后续能够利用各个像素点对应的辐照度,分别针对相应像素点在上文待处理数据中所具有的亮度进行调整处理。
步骤232:依据上文第一光线亮度表征信息,对待处理数据进行亮度调整处理,得到第一调整结果。
其中,第一调整结果是指针对待处理数据的亮度调整处理结果,以使该第一调整结果用于记录该待处理数据中各个像素点对应的重打光后亮度。
另外,本公开不限定上文步骤232中“亮度调整处理”的实施方式,例如,其具体可以为:将一个像素点对应的辐照度与该像素点在上文待处理数据中所具有的亮度进行相乘,得到该像素点对应的重打光后亮度。
步骤233:根据上文第一调整结果确定调整后数据。
需要说明的是,本公开不限定步骤233的实施方式,例如,其具体可以为:直接将上文第一调整结果确定为调整后数据。其中,因上文待处理数据属于普通数据(例如,普通图像或者普通视频),使得基于该待处理数据调整所得的第一调整结果也属于普通数据,从而使得该调整后数据也属于普通数据。
实际上,上文待处理数据中不仅存在目标对象,还存在一些与该目标对象距离比较近的背景区域,故为了避免该背景区域造成干扰,本公开还提供了上文调整后数据的确定过程的一种可能的实施方式,其具体可以包括下文步骤41-步骤42。
步骤41:对上文待处理数据进行对象位置确定处理,得到对象位置表征数据。
其中,对象位置表征数据用于表征上文目标对象在上文待处理数据中所处位置;而且本公开不限定该对象位置表征数据的实施方式,例如,其可以
采用人像抠图(matting)进行实施。
另外,本公开不限定上文对象位置表征数据的数据类型,例如,当上文待处理数据为一帧图像数据时,该对象位置表征数据也为一帧图像数据。又如,当该待处理数据为图像序列(比如,视频数据)时,该对象位置表征数据也为图像序列(比如,视频数据)。
此外,本公开不限定步骤41中“对象位置确定处理”的实施方式,例如,其可以采用任意一种能够针对一个图像数据进行对象matting估计处理的方法(例如,基于MobileNetV3+ASPP网络等)进行实施。
还有,当上文对象位置表征数据为图像序列(比如,视频数据)时,为了更好地提高对象位置表征数据的时序稳定性,本公开还提供了上文“对象位置表征数据”的确定过程的一种可能的实施方式,其具体可以包括为:先对待处理数据进行对象位置确定处理,得到对象位置表征数据;再利用预设平滑算法,对该对象位置表征数据中相邻帧进行平滑处理。需要说明的是,该预设平滑算法的相关内容请参见上文步骤22的相关内容,为了简要起见,在此不再赘述。
可见,在一种可能的实施方式下,当获取到待处理数据之后,可以先针对该待处理数据进行对象位置确定处理,得到对象位置表征数据;再利用光流算法针对该对象位置表征数据中相邻帧进行平滑处理,以得到该待处理数据对应的最终的对象位置表征数据,以使该对象位置表征数据具有更好的时序稳定性,如此有利于提高重打光效果。
基于上文步骤41的相关内容可知,在一些应用场景(例如,针对VR直播视频进行背景更换处理等场景)下,如果需要针对某个VR直播视频(例如,图2所示的VR直播视频等)进行背景更换处理,则在将该VR直播视频转换为普通视频(例如,图2所示的人像普通视频)之后,可以先针对该普通视频进行对象位置确定处理(例如,人像matting估计处理),得到对象位置表征数据(例如,人像matting结果);再利用光流算法针对该对象位置表征数据中相邻帧进行平滑处理,以得到该待处理数据对应的最终的对象位置表征数据,以使该对象位置表征数据具有更好的时序稳定性,以便后续能够基于该对象位置表征数据实现重打光处理。
需要说明的是,本公开不限定步骤41的执行时间,例如,其早于下文步
骤42的执行时间即可。
步骤42:根据上文对象位置表征数据和上文第一调整结果,确定调整后数据,以使该调整后数据中记录有上文目标对象对应的调整后亮度。
需要说明的是,本公开不限定步骤42的实施方式,例如,其具体可以为:按照上文对象位置表征数据,从上文第一调整结果中提取出调整后数据,以使该调整后数据中只记录了目标对象对应的重打光后亮度。为了便于理解,下面结合示例进行说明。
作为示例,当上文调整后数据属于虚拟现实VR图像数据,而且上文待处理数据为利用原始VR图像数据转化所得的图像数据(比如,一个普通图像数据或者一个普通视频等)时,上文步骤42具体可以包括下文步骤421-步骤423。
步骤421:将上文第一调整结果转化为第一VR图像数据。
其中,第一VR图像数据是指按照VR图像数据格式表示上文第一调整结果;而且本公开不限定该第一VR图像数据的确定过程,例如,其具体可以为:将第一调整结果进行逆向重投影处理,以得到该第一VR图像数据,以使该第一VR图像数据符合VR图像数据格式(尤其是,符合上文“原始VR图像数据”的图像数据格式)。其中,该“逆向重投影处理”与上文“投影去畸变”是两个相反的处理过程;而且本公开不限定该“逆向重投影处理”的实施方式。
又如,在一些应用场景下,上文第一VR图像数据的确定过程,具体可以为:按照预设VR图像数据格式,将上文第一调整结果转化为第一VR图像数据,以使该第一VR图像数据符合该预设VR图像数据格式。其中,该预设VR图像数据格式是根据下文目标数据的格式需求所确定的;而且该预设VR图像数据格式可以预先设定,例如,该预设VR图像数据格式与上文“原始VR图像数据”所具有的图像数据格式相同。
基于上文步骤421的相关内容可知,在一些应用场景(例如,针对VR直播视频进行背景更换处理等场景)下,如果需要针对某个VR直播视频(例如,图2所示的VR直播视频等)进行背景更换处理,则当该VR直播视频转换为普通视频(例如,图2所示的人像普通视频),而且获取到该普通视频对应的第一调整结果(例如,图2所示的“法向估计并重打光”模块的输出
数据)之后,将该第一调整结果进行逆向重投影处理,以得到符合上文“原始VR视频”的视频格式的VR视频,以使该VR视频能够按照VR视频格式更好地表示该第一调整结果。
步骤422:将上文对象位置表征数据转化为第二VR图像数据。
其中,第二VR图像数据是指按照VR图像数据格式表示上文对象位置表征数据;而且本公开不限定该第二VR图像数据的确定过程类似于上文“第一VR图像数据的确定过程”,为了便于理解,下面结合示例进行说明。
作为示例,上文“第二VR图像数据的确定过程”具体可以为:将对象位置表征数据进行逆向重投影处理,以得到该第二VR图像数据,以使该第二VR图像数据符合VR图像数据格式(尤其是,符合上文“原始VR图像数据”的图像数据格式)。其中,该“逆向重投影处理”与上文“投影去畸变”是两个相反的处理过程;而且本公开不限定该“逆向重投影处理”的实施方式。
基于上文步骤422的相关内容可知,在一些应用场景(例如,针对VR直播视频进行背景更换处理等场景)下,如果需要针对某个VR直播视频(例如,图2所示的VR直播视频等)进行背景更换处理,则当该VR直播视频转换为普通视频(例如,图2所示的人像普通视频),而且获取到该普通视频对应的对象位置表征数据(例如,图2所示的“人像抠图”模块的输出数据)之后,将该对象位置表征数据进行逆向重投影处理,以得到符合上文“原始VR视频”的视频格式的VR视频,以使该VR视频能够按照VR视频格式更好地表示该对象位置表征数据。
步骤423:根据上文第一VR图像数据和上文第二VR图像数据,确定调整后数据。
本公开中,在获取到上文第一VR图像数据与上文第二VR图像数据之后,可以根据该第一VR图像数据和该第二VR图像数据,确定调整后数据(比如,将该第一VR图像数据与该第二VR图像数据进行融合处理,得到调整后数据;或者,按照该第二VR图像数据,针对该第一VR图像数据进行抠图处理,得到调整后的数据),以使该调整后数据能够按照上文“原始VR图像数据”所具有的图像数据格式表示出重打光后的目标对象,从而使得该调整后数据中目标对象所处的亮度状态不同于该待处理数据中目标对象所
处的亮度状态。
需要说明的是,本公开不限定步骤423中“融合处理”的实施方式,例如,其可以采用任意一种能够将两个VR图像数据进行融合的方法进行实施。
基于上文步骤231至步骤233的相关内容可知,在一些应用场景下,在获取到法向估计结果之后,可以先利用该法向估计结果,从上文光照贴图中索引出相应的辐照度,并利用该辐照度针对上文待处理数据进行亮度调整处理,得到第一调整结果,以便后续能够基于该第一调整结果,确定该待处理数据对应的调整后数据,以使该调整后数据中目标对象所处的亮度状态不同于该待处理数据中目标对象所处的亮度状态。
基于上文步骤21至步骤23的相关内容可知,在一些应用场景(例如,针对VR直播视频进行背景更换处理等场景)下,如果需要基于HDR环境贴图,针对某个VR直播视频(例如,图2所示的VR直播视频等)进行背景更换处理,则在将该VR直播视频转换为普通视频(例如,图2所示的人像普通视频)之后,可以先针对该普通视频进行人像matting估计处理以及法向估计处理;再利用法向估计结果,从该HDR环境贴图对应的光照贴图中索引其对应的辐照度,并将该辐照度针对该普通视频中的亮度进行相乘,以得到重打光后结果;然后,分别将人像matting估计结果(也就是,上文“对象位置表征数据”)和该重打光后结果分别逆映射投影至上文VR直播视频所具有的视频格式,以得到该人像matting估计结果对应的VR视频(也就是,上文“第二VR视频”)和该重打光后结果对应的VR视频(也就是,上文“第一VR视频”);最后,将该人像matting估计结果对应的VR视频与该重打光后结果对应的VR视频进行融合,得到该VR直播视频对应的人像重打光VR视频(也就是,上文“调整后数据”),以使该人像重打光VR视频中人像所处的亮度状态不同于该VR直播视频中人像所处的亮度状态,如此能够实现针对一个VR视频中对象进行重打光处理的目的。
实际上,本公开还提供了上文“调整后数据”的确定过程另一种可能的实施方式,其具体可以包括下文步骤51-步骤54。
步骤51:将上文目标环境贴图中记录的光线信息转化为光照贴图(irradiance map)。
需要说明的是,步骤51的相关内容请参见上文步骤21。
步骤52:对上文待处理数据进行法向估计处理,得到法向估计结果。
需要说明的是,步骤51的相关内容请参见上文步骤22。
步骤53:对上文待处理数据进行对象位置确定处理,得到对象位置表征数据。
需要说明的是,步骤53的相关内容请参见上文步骤41。
步骤54:根据上文对象位置表征数据、上文法向估计结果、上文待处理数据以及上文光照贴图,确定调整后数据。
需要说明的是,本公开不限定步骤54的实施方式,例如,其具体可以包括下文步骤541-步骤543。
步骤541:按照上文对象位置表征数据,从待处理数据中提取对象表征数据。
其中,对象表征数据用于表征上文目标对象在上文待处理数据中所处状态(例如,处于什么位置、具有什么样的亮度等)。
另外,本公开不限定上文步骤541的实施方式,例如,当上文对象位置表征数据以及上文待处理数据均属于视频数据时,步骤541具体可以为:将该对象位置表征数据与该待处理数据进行融合处理,得到该对象表征数据,以使该对象表征数据也属于视频数据,从而使得该对象表征数据能够表示出在该待处理数据中所记录的目标对象。
步骤542:按照上文对象位置表征数据,从上文法向估计结果中提取对象法向结果。其中,该对象法向结果用于表征上文目标对象在上文待处理数据中所具有的法向状态。
步骤543:根据上文对象法向结果、上文对象表征数据以及上文光照贴图,确定调整后数据。
需要说明的是,本公开不限定上文步骤543的实施方式,例如,其具体可以包括下文步骤5431-步骤5433。
步骤5431:依据上文对象法向结果,从上文光照贴图中确定第二光线亮度表征信息。
其中,第二光线亮度表征信息用于记录上文对象表征数据中各个像素点(也就是,上文待处理数据中目标对象所占区域内各像素点)在上文目标环境贴图中所对应的辐照度。
另外,本公开不限定步骤5431的实施方式。
基于上文步骤5431的相关内容可知,在获取到上文光照贴图之后,可以利用上文对象法向结果中针对目标对象所占区域内各像素点所记录的法向,对该光照贴图进行索引,以得到各像素点对应的辐照度,以便后续能够利用各个像素点对应的辐照度,分别针对相应像素点在上文对象表征数据中所具有的亮度进行调整处理。
步骤5432:依据上文第二光线亮度表征信息,对上文对象表征数据进行亮度调整处理,得到第二调整结果。
其中,第二调整结果是指针对上文对象表征数据(也就是,目标对象)的亮度调整处理结果,以使该第二调整结果用于记录该目标对象所占区域内各个像素点对应的重打光后亮度。
另外,本公开不限定上文步骤5432中“亮度调整处理”的实施方式类似于上文步骤232中“亮度调整处理”的实施方式。
步骤5433:根据上文第二调整结果确定调整后数据。
需要说明的是,本公开不限定步骤5433的实施方式,例如,其具体可以为:直接将上文第二调整结果确定为调整后数据,以使该调整后数据属于普通数据(例如,普通图像或者普通视频)。又如,其具体可以为:将该第二调整结果进行逆向重投影处理,以得到该调整后数据,以使该调整后数据属于VR图像数据(例如,VR图像或者VR视频)。
还如,当上文第二调整结果为一帧图像数据或者图像序列(比如,视频)时,上文步骤5433具体可以为:按照预设VR图像数据格式,将该第二调整结果转化为调整后数据,以使该调整后数据符合该预设VR图像数据格式。
基于上文步骤51至步骤54的相关内容可知,在一些应用场景(例如,针对VR直播视频进行背景更换处理等场景)下,如果需要基于HDR环境贴图,针对某个VR直播视频(例如,图2所示的VR直播视频等)进行背景更换处理,则在将该VR直播视频转换为普通视频(例如,图2所示的人像普通视频)之后,可以先针对该普通视频进行人像matting估计处理以及法向估计处理;再利用人像matting估计结果(也就是,上文“对象位置表征数据”),针对该普通视频以及法向估计结果进行人像提取处理,得到对象表征数据以及对象法向结果;然后,借助对象表征数据、该对象法向结果、以
及该HDR环境贴图对应的光照贴图,直接确定该VR直播视频对应的人像重打光VR视频(也就是,上文“调整后数据”),以使该人像重打光VR视频中人像所处的亮度状态不同于该VR直播视频中人像所处的亮度状态,如此能够实现针对一个VR视频中对象进行重打光处理的目的。
基于上文S3的相关内容可知,在一些应用场景(例如,针对VR直播视频进行背景更换处理等场景)下,如果需要基于HDR环境贴图,针对某个VR直播视频(例如,图2所示的VR直播视频等)进行背景更换处理,则在将该VR直播视频转换为普通视频(例如,图2所示的人像普通视频)之后,可以借助该HDR环境贴图所记录的光线信息,针对该普通视频中人像进行重打光处理,以得到该VR直播视频对应的人像重打光VR视频(也就是,上文“调整后数据”),以使该人像重打光VR视频中人像所处的亮度状态符合该HDR环境贴图中所记录的光线分布规律,从而使得该人像重打光VR视频中人像所处的亮度状态不同于该VR直播视频中人像所处的亮度状态,如此能够实现针对一个VR视频中对象进行重打光处理的目的。
S4:根据目标背景表征数据与调整后数据,生成目标数据;该目标数据用于描述目标对象在目标背景下所处状态;该目标数据中目标对象所处的亮度状态与该调整后数据中目标对象所处的亮度状态保持一致。
其中,目标数据是指针对上文待处理数据的背景更换结果,以使该目标数据用于描述上文目标对象在上文目标背景下所处状态。
实际上,对于上文目标数据来说,该目标数据与上文待处理数据之间不仅存在下文①-②所示的区别,还存在下文③-④所示的共同点。
①区别点一:目标数据所使用的背景(也就是,上文目标背景)不同于上文待处理数据所使用的背景(也就是,上文原始背景)。
②区别点二:目标数据中目标对象所处的亮度状态不同于上文待处理数据中目标对象所处的亮度状态。
③共同点一:目标数据所描述的对象与上文待处理数据所描述的对象相同,均是上文目标对象。
④共同点二:目标数据中图像个数与上文待处理数据中图像个数相同,也就是,如果该待处理数据为一帧图像数据,则该目标数据也为一帧图像数据;如果该待处理数据属于视频数据(也就是,图像数据序列),则该目标数
据也属于视频数据。
基于上文①至④可知,相较于上文待处理数据来说,上文目标数据不仅更换了背景,还调整了目标对象的亮度,以使该目标数据中目标对象所处的亮度状态与该目标数据中背景所处的亮度状态比较协调,如此能够有效地保证该目标数据呈现出比较自然的光线分布,从而能够有效地提高背景更换效果。
另外,本公开不限定上文目标数据的实施方式,例如,其可以包括至少一个VR图像。可见,在一种可能的实施方式下,该目标数据可以为一帧VR图像数据,也可以为一个VR图像序列(比如,VR视频),本公开对此不做具体限定。
此外,本公开不限定上文目标数据的确定过程,为了便于理解,下面结合两种情况进行说明。
情况1,在一些应用场景下,当上文目标背景表征数据、上文调整后数据、以及上文目标数据均属于VR图像数据时,该目标数据的确定过程具体可以为:将该目标背景表征数据与该调整后数据进行融合处理,得到该目标数据。
可见,在一种可能的实施方式下,在获取到上文目标背景表征数据以及上文调整后数据之后,如果确定该目标背景表征数据与该调整后数据均属于VR图像数据,而且上文目标数据也属于VR图像数据,则可以确定该目标数据可以借助两个VR图像数据的融合过程进行生成,故可以直接将该目标背景表征数据与该调整后数据进行融合处理,以得到该目标数据,以使该目标数据能够按照VR图像数据格式综合表示出由该目标背景表征数据所表征的背景以及由该调整后数据所表征的目标对象的重打光结果。
情况2,在一些应用场景下,当上文目标背景表征数据以及上文目标数据均属于VR图像数据,但是上文调整后数据属于普通图像数据时,该目标数据的确定过程具体可以包括下文步骤61-步骤62。
步骤61:将上文调整后数据转化为第三VR图像数据。
其中,第三VR图像数据用于按照VR图像数据格式描述出上文目标对象的重打光结果。
另外,本公开不限定上文步骤61的实施方式,例如,其可以采用上文步
骤421的任一实施方式进行实施。
步骤62:将上文目标背景表征数据与上文第三VR图像数据进行融合处理,得到目标数据。
需要说明的是,本公开不限定步骤62中“融合处理”的实施方式,例如,其可以采用任意一种能够将两个VR图像数据进行融合的方法进行实施。
基于上文步骤61至步骤62的相关内容可知,在一种可能的实施方式下,在获取到上文目标背景表征数据以及上文调整后数据之后,如果确定该目标背景表征数据所具有的图像数据格式不同于该调整后数据所具有的图像数据格式,则可以依据上文目标数据的图像数据格式需求,将该目标背景表征数据或者该调整后数据进行图像数据格式转化处理,以便后续能够借助针对两个具有相同图像数据格式的图像数据进行融合的方式,确定出该目标数据,如此有利于更好地满足该目标数据的格式需求。
基于上文S1至S4的相关内容可知,对于包括至少一个图像数据的待处理数据(例如,一个VR图像数据)来说,针对该待处理数据的背景更换过程具体为:在获取到目标环境贴图(例如,HDR环境贴图)之后,先利用该目标环境贴图,生成该待处理数据对应的目标背景表征数据,以使该目标背景表征数据所描述的目标背景符合该目标环境贴图中的景物呈现状态(例如,景物姿态、景物所处的亮度状态等),从而使得该目标背景表征数据所描述的目标背景不同于该待处理数据中所出现的原始背景,并依据该目标环境贴图,对该待处理数据中目标对象进行亮度调整处理,得到调整后数据,以使该调整后数据中目标对象所处的亮度状态符合该目标环境贴图中所记录的光线分布规律,从而使得该调整后数据中目标对象所处的亮度状态不同于该待处理数据中目标对象所处的亮度状态;再根据该目标背景表征数据与该调整后数据,生成目标数据,以使该目标数据中目标对象所处的亮度状态与该调整后数据中目标对象所处的亮度状态保持一致。其中,因该目标数据中目标对象所处的亮度状态以及该目标数据中背景所处的亮度状态均符合该目标环境贴图中所记录的光线分布规律,使得该目标数据中目标对象所处的亮度状态与该目标数据中背景所处的亮度状态之间比较协调(例如,亮度过渡比较自然),从而使得该目标数据能够更自然地描述出该目标对象在该目标背景下所处状态(例如,光线亮度状态),如此能够实现在确保对象与背景光线相
协调的前提下进行背景更换的目的,以克服因VR图像数据无法满足一些应用场景的背景需求而导致的不良影响。
另外,本公开不限定本公开实施例提供的数据生成方法的执行主体,例如,本公开实施例提供的数据生成方法可以应用于终端设备或服务器等具有数据处理功能的设备。又如,本公开实施例提供的数据生成方法也可以借助不同设备(例如,终端设备与服务器、两个终端设备、或者两个服务器)之间的数据通信过程进行实现。其中,终端设备可以为智能手机、计算机、个人数字助理(Personal Digital Assistant,PDA)或平板电脑等。服务器可以为独立服务器、集群服务器或云服务器。
此外,本公开不限定数据生成方法的应用场景,为了便于理解,下面结合两个场景进行说明。
场景一:本公开提供的数据生成方法可以应用于VR直播场景,而且该场景下的实现过程可以包括下文步骤71-步骤73。
步骤71:接收目标用户针对原始VR图像数据触发的背景更换请求;该背景更换请求携带有目标环境贴图和该原始VR图像数据。
其中,目标用户是指请求触发者;而且本公开不限定该目标用户,例如,当上文原始VR图像数据为VR直播图像数据或者VR直播视频数据时,该目标用户可以是指直播发布者(比如,直播主播或者直播间的其他工作人员),以满足直播发布者针对直播背景的更换需求。又如,当上文原始VR图像数据为VR直播图像数据或者VR直播视频数据时,该目标用户可以是指直播观看者,以满足该直播观看者针对直播背景的更换需求。
背景更换请求用于请求将原始VR图像数据中的原始背景替换为目标环境贴图所描述的目标背景。其中,该原始VR图像数据用于描述目标对象在原始背景下所处状态。
另外,本公开不限定上文背景更换请求的触发方式,例如,其具体可以为:在目标用户通过一些操作将原始VR图像数据设置为需要进行背景更换处理的图像数据,并通过另一些操作将目标环境贴图设置为背景更换处理时所需使用的目标背景之后,该目标用户可以通过执行预设操作(比如,点击“背景更换”这一按钮等操作),以触发该背景更换请求。
步骤72:对原始VR图像数据进行转化处理,得到待处理数据。
需要说明的是,步骤72中“待处理数据”的获取过程请参见上文。
步骤73:依据待处理数据以及背景更换请求所携带的目标环境贴图,实现针对该背景更换请求所描述的原始VR图像数据进行背景更换处理以及对象重打光处理,得到目标数据,以使该目标数据能够描述出目标对象在目标背景下所处状态,从而使得该目标数据能够满足目标用户的背景更换需求。
需要说明的是,步骤73中“目标数据”的获取过程(也就是,背景更换处理以及对象重打光处理)请参见上文S1-S4,为了简要起见,在此不再赘述。
基于上文步骤71至步骤73的相关内容可知,对于VR直播场景中的VR图像数据(比如,VR直播视频)来说,用户可以针对已经录制好的VR图像数据进行背景更换处理以及对象重打光处理,以使处理后的VR图像数据能够更自然的表示出直播对象在新背景下所处状态,从而使得VR直播图像数据不再拘泥于VR直播录制场景中的真实背景,如此能够有效地克服因VR图像数据无法满足一些直播背景需求而导致的不良影响。
场景二:本公开提供的数据生成方法可以应用于双目场景,而且该场景下的实现过程可以包括下文步骤81-步骤83。
步骤81:若原始VR图像数据属于双目数据(比如,双目全景图像数据或者双目半全景图像数据),则在获取到原始VR图像数据之后,将该原始VR图像数据转化为双目的待处理数据,以使该双目的待处理数据包括左目的待处理数据和右目的待处理处理(比如,左目普通图像数据+右目普通图像数据)。
步骤82:在获取到目标环境贴图之后,利用该目标环境贴图,生成该待处理数据对应的双目的目标背景表征数据,以使该双目的目标背景表征数据包括左目对应的目标背景表征数据和右目对应的目标背景表征数据。
步骤83:在获取到目标环境贴图之后,依据该目标环境贴图,对双目的待处理数据中目标对象进行亮度调整处理,得到双目的调整后数据,以使该双目的调整后数据包括左目对应的调整后数据+右目对应的调整后的数据。
步骤84:根据双目的目标背景表征数据与双目的调整后数据,生成双目的目标数据(比如,根据左目对应的目标背景表征数据与左目对应的调整后数据,生成左目对应的目标数据,并且根据右目对应的目标背景表征数据与
右目对应的调整后数据,生成右目对应的目标数据)。其中,该左目对应的目标数据用于描述左目视角下目标对象在目标背景下所处状态;该右目对应的目标数据用于描述右目视角下目标对象在目标背景下所处状态。
需要说明的是,对于上文步骤81-步骤84来说,左目对应的某种数据以及右目对应的某种数据均可以采用上文S1-S4中针对该某种数据所示的数据处理过程进行确定,为了简要起见,在此不再赘述。
基于上文步骤81至步骤84的相关内容可知,对于双目场景来说,在获取到双目的VR图像数据之后,可以针对左目的VR图像数据进行处理,以得到该左目对应的一系列数据(比如,目标背景表征数据、调整后数据、目标数据),并且针对右目的VR图像数据进行处理,以得到该右目对应的一系列数据(比如,目标背景表征数据、调整后数据、目标数据),以便后续能够基于这些数据,确定最终生成的双目的目标数据,以使该双目的目标数据能够表示出双目视角下目标对象在目标背景下所处状态,如此能够实现双目场景下的背景更换处理以及对象重打光处理。
基于本公开实施例提供的数据生成方法,本公开实施例还提供了一种数据生成装置,下面结合图3进行解释和说明。其中,图3为本公开实施例提供的一种数据生成装置的结构示意图。需要说明的是,本公开实施例提供的数据生成装置的技术详情,请参照上文数据生成方法的相关内容。
如图3所示,本公开实施例提供的数据生成装置300,包括:
信息获取单元301,被配置为获取目标环境贴图和待处理数据;待处理数据包括至少一个图像数据;待处理数据用于描述目标对象在原始背景下所处状态;
背景生成单元302,被配置为利用目标环境贴图,生成待处理数据对应的目标背景表征数据;目标背景表征数据所描述的目标背景不同于原始背景;
对象重打光单元303,被配置为依据目标环境贴图,对待处理数据中目标对象进行亮度调整处理,得到调整后数据;调整后数据中目标对象所处的亮度状态不同于待处理数据中目标对象所处的亮度状态;
数据生成单元304,被配置为根据目标背景表征数据与调整后数据,生成目标数据;目标数据用于描述目标对象在目标背景下所处状态;目标数据中目标对象所处的亮度状态与调整后数据中目标对象所处的亮度状态保持
一致。
在一种可能的实施方式下,对象重打光单元303,包括:
信息转化子单元,被配置为将目标环境贴图中记录的光线信息转化为光照贴图;
法向估计子单元,被配置为对待处理数据进行法向估计处理,得到法向估计结果;
第一确定子单元,被配置为根据法向估计结果、待处理数据以及光照贴图,确定调整后数据。
在一种可能的实施方式下,第一确定子单元,包括:
第一确定子单元,被配置为依据法向估计结果,从光照贴图中确定第一光线亮度表征信息;
第一调整子单元,被配置为依据第一光线亮度表征信息,对待处理数据进行亮度调整处理,得到第一调整结果;
第二确定子单元,被配置为根据第一调整结果确定调整后数据。
在一种可能的实施方式下,数据生成装置300还包括:
位置确定单元,被配置为对待处理数据进行对象位置确定处理,得到对象位置表征数据;
第二确定子单元,具体被配置为:根据对象位置表征数据和第一调整结果,确定调整后数据,以使调整后数据中记录有目标对象对应的调整后亮度。
在一种可能的实施方式下,调整后数据属于虚拟现实VR图像数据;待处理数据为利用原始VR图像数据转化所得的图像数据;
第二确定子单元,具体被配置为:将第一调整结果转化为第一VR图像数据;将对象位置表征数据转化为第二VR图像数据;根据第一VR图像数据和第二VR图像数据,确定调整后数据。
在一种可能的实施方式下,数据生成装置300还包括:
位置确定单元,被配置为对待处理数据进行对象位置确定处理,得到对象位置表征数据;
第一确定子单元,具体被配置为:根据对象位置表征数据、法向估计结果、待处理数据以及光照贴图,确定调整后数据。
在一种可能的实施方式下,第一确定子单元,包括:
第一提取子单元,被配置为按照对象位置表征数据,从待处理数据中提取对象表征数据;
第二提取子单元,被配置为按照对象位置表征数据,从法向估计结果中提取对象法向结果;
第三确定子单元,被配置为根据对象法向结果、对象表征数据以及光照贴图,确定调整后数据。
在一种可能的实施方式下,第三确定子单元,具体被配置为:依据对象法向结果,从光照贴图中确定第二光线亮度表征信息;依据第二光线亮度表征信息,对对象表征数据进行亮度调整处理,得到第二调整结果;根据第二调整结果确定调整后数据。
在一种可能的实施方式下,待处理数据、对象位置表征数据、以及法向估计结果均为图像序列;
数据生成装置300还包括:
第一平滑单元,被配置为利用预设平滑算法,对对象位置表征数据中相邻帧进行平滑处理;
和/或,
第二平滑单元,被配置为利用预设平滑算法,对法向估计结果中相邻帧进行平滑处理。
在一种可能的实施方式下,待处理数据为利用原始VR图像数据转化所得的数据;
信息获取单元301,包括:
对象检测子单元,被配置为对原始VR图像数据进行对象检测处理,得到对象检测结果;
区域截取子单元,被配置为按照对象检测结果,对原始VR图像数据进行对象区域截取处理,得到区域表征数据;
投影处理子单元,被配置为对区域表征数据进行投影去畸变处理,得到待处理数据。
在一种可能的实施方式下,目标背景表征数据、调整后数据、以及目标数据均属于VR图像数据;
数据生成单元304,具体被配置为:将目标背景表征数据与调整后数据
进行融合处理,得到目标数据。
在一种可能的实施方式下,待处理数据为利用原始VR图像数据转化所得的图像数据;原始VR图像数据为单目全景图像数据、双目全景图像数据、单目半全景图像数据或者双目半全景图像数据。
在一种可能的实施方式下,目标环境贴图为高动态范围成像HDR环境贴图。
在一种可能的实施方式下,信息获取单元301,具体被配置为:接收目标用户针对原始VR图像数据触发的背景更换请求;背景更换请求携带有目标环境贴图和原始VR图像数据;对原始VR图像数据进行转化处理,得到待处理数据。
基于上述数据生成装置300的相关内容可知,对于包括至少一个图像数据的待处理数据(例如,一个VR图像数据或者一个VR视频数据)来说,在获取到目标环境贴图(例如,HDR环境贴图)之后,先利用该目标环境贴图,生成该待处理数据对应的目标背景表征数据,以使该目标背景表征数据所描述的目标背景符合该目标环境贴图中的景物呈现状态(例如,景物姿态、景物所处的亮度状态等),从而使得该目标背景表征数据所描述的目标背景不同于该待处理数据中所出现的原始背景,并依据该目标环境贴图,对该待处理数据中目标对象进行亮度调整处理,得到调整后数据,以使该调整后数据中目标对象所处的亮度状态符合该目标环境贴图中所记录的光线分布规律,从而使得该调整后数据中目标对象所处的亮度状态不同于该待处理数据中目标对象所处的亮度状态;再根据该目标背景表征数据与该调整后数据,生成目标数据,以使该目标数据中目标对象所处的亮度状态与该调整后数据中目标对象所处的亮度状态保持一致。其中,因该目标数据中目标对象所处的亮度状态以及该目标数据中背景所处的亮度状态均符合该目标环境贴图中所记录的光线分布规律,使得该目标数据中目标对象所处的亮度状态与该目标数据中背景所处的亮度状态之间比较协调(例如,亮度过渡比较自然),从而使得该目标数据能够更自然地描述出该目标对象在该目标背景下所处状态(例如,光线亮度状态),如此能够实现在确保对象与背景光线相协调的前提下进行背景更换的目的,以克服因VR图像数据无法满足一些应用场景的背景需求而导致的不良影响。
另外,本公开实施例还提供了一种电子设备,该电子设备包括处理器以及存储器:存储器,被配置为存储指令或计算机程序;处理器,被配置为执行存储器中的指令或计算机程序,以使得电子设备执行本公开实施例提供的数据生成方法的任一实施方式。
参见图4,其示出了适于用来实现本公开实施例的电子设备400的结构示意图。本公开实施例中的终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、PDA(个人数字助理)、PAD(平板电脑)、PMP(便携式多媒体播放器)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字TV、台式计算机等等的固定终端。图4示出的电子设备仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图4所示,电子设备400可以包括处理装置(例如中央处理器、图形处理器等)401,其可以根据存储在只读存储器(ROM)402中的程序或者从存储装置408加载到随机访问存储器(RAM)403中的程序而执行各种适当的动作和处理。在RAM403中,还存储有电子设备400操作所需的各种程序和数据。处理装置401、ROM 402以及RAM 403通过总线404彼此相连。输入/输出(I/O)接口405也连接至总线404。
通常,以下装置可以连接至I/O接口405:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置406;包括例如液晶显示器(LCD)、扬声器、振动器等的输出装置407;包括例如磁带、硬盘等的存储装置408;以及通信装置409。通信装置409可以允许电子设备400与其他设备进行无线或有线通信以交换数据。虽然图4示出了具有各种装置的电子设备400,但是应理解的是,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算机程序产品,其包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置409从网络上被下载和安装,或者从存储装置408被安装,或者从ROM402被安装。在该计算机程序被处理装置401执行时,执行本公开实施例的方法中限定的上述功能。
本公开实施例提供的电子设备与上述实施例提供的方法属于同一发明构思,未在本实施例中详尽描述的技术细节可参见上述实施例,并且本实施例与上述实施例具有相同的有益效果。
本公开实施例还提供了一种计算机可读介质,计算机可读介质中存储有指令或计算机程序,当指令或计算机程序在设备上运行时,使得设备执行本公开实施例提供的数据生成方法的任一实施方式。
需要说明的是,本公开上述的计算机可读介质可以是计算机可读信号介质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本公开中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本公开中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算机可读存储介质以外的任何计算机可读介质,该计算机可读信号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、RF(射频)等等,或者上述的任意合适的组合。
在一些实施方式中,客户端、服务器可以利用诸如HTTP(Hyper Text Transfer Protocol,超文本传输协议)之类的任何当前已知或未来研发的网络协议进行通信,并且可以与任意形式或介质的数字数据通信(例如,通信网络)互连。通信网络的示例包括局域网(“LAN”),广域网(“WAN”),网际网(例如,互联网)以及端对端网络(例如,ad hoc端对端网络),以及任何当前已知或未来研发的网络。
上述计算机可读介质可以是上述电子设备中所包含的;也可以是单独存在,而未装配入该电子设备中。
上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备可以执行上述方法。
可以以一种或多种程序设计语言或其组合来编写用于执行本公开的操作的计算机程序代码,上述程序设计语言包括但不限于面向对象的程序设计语言—诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括局域网(LAN)或广域网(WAN)—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本公开各种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本公开实施例中所涉及到的单元可以通过软件的方式实现,也可以通过硬件的方式来实现。其中,单元/模块的名称在某种情况下并不构成对该单元本身的限定。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(FPGA)、专用集成电路(ASIC)、专用标准产品(ASSP)、片上系统(SOC)、复杂可编程逻辑设备(CPLD)等等。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的更具体示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦除可编程只读存储器(EPROM或快闪存储器)、光纤、便捷式紧凑盘只读存储器(CD-ROM)、光学储存设备、磁储存设备、或上述内容的任何合适组合。
需要说明的是,本说明书中各个实施例采用递进的方式描述,每个实施例重点说明的都是与其他实施例的不同之处,各个实施例之间相同相似部分互相参见即可。对于实施例公开的系统或装置而言,由于其与实施例公开的方法相对应,所以描述的比较简单,相关之处参见方法部分说明即可。
应当理解,在本公开中,“至少一个(项)”是指一个或者多个,“多个”是指两个或两个以上。“和/或”,用于描述关联对象的关联关系,表示可以存在三种关系,例如,“A和/或B”可以表示:只存在A,只存在B以及同时存在A和B三种情况,其中A,B可以是单数或者复数。字符“/”一般表示前后关联对象是一种“或”的关系。“以下至少一项(个)”或其类似表达,是指这些项中的任意组合,包括单项(个)或复数项(个)的任意组合。例如,a,b或c中的至少一项(个),可以表示:a,b,c,“a和b”,“a和c”,“b和c”,或“a和b和c”,其中a,b,c可以是单个,也可以是多个。
还需要说明的是,在本文中,诸如第一和第二等之类的关系术语仅仅用来将一个实体或者操作与另一个实体或操作区分开来,而不一定要求或者暗示这些实体或操作之间存在任何这种实际的关系或者顺序。而且,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、物品或者设备不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、物品或者设备所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括所述要素的过程、方法、物品或者设备中还存在另外的相同要素。
结合本文中所公开的实施例描述的方法或算法的步骤可以直接用硬件、处理器执行的软件模块,或者二者的结合来实施。软件模块可以置于随机存储器(RAM)、内存、只读存储器(ROM)、电可编程ROM、电可擦除可编程ROM、寄存器、硬盘、可移动磁盘、CD-ROM、或技术领域内所公知的任意其它形式的存储介质中。
对所公开的实施例的上述说明,使本领域专业技术人员能够实现或使用本公开。对这些实施例的多种修改对本领域的专业技术人员来说将是显而易见的,本文中所定义的一般原理可以在不脱离本公开的精神或范围的情况下,在其它实施例中实现。因此,本公开将不会被限制于本文所示的这些实施例,而是要符合与本文所公开的原理和新颖特点相一致的最宽的范围。
Claims (17)
- 一种数据生成方法,包括:获取目标环境贴图和待处理数据,其中,所述待处理数据包括至少一个图像数据,所述待处理数据用于描述目标对象在原始背景下所处状态;利用所述目标环境贴图,生成所述待处理数据对应的目标背景表征数据,其中,所述目标背景表征数据所描述的目标背景不同于所述原始背景;依据所述目标环境贴图,对所述待处理数据中目标对象进行亮度调整处理,得到调整后数据,其中,所述调整后数据中目标对象所处的亮度状态不同于所述待处理数据中目标对象所处的亮度状态;根据所述目标背景表征数据与所述调整后数据,生成目标数据,其中,所述目标数据用于描述所述目标对象在所述目标背景下所处状态,所述目标数据中目标对象所处的亮度状态与所述调整后数据中目标对象所处的亮度状态保持一致。
- 根据权利要求1所述的方法,其中,所述依据所述目标环境贴图,对所述待处理数据中目标对象进行亮度调整处理,得到调整后数据,包括:将所述目标环境贴图中记录的光线信息转化为光照贴图;对所述待处理数据进行法向估计处理,得到法向估计结果;根据所述法向估计结果、所述待处理数据以及所述光照贴图,确定所述调整后数据。
- 根据权利要求2所述的方法,其中,所述根据所述法向估计结果、所述待处理数据以及所述光照贴图,确定所述调整后数据,包括:依据所述法向估计结果,从所述光照贴图中确定第一光线亮度表征信息;依据所述第一光线亮度表征信息,对所述待处理数据进行亮度调整处理,得到第一调整结果;根据所述第一调整结果确定所述调整后数据。
- 根据权利要求3所述的方法,其中,所述根据所述第一调整结果确定所述调整后数据之前,所述方法还包括:对所述待处理数据进行对象位置确定处理,得到对象位置表征数据;所述根据所述第一调整结果确定所述调整后数据,包括:根据所述对象位置表征数据和所述第一调整结果,确定所述调整后数据,以使所述调整后数据中记录有所述目标对象对应的调整后亮度。
- 根据权利要求4所述的方法,其中,所述调整后数据属于虚拟现实VR图像数据;所述待处理数据为利用原始VR图像数据转化所得的图像数据;所述根据所述对象位置表征数据和所述第一调整结果,确定所述调整后数据,包括:将所述第一调整结果转化为第一VR图像数据;将所述对象位置表征数据转化为第二VR图像数据;根据所述第一VR图像数据和所述第二VR图像数据,确定所述调整后数据。
- 根据权利要求2所述的方法,其中,所述根据所述法向估计结果、所述待处理数据以及所述光照贴图,确定所述调整后数据之前,所述方法还包括:对所述待处理数据进行对象位置确定处理,得到对象位置表征数据;所述根据所述法向估计结果、所述待处理数据以及所述光照贴图,确定所述调整后数据,包括:根据所述对象位置表征数据、所述法向估计结果、所述待处理数据以及所述光照贴图,确定所述调整后数据。
- 根据权利要求6所述的方法,其中,所述根据所述对象位置表征数据、所述法向估计结果、所述待处理数据以及所述光照贴图,确定所述调整后数据,包括:按照所述对象位置表征数据,从所述待处理数据中提取对象表征数据;按照所述对象位置表征数据,从所述法向估计结果中提取对象法向结果;根据所述对象法向结果、所述对象表征数据以及所述光照贴图,确定所述调整后数据。
- 根据权利要求7所述的方法,其中,所述根据所述对象法向结果、所述对象表征数据以及所述光照贴图,确定所述调整后数据,包括:依据所述对象法向结果,从所述光照贴图中确定第二光线亮度表征信息;依据所述第二光线亮度表征信息,对所述对象表征数据进行亮度调整处理,得到第二调整结果;根据所述第二调整结果确定所述调整后数据。
- 根据权利要求4或者6所述的方法,其中,所述待处理数据、所述对象位置表征数据、以及所述法向估计结果均为图像序列;所述方法还包括:利用预设平滑算法,对所述对象位置表征数据中相邻帧进行平滑处理;和/或,利用预设平滑算法,对所述法向估计结果中相邻帧进行平滑处理。
- 根据权利要求1所述的方法,其中,所述待处理数据为利用原始VR图像数据转化所得的数据;所述待处理数据的获取过程,包括:对所述原始VR图像数据进行对象检测处理,得到对象检测结果;按照所述对象检测结果,对所述原始VR图像数据进行对象区域截取处理,得到区域表征数据;对所述区域表征数据进行投影去畸变处理,得到所述待处理数据。
- 根据权利要求1所述的方法,其中,所述目标背景表征数据、所述调整后数据、以及所述目标数据均属于VR图像数据;所述根据所述目标背景表征数据与所述调整后数据,生成目标数据,包括:将所述目标背景表征数据与所述调整后数据进行融合处理,得到所述目标数据。
- 根据权利要求1所述的方法,其中,所述待处理数据为利用原始VR图像数据转化所得的图像数据;所述原始VR图像数据为单目全景图像数据、双目全景图像数据、单目半全景图像数据或者双目半全景图像数据。
- 根据权利要求1所述的方法,其中,所述目标环境贴图为高动态范围成像HDR环境贴图。
- 根据权利要求1所述的方法,其中,所述获取目标环境贴图和待处理数据,包括:接收目标用户针对原始VR图像数据触发的背景更换请求,其中,所述背景更换请求携带有所述目标环境贴图和所述原始VR图像数据;对所述原始VR图像数据进行转化处理,得到所述待处理数据。
- 一种数据生成装置,包括:信息获取单元,被配置为获取目标环境贴图和待处理数据,其中,所述待处理数据包括至少一个图像数据,所述待处理数据用于描述目标对象在原始背景下所处状态;背景生成单元,被配置为利用所述目标环境贴图,生成所述待处理数据对应的目标背景表征数据,其中,所述目标背景表征数据所描述的目标背景不同于所述原始背景;对象重打光单元,被配置为依据所述目标环境贴图,对所述待处理数据中目标对象进行亮度调整处理,得到调整后数据,其中,所述调整后数据中目标对象所处的亮度状态不同于所述待处理数据中目标对象所处的亮度状态;以及数据生成单元,被配置为根据所述目标背景表征数据与所述调整后数据,生成目标数据,其中,所述目标数据用于描述所述目标对象在所述目标背景下所处状态;所述目标数据中目标对象所处的亮度状态与所述调整后数据中目标对象所处的亮度状态保持一致。
- 一种电子设备,包括:存储器,被配置为存储指令或计算机程序;以及处理器,被配置为执行所述存储器中的所述指令或计算机程序,以使得所述电子设备执行权利要求1-14任一项所述的方法。
- 一种计算机可读介质,其中,所述计算机可读介质中存储有指令或计算机程序,当所述指令或计算机程序在设备上运行时,使得所述设备执行权利要求1-14任一项所述的方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310436249.6 | 2023-04-21 | ||
| CN202310436249.6A CN118823176A (zh) | 2023-04-21 | 2023-04-21 | 一种数据生成方法、装置、电子设备、计算机可读介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024217237A1 true WO2024217237A1 (zh) | 2024-10-24 |
Family
ID=93080635
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/084022 Ceased WO2024217237A1 (zh) | 2023-04-21 | 2024-03-27 | 数据生成方法、装置、电子设备、计算机可读介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN118823176A (zh) |
| WO (1) | WO2024217237A1 (zh) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109887066A (zh) * | 2019-02-25 | 2019-06-14 | 网易(杭州)网络有限公司 | 光照效果处理方法及装置、电子设备、存储介质 |
| CN110113534A (zh) * | 2019-05-13 | 2019-08-09 | Oppo广东移动通信有限公司 | 一种图像处理方法、图像处理装置及移动终端 |
| CN115082639A (zh) * | 2022-06-15 | 2022-09-20 | 北京百度网讯科技有限公司 | 图像生成方法、装置、电子设备和存储介质 |
-
2023
- 2023-04-21 CN CN202310436249.6A patent/CN118823176A/zh active Pending
-
2024
- 2024-03-27 WO PCT/CN2024/084022 patent/WO2024217237A1/zh not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN109887066A (zh) * | 2019-02-25 | 2019-06-14 | 网易(杭州)网络有限公司 | 光照效果处理方法及装置、电子设备、存储介质 |
| CN110113534A (zh) * | 2019-05-13 | 2019-08-09 | Oppo广东移动通信有限公司 | 一种图像处理方法、图像处理装置及移动终端 |
| CN115082639A (zh) * | 2022-06-15 | 2022-09-20 | 北京百度网讯科技有限公司 | 图像生成方法、装置、电子设备和存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN118823176A (zh) | 2024-10-22 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20240046557A1 (en) | Method, device, and non-transitory computer-readable storage medium for reconstructing a three-dimensional model | |
| CN115690382B (zh) | 深度学习模型的训练方法、生成全景图的方法和装置 | |
| CN111311756A (zh) | 增强现实ar显示方法及相关装置 | |
| CN110874818A (zh) | 图像处理和虚拟空间构建方法、装置、系统和存储介质 | |
| CN114742703B (zh) | 双目立体全景图像的生成方法、装置、设备和存储介质 | |
| US11836887B2 (en) | Video generation method and apparatus, and readable medium and electronic device | |
| WO2023169281A1 (zh) | 图像配准方法、装置、存储介质及电子设备 | |
| CN115002442B (zh) | 一种图像展示方法、装置、电子设备及存储介质 | |
| WO2022037484A1 (zh) | 图像处理方法、装置、设备及存储介质 | |
| CN111325792A (zh) | 用于确定相机位姿的方法、装置、设备和介质 | |
| See et al. | Virtual reality 360 interactive panorama reproduction obstacles and issues | |
| WO2025092175A1 (zh) | 虚拟对象生成方法、装置、计算机设备及存储介质 | |
| CN117689804A (zh) | 一种三维重建方法、装置、设备和存储介质 | |
| JP2023550970A (ja) | 画面の中の背景を変更する方法、機器、記憶媒体、及びプログラム製品 | |
| CN114241127A (zh) | 全景图像生成方法、装置、电子设备和介质 | |
| CN110111241A (zh) | 用于生成动态图像的方法和装置 | |
| WO2024055837A1 (zh) | 一种图像处理方法、装置、设备及介质 | |
| CN114202617A (zh) | 视频图像处理方法、装置、电子设备及存储介质 | |
| WO2024217237A1 (zh) | 数据生成方法、装置、电子设备、计算机可读介质 | |
| CN113537194A (zh) | 光照估计方法、光照估计装置、存储介质与电子设备 | |
| CN111818265A (zh) | 基于增强现实模型的交互方法、装置、电子设备及介质 | |
| CN117152393A (zh) | 一种增强现实的呈现方法、系统、装置、设备及介质 | |
| KR20200114348A (ko) | 증강현실의 공간맵을 이용한 콘텐츠 공유 장치 및 그 방법 | |
| KR102534449B1 (ko) | 이미지 처리 방법, 장치, 전자 장치 및 컴퓨터 판독 가능 저장 매체 | |
| WO2023216822A1 (zh) | 图像校正方法、装置、电子设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24791823 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 24791823 Country of ref document: EP Kind code of ref document: A1 |