WO2024124653A1 - 三维场景重建、在线家装与商品获取方法、设备及介质 - Google Patents

三维场景重建、在线家装与商品获取方法、设备及介质 Download PDF

Info

Publication number
WO2024124653A1
WO2024124653A1 PCT/CN2023/072122 CN2023072122W WO2024124653A1 WO 2024124653 A1 WO2024124653 A1 WO 2024124653A1 CN 2023072122 W CN2023072122 W CN 2023072122W WO 2024124653 A1 WO2024124653 A1 WO 2024124653A1
Authority
WO
WIPO (PCT)
Prior art keywords
target
model
camera
target image
image
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2023/072122
Other languages
English (en)
French (fr)
Inventor
宋瑾
马林
费义云
胡晓航
蒋健安
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Alibaba China Co Ltd
Original Assignee
Alibaba China Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Alibaba China Co Ltd filed Critical Alibaba China Co Ltd
Publication of WO2024124653A1 publication Critical patent/WO2024124653A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T17/00Three-dimensional [3D] modelling for computer graphics
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T19/00Manipulating three-dimensional [3D] models or images for computer graphics
    • G06T19/20Editing of three-dimensional [3D] images, e.g. changing shapes or colours, aligning objects or positioning parts
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06TIMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T3/00Geometric image transformations in the plane of the image
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02TCLIMATE CHANGE MITIGATION TECHNOLOGIES RELATED TO TRANSPORTATION
    • Y02T10/00Road transport of goods or passengers
    • Y02T10/10Internal combustion engine [ICE] based vehicles
    • Y02T10/40Engine management systems

Definitions

  • the present application relates to the field of three-dimensional reconstruction technology, and in particular to a three-dimensional scene reconstruction, online home improvement and product acquisition method, device and medium.
  • APPs home decoration applications
  • the existing 3D model room is a three-dimensional house model made by using 3D mapping technology, which lacks the user's real house scene information. Facing the home decoration needs that need to be matched with the real house environment, the home decoration matching effect is not very ideal. Therefore, a solution that can reflect the real house scene information in the house three-dimensional model is urgently needed to improve the matching effect of online home decoration.
  • Multiple aspects of the present application provide a three-dimensional scene reconstruction, online home improvement and product acquisition method, device and medium, which are used to reflect real three-dimensional scene information in a three-dimensional scene model and improve the application effect based on the three-dimensional scene model, such as home improvement matching effect.
  • An embodiment of the present application provides a three-dimensional scene reconstruction method, including: acquiring a target image corresponding to a first spatial object, where the first spatial object is at least a part of a spatial object in a target physical space; detecting multiple boundary lines and vanishing points in at least two orthogonal main directions in the target image, where the boundary lines are intersection lines between adjacent physical main structures in the first spatial object; determining, based on the vanishing points in the at least two orthogonal main directions, intrinsic camera parameters of a target camera and a gravity direction in a camera coordinate system, where the target camera refers to a camera used to capture the target image; reconstructing an initial three-dimensional model corresponding to the first spatial object based on the multiple boundary lines and the gravity direction in the camera coordinate system, where the initial three-dimensional model includes adjacent model main structures corresponding to the adjacent physical main structures; optimizing the initial three-dimensional model based on the constraint relationship between the multiple boundary lines, the constraint relationship between the vanishing points in the at least two orthogonal main directions, and the
  • the embodiment of the present application also provides an online home improvement method, including: responding to an image upload operation, obtaining a first space A target image corresponding to a spatial object in a target space, wherein the first spatial object is at least a part of a spatial object in a target physical space; in response to a placement operation of a target home decoration object on the target image, the target home decoration object is integrated into a target three-dimensional model corresponding to the first spatial object to obtain a target three-dimensional model integrating the target home decoration object; the target three-dimensional model integrating the target home decoration object is projected onto the target image to obtain a home decoration rendering containing the target home decoration object; wherein the target three-dimensional model is constructed according to the steps in the three-dimensional scene reconstruction method provided in an embodiment of the present application.
  • An embodiment of the present application also provides a product selection method, including: responding to a selection operation on a product page, determining a selected target product, the target product having a product three-dimensional model; responding to a matching effect viewing operation, selecting a target image corresponding to a first spatial object to be matched with the target product; adding the product three-dimensional model to the target three-dimensional model corresponding to the first spatial object to obtain a target three-dimensional model that integrates the target product; projecting the target three-dimensional model that integrates the target product onto the target image to obtain a matching effect diagram of the target product and the first spatial object; wherein the target three-dimensional model is constructed according to the steps in the three-dimensional scene reconstruction method provided in an embodiment of the present application.
  • An embodiment of the present application also provides an electronic device, comprising: a memory and a processor; the memory is used to store a computer program; the processor is coupled to the memory and is used to execute the computer program to execute the steps in the three-dimensional scene reconstruction method, online home improvement method or product selection method provided in the embodiment of the present application.
  • An embodiment of the present application also provides a computer-readable storage medium storing a computer program.
  • the processor When the computer program is executed by a processor, the processor is enabled to implement the steps in the three-dimensional scene reconstruction method, online home improvement method or product selection method provided in the embodiment of the present application.
  • the 3D structural information corresponding to the first space is obtained, that is, the target three-dimensional model.
  • the real scene information corresponding to the first spatial object can be reflected in the target three-dimensional model, which is conducive to improving the application effect based on the three-dimensional model. For example, in a home decoration scene, the home decoration matching effect based on the three-dimensional model can be improved.
  • 2D visual information such as boundary lines and vanishing points are directly detected from the image, and then 3D structural information is generated using these 2D visual information and related constraint relationships, rather than directly regressing the 3D structural information from the image. Therefore, the three-dimensional model generated in the embodiment of the present application is more stable and more interpretable, and is more robust to situations where the shooting angle is large and the scene is cluttered. At the same time, combined with the constraint relationship between the boundary lines and the vanishing points, the main structures of the model in the generated three-dimensional model have a higher degree of fit.
  • FIG. 1a is a schematic structural diagram of a three-dimensional scene reconstruction system provided by an exemplary embodiment of the present application
  • FIG1b is a schematic flow chart of a three-dimensional scene reconstruction method provided by an exemplary embodiment of the present application.
  • FIG2a is a schematic diagram of a boundary line in a first space object provided by an exemplary embodiment of the present application.
  • FIG2b is a schematic structural diagram of a vanishing point provided by an exemplary embodiment of the present application.
  • FIG3a is a schematic structural diagram of a boundary line detection model based on Hough transform provided by an exemplary embodiment of the present application
  • FIG3 b is a schematic structural diagram of a vanishing point detection model based on Hough transform provided by an exemplary embodiment of the present application;
  • FIG3c is a schematic diagram of the intersection position between boundary lines provided by an exemplary embodiment of the present application.
  • FIG3d is a schematic structural diagram of a reference model main structure provided by an exemplary embodiment of the present application.
  • FIG3e is a schematic structural diagram of another model main structure provided by an exemplary embodiment of the present application.
  • FIG4a is a schematic flow chart of an online home improvement method provided by an exemplary embodiment of the present application.
  • FIG4b is a schematic diagram of a process of three-dimensional scene reconstruction provided by an exemplary embodiment of the present application.
  • FIG4c is a schematic diagram of an online home improvement effect based on a target three-dimensional model provided by an exemplary embodiment of the present application.
  • FIG4d is another schematic diagram of online home improvement effects based on a target three-dimensional model provided by an exemplary embodiment of the present application.
  • FIG4e is a schematic diagram of a flow chart of a commodity selection method provided by an exemplary embodiment of the present application.
  • FIG5 is a schematic structural diagram of a three-dimensional scene reconstruction device provided by an exemplary embodiment of the present application.
  • FIG. 6 is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present application.
  • the embodiment of the present application provides a three-dimensional scene reconstruction method.
  • the method based on the image corresponding to the first spatial object, by detecting 2D visual information such as intersection lines and vanishing points in the image, based on these 2D visual information and combined with the constraint relationship between the intersection lines and the vanishing points, the 3D structural information corresponding to the first space is obtained, that is, the target three-dimensional model.
  • the real scene information corresponding to the first spatial object can be reflected in the target three-dimensional model, which is conducive to improving the application effect based on the three-dimensional model.
  • the home improvement matching effect based on the three-dimensional model can be improved.
  • the shopping experience can be improved and the probability of return and exchange can be reduced.
  • Fig. 1a is a schematic diagram of a structure of a 3D scene reconstruction system provided by an exemplary embodiment of the present application. As shown in Fig. 1a, the system includes: a terminal device 10 and a server device 20. The terminal device 10 and the server device 20 are in communication connection.
  • the terminal device 10 may be a mobile phone, a laptop computer, or a desktop computer, etc.
  • the server device 20 It can be a physical server, a cloud server or a server array, etc.
  • the terminal device 101 is a smart phone and the server device 20 is a physical server as an example, but it is not limited to this.
  • the terminal device 10 can obtain the target image corresponding to the first spatial object.
  • the terminal device has its own camera, and the terminal device can capture the target image in the first spatial object through its own camera; for another example, the target image corresponding to the first spatial object is captured by a camera independent of the terminal device 10, and the camera independent of the terminal device provides the captured target image to the terminal device 10.
  • the terminal device 10 can provide the acquired target image corresponding to the first spatial object to the server device 20, and the server device 20 generates a target three-dimensional model corresponding to the first spatial object.
  • the process of the server device 20 generating the target three-dimensional model corresponding to the first spatial object includes: detecting multiple boundary lines and at least two vanishing points in the orthogonal main directions in the target image, the boundary line is the intersection line between adjacent physical main structures in the first spatial object; according to the vanishing points in at least two orthogonal main directions, determining the camera internal parameters of the target camera and the gravity direction in the camera coordinate system, the target camera refers to the camera used to shoot the target image; according to the multiple boundary lines and the gravity direction in the camera coordinate system, reconstructing the initial three-dimensional model corresponding to the first spatial object, the initial three-dimensional model includes adjacent model main structures corresponding to adjacent physical main structures; according to the constraint relationship between the multiple boundary lines, the constraint relationship between the vanishing points in at least two orthogonal main directions and the camera internal parameters,
  • the user can also add target objects required for home improvement to the target image on the terminal device 10, such as sofas, paintings, office chairs or desks;
  • the terminal device 10 can obtain the position range information of the target object in the target image, and provide the target object information and the position range information of the target object in the target image to the server device 20;
  • the server device 20 fuses the three-dimensional model of the target object into the target three-dimensional model corresponding to the first spatial object according to the position range information of the target object in the target image, and projects the target three-dimensional model after the fusion of the target object onto the target image to obtain a target image containing the target object, and returns the target image containing the target object to the terminal device 10, which is displayed to the user by the terminal device 10 to achieve the display of online home improvement matching effects.
  • the 3D scene reconstruction method provided in the embodiment of the present application can not only be applied to the 3D scene reconstruction system shown in FIG1a, which is completed by the cooperation between the terminal device and the server device, but can also be implemented independently by the terminal device.
  • the action originally performed by the server device can be performed by the terminal device in this embodiment, and the other contents are the same or similar to those in the system shown in FIG1a, and will not be repeated here.
  • the process of 3D scene reconstruction please refer to the description in the following method embodiment.
  • FIG1b is a flow chart of a three-dimensional scene reconstruction method provided by an exemplary embodiment of the present application. As shown in FIG1b , the method includes:
  • the initial three-dimensional model is optimized to obtain the target three-dimensional model.
  • the target physical space refers to a spatial area in a specific scene.
  • the target physical space can be a shopping mall, a supermarket, an airport, or a house, etc., and various areas with spatial concepts.
  • the target physical space contains at least one spatial object.
  • at least one spatial object constitutes the target physical space.
  • the target physical space can be a physical house, and the physical house includes multiple spatial objects, such as a kitchen, a bedroom, a living room, and a bathroom.
  • the target physical space can be a shopping mall, which includes multiple floors, and different floors have different merchants, and each merchant can be regarded as a spatial object.
  • the first spatial object can be a spatial object in the target physical space, or it can be multiple spatial objects in the target physical space, such as 2 or 3; for another example, the first spatial object can be a local space in a certain spatial object.
  • the first spatial object can be the entire physical house, or an independent spatial object such as the master bedroom, second bedroom or living room therein, or a local space in an independent spatial object such as the master bedroom, second bedroom or living room, or a local space that includes multiple independent spatial objects at the same time, such as including part of the living room space and part of the balcony space at the same time.
  • the target image corresponding to the first spatial object can be obtained, for example, by capturing an image in the first spatial object through a camera, and obtaining the target image corresponding to the first spatial object from the camera, or by capturing the target image corresponding to the first spatial object through the camera on the terminal device, where the target image is an environmental image of the first spatial object and contains real environmental information of the first spatial object.
  • the boundary lines are the intersection lines between adjacent physical main structures in the first spatial object, wherein the physical main structures may include but are not limited to: walls, ceilings or floors, etc.
  • the intersection lines between walls can be called wall lines
  • the intersection lines between walls and floors can be called ground lines
  • the intersection lines between walls and ceilings can be called ceiling lines.
  • the first spatial object includes: ceiling, floor and wall
  • the extracted ground line is C1
  • the extracted wall line is C2
  • the extracted ceiling line is C3.
  • the ground line C1, the wall line C2 and the ceiling line C3 in FIG2a are all examples of boundary lines in the embodiments of the present application, but are not limited to this.
  • the first spatial object may conform to a Manhattan structure, wherein the Manhattan structure includes three mutually orthogonal main directions, for example, the three orthogonal main directions are the direction of gravity, the direction facing the main wall, and the direction facing the side wall.
  • the direction of gravity can be considered as a direction perpendicular to the ground or the ceiling;
  • the main wall refers to the wall with the smallest angle with the optical axis of the target camera, the target camera is the camera used to capture the target image, the optical axis of the target camera refers to the line passing through the center of the lens, the direction facing the main wall is the direction perpendicular to the main wall, that is, the direction of the normal vector of the main wall;
  • the side wall is relative to the main wall, and mainly refers to the wall adjacent to the main wall.
  • the direction facing the side wall refers to the direction perpendicular to the side wall.
  • there can be multiple side walls in the target image for example, 2, multiple side walls are relative or parallel, and the directions corresponding to the multiple side walls are the same, that is, the normal vector direction of the side wall.
  • the target image may contain physical main structures corresponding to three orthogonal main directions at the same time, for example, it contains the main wall, the side wall, and the ground or ceiling at the same time; it may also contain only physical main structures corresponding to two orthogonal main directions, for example, it contains only the main wall and the ground, or it contains only the side wall and the ceiling, or it contains only the main wall, the ceiling and the ground.
  • the target image contains physical main structures corresponding to three orthogonal main directions or physical main structures corresponding to two orthogonal main directions, it depends on the shooting angle of the target image. In short, there are at least physical main structures corresponding to two orthogonal main directions in the target image.
  • parallel straight lines in a three-dimensional space intersect at a point in a two-dimensional image corresponding to the three-dimensional space after perspective change, and the point is the vanishing point (VP), and the direction of the vanishing point represents the direction of the parallel straight lines in the three-dimensional space that intersect to obtain the vanishing point.
  • VP vanishing point
  • vanishing points in the target image including vanishing points located in at least two orthogonal main directions; for the vanishing points located in a certain orthogonal main direction, the parallel straight lines in the three-dimensional space that intersect to obtain the vanishing point can be perpendicular to the physical main structure corresponding to the orthogonal main direction in the three-dimensional space, and the physical main structure in the three-dimensional space conforms to the Manhattan structure, and each physical main structure corresponds to a main direction, that is, the direction of the vanishing point located in a certain orthogonal main direction in the target image corresponds to a main direction in the three-dimensional space, that is, the direction of the vanishing point located in a certain orthogonal main direction represents the direction of the normal vector of a certain physical main structure.
  • line 1 is parallel to the right wall (i.e., the wall where the painting is hung in Figure 2b), and the direction of the vanishing point D1 obtained by the intersection of line 1 represents the normal vector direction of the left wall (i.e., the wall with the window in Figure 2b);
  • line 2 is parallel to the left wall (i.e., the wall with the window in Figure 2b), and the direction of the vanishing point D2 obtained by the intersection of line 2 represents the normal vector direction of the right wall (i.e., the wall where the painting is hung in Figure 2b);
  • line 3 is parallel to the ground (or ceiling), and the direction of the vanishing point D3 obtained by the intersection of line 3 represents the normal vector direction of the ground (or ceiling), i.e., the direction parallel to the direction of gravity.
  • multiple boundary lines and vanishing points in at least two orthogonal main directions in the target image can be detected.
  • detection is performed directly on the basis of the target image without being disturbed by image segmentation. Therefore, the detected boundary lines and vanishing points have high accuracy and precision.
  • the position of the vanishing point in the target image is determined by the intrinsic parameters of the target camera and the 3D direction corresponding to the vanishing point (i.e., the main direction where the vanishing point is located).
  • the intrinsic parameters of the target camera are parameters related to the characteristics of the target camera itself, such as the focal length, pixel size, and principal point position (i.e., the position of the optical center of the camera) of the target camera.
  • the intrinsic camera parameters of the target camera and the gravity direction in the camera coordinate system can be determined based on the vanishing points in at least two orthogonal main directions.
  • the intrinsic camera parameters of the target camera can be determined based on the vanishing points in at least two orthogonal main directions
  • the gravity direction in the camera coordinate system can be determined based on the intrinsic camera parameters and the vanishing points in the gravity direction.
  • the initial three-dimensional model corresponding to the first spatial object can be reconstructed according to the multiple boundary lines existing in the target image and the gravity direction in the camera coordinate system, and the initial three-dimensional model includes the adjacent model main structure corresponding to the adjacent physical main structure.
  • the physical main structure is the main structure constituting the first spatial object
  • the model main structure is the main structure constituting the initial three-dimensional model
  • the model main structure is the embodiment of the physical main structure in the initial three-dimensional model.
  • the initial three-dimensional model corresponding to the first spatial object is constructed according to the plane equations of each physical main structure, for example, the corresponding three-dimensional model of the house can be constructed according to the plane equations of each wall, ground or ceiling contained in the house.
  • the plane equation of the physical main structure can uniquely represent the physical main structure, and optionally, it can be represented by the normal vector of the physical main structure and the distance to the optical center of the camera, but it is not limited to this.
  • the model main structure in the initial three-dimensional model also has a plane equation.
  • the plane equation of the model main structure should be the same as the plane direction of the corresponding physical main structure.
  • the general principle of reconstructing the initial three-dimensional model corresponding to the first spatial object according to the multiple boundary lines existing in the target image and the gravity direction in the camera coordinate system is as follows: according to the gravity direction in the camera coordinate system, a ground plane in the three-dimensional space is constructed, and the ground plane is perpendicular to the gravity direction in the camera coordinate system; based on the multiple boundary lines existing in the target image, the boundary formed by the first spatial object on the ground plane and the main structure of the model existing on the boundary can be determined; further, the height of the camera optical center and the ground, that is, the camera height, can be assumed, and based on the assumed camera height and the pre-assumed scaling ratio, the height of the main structure of the model (such as the ceiling) can be determined, thereby obtaining the initial three-dimensional model corresponding to the first spatial object.
  • the initial three-dimensional model corresponding to the first spatial object is constructed based on the boundary line in the target image.
  • the perspective relationship in the initial three-dimensional model will be incorrect, such as there is no vertical relationship between adjacent walls. Further, assuming the camera height and scaling ratio, the height of the generated wall will also be inaccurate.
  • the initial three-dimensional model after obtaining the initial three-dimensional model, the initial three-dimensional model will be optimized according to the constraint relationship between multiple boundary lines and the constraint relationship between the vanishing points in at least two orthogonal main directions, combined with the camera internal parameters, to obtain the target three-dimensional model.
  • the constraint relationship between the multiple boundary lines is: the adjacent physical main structures forming the boundary lines are perpendicular to each other.
  • This constraint relationship can also be called Manhattan constraint, that is, the walls, ground and ceiling in the first spatial object are perpendicular to each other.
  • the constraint relationship between the vanishing points in at least two orthogonal main directions is: the physical main structures corresponding to the vanishing points in any two orthogonal main directions are perpendicular to each other, and the normal vector direction of the physical main structure corresponding to the same vanishing point is parallel to the direction of the parallel straight line represented by the vanishing point. Based on these constraint relationships, the normal vector direction and the distance to the camera optical center corresponding to the model main structure in each initial three-dimensional model can be optimized to obtain a model main structure whose position relationship, height, etc. are highly consistent with the physical main structure.
  • the 3D structural information corresponding to the first space is obtained, that is, the target three-dimensional model.
  • the real scene information corresponding to the first spatial object can be reflected in the target three-dimensional model, which is conducive to improving the application effect based on the three-dimensional model, for example In the home decoration scene, the home decoration matching effect based on the three-dimensional model can be improved.
  • the target three-dimensional model generated by this embodiment has higher stability and stronger interpretability, and is more robust for situations where the shooting angle is tilted and the scene is cluttered; at the same time, combined with the constraint relationship between the intersection lines and the vanishing points, the degree of fit between the main structure of the model in the generated target three-dimensional model is higher.
  • multiple boundary lines existing in the target image and vanishing points in at least two orthogonal main directions can be detected based on Hough transform.
  • Hough transform is a feature extraction technology, which is widely used in image analysis, computer vision, digital image processing and other fields.
  • Hough transform can be used to detect multiple boundary lines existing in the target image, such as ground lines, wall lines and ceiling lines.
  • vanishing points in at least two orthogonal main directions contained in the target image can also be detected.
  • a deep neural network model based on Hough transform can be used to detect multiple boundary lines and vanishing points in at least two orthogonal main directions in the target image.
  • the deep neural network model based on Hough transform includes: a boundary line detection model based on Hough transform and a vanishing point detection model based on Hough transform.
  • the boundary line detection model based on Hough transform can be any neural network model that can detect boundary lines from an image
  • the vanishing point detection model based on Hough transform can be any neural network model that can detect vanishing points from an image.
  • the target image can be input into the boundary line detection model based on Hough transform to detect boundary lines, so as to obtain multiple boundary lines in the target image; on the other hand, the target image can be input into the vanishing point detection model based on Hough transform to detect vanishing points, so as to obtain vanishing points in at least two orthogonal main directions in the target image.
  • the architecture and detection process of the neural network model for boundary line detection and vanishing point detection are exemplarily described below.
  • the boundary detection model based on the Hough transform includes: a first feature extraction network, which is a neural network that integrates skip connections and multi-scale features.
  • multi-scale features refer to the technology of extracting multiple feature maps of different scales for the target image.
  • Feature maps of different scales contain different feature information. The smaller the size of the feature map, the greater the depth, which belongs to the deep network feature; conversely, the larger the size of the feature map, the smaller the depth, which belongs to the shallow network feature.
  • the receptive field (Receptive Field) of the deep network feature is relatively large, the resolution of the feature map is low, the semantic information representation ability is strong, and the representation ability of geometric detail information is weak; the receptive field of the shallow network feature is relatively small, the resolution of the feature map is high, the semantic information representation ability is weak, and the geometric detail information representation ability is strong.
  • the receptive field is the size of the area mapped on the input image by the pixel points on the feature map (feature map) output by each layer of the convolutional neural network.
  • the skip connection is a simple and effective operation to integrate the deep network features and the shallow network features. The skip connection can add the shallow network features to the deep network features of the same scale at the element level.
  • the target image can be input into the boundary detection model based on Hough transform, and the first feature extraction network integrating jump connection and multi-scale features can be used to extract features of the target image to obtain feature maps of multiple scales.
  • the feature maps of multiple scales finally output by the first feature extraction network are referred to as first target feature maps of multiple scales.
  • the first target feature maps of multiple scales belong to the features in the image pixel space.
  • the features in the image pixel space can be converted into features in the Hough space, and the features in the Hough space are classified into two categories, that is, to distinguish whether the features in the Hough space are features corresponding to the boundary lines in the image pixel space.
  • the boundary line detection model based on the Hough transform also includes: a Hough transform network.
  • the first target feature maps of multiple scales can be Hough transformed using the polar coordinate-based Hough transform to map the straight lines in the target image to points in the Hough space, and the converted points are classified into two categories in the Hough space, for example, to distinguish which points in the Hough space are points formed by the boundary lines in the image pixel space and which points are not points formed by the boundary lines in the image pixel space.
  • the first target feature map of any scale it belongs to the image space, that is, it uses the pixel coordinate system, which is a coordinate system with the center of the target image as the origin, the horizontal direction as the x-axis, and the vertical direction as the y-axis, and its unit length is pixel.
  • the Hough space uses the polar coordinate system.
  • a line in the target image can be represented by ( ⁇ , ⁇ ), ⁇ represents the distance from the end of the line to the origin of the pixel coordinate system, ⁇ represents the angle between the line and the x-axis of the pixel coordinate system, and the coordinates of the first target feature map in the Hough space are ( ⁇ , ⁇ ), so a line of the target image represented by polar coordinates will become a point.
  • the straight lines in the target image will be mapped to points in the Hough space, including points mapped by the boundary line, that is, points matching the boundary line features, and points mapped by the non-boundary line, that is, points not matching the boundary line features.
  • the points matching the boundary line features in the Hough space are called target points. Based on this, multiple target points matching the boundary line features can be selected in the Hough space, and the multiple target points can be remapped into the image space to obtain multiple boundary lines existing in the target image. As shown in FIG3a, multiple target points can be remapped into the image space by reverse Hough transform (RHT) to obtain multiple boundary lines existing in the target image.
  • RHT reverse Hough transform
  • the first target feature extraction network includes a plurality of downsampling modules, a plurality of upsampling modules, and a jump connection module.
  • feature extraction is performed on the target image to obtain an initial feature map, and the initial feature map is used as the first intermediate feature map of the maximum scale, and the first intermediate feature map of the maximum scale is downsampled (undersampling) processed N times by the downsampling module to obtain the first intermediate feature map of other scales, where N is a positive integer, for example, N can be 2, 3, 4 or 6, etc.; the first intermediate feature map of the minimum scale is used as the first target feature map of the minimum scale, and the first target feature map of the minimum scale is upsampled (upsampling) processed N times by the upsampling module, and in each upsampling process, a jump connection is performed with the first intermediate feature map of the same scale obtained by the downsampling process to obtain the first target feature map of other scales.
  • the second intermediate feature map of the scale of 128*128*3 is skipped and connected to the first intermediate feature map of the scale of 128*128*3 to obtain the first target feature map of the scale of 128*128*3; the first target feature map of the scale of 128*128*3 is upsampled once to obtain the second intermediate feature map of the scale of 256*256*3, the second intermediate feature map of the scale of 256*256*3 is skipped and connected to the first intermediate feature map of the scale of 256*256*3 to obtain the first target feature map of the scale of 256*256*3; the first target feature map of the scale of 256*256*3 is upsampled once to obtain the second intermediate feature map of the scale of 512*512*3, the second intermediate feature map of the scale of 512*512*3 is skipped and connected to the first intermediate feature map of the scale of 512*512*3 to obtain the first target feature map of the scale of 512*512*3.
  • the Hough transform network includes: a Hough transform module and a feature fusion module.
  • an implementation method of using a polar coordinate-based Hough transform to perform a Hough transform on a first target feature map of multiple scales to map a straight line in a target image to a point in a Hough space includes: using a Hough transform module to perform a Hough transform on a first target feature map X of multiple scales to obtain a second target feature map Y of multiple scales in a Hough space based on polar coordinates, wherein a Hough transform can be performed on the first target feature map X of each scale to obtain a second target feature map Y in a Hough space based on polar coordinates at the scale; using a feature fusion module to perform a scale transformation on the second target feature map of multiple scales to obtain multiple feature maps of the same scale, and splicing multiple feature maps of the same scale to obtain a third target feature map Z; performing convolution dimensionality reduction on
  • this embodiment does not limit the scale of the third target feature map.
  • the scale of the third target feature map can be any scale of the multiple scales corresponding to the second target feature map.
  • the scale of the third target feature map is the maximum scale of the multiple scales corresponding to the second target feature map. Based on this, when the second target feature maps of multiple scales are scaled, the second target feature map of the non-maximum scale can be upsampled to change its size to the maximum scale.
  • the feature fusion module is represented by the circled c, and the feature fusion module further includes an upsampling unit (upsample) and a concatenation unit (Concat).
  • the upsampling unit is used to scale the second target feature map of multiple scales to obtain multiple feature maps of the same scale; the concatenation unit (Concat) is used to concatenate multiple feature maps of the same scale to obtain the third target feature map Z.
  • model training can be performed in advance to obtain the boundary detection model based on Hough transform.
  • the combination of deep Hough transform and neural network model is innovatively applied to the detection of boundary lines in images. Specifically, a large number of pictures of spatial objects (such as indoor scenes) are obtained, and the wall lines, ground lines and ceiling lines in the pictures are annotated to obtain a sample data set. The sample data set is used to train a basic neural network based on Hough transform.
  • the basic neural network includes a first feature extraction network and a Hough transform network.
  • the first feature extraction network includes an upsampling module, a downsampling module and a jump connection module, which are used to extract features from sample images in pixel space to obtain multi-scale sample feature maps.
  • the Hough transform network includes a Hough transform The module and feature fusion module are used to transform the multi-scale sample feature map into the Hough space through the Hough transform and obtain the sample image in the Hough space. Then, the sample image in the Hough space is classified by using the binary classification method in the Hough space, specifically, the points in the sample image are divided into the first type of points corresponding to the boundary line and the second type of points corresponding to other straight lines; then, the loss function is generated according to the labeling results and the classification results.
  • the cross entropy loss (Binary CrossEntropy Loss, BCELoss) function can be used. If the loss function does not meet the standard, the training continues until the loss function meets the standard or the model training time or number reaches the set time or number, and the boundary detection model based on the Hough transform is obtained. It should be noted that in Figure 3a, the part of calculating the loss function based on the labeling results and the classification results is shown, which is used in the model training (Training only) stage; in addition, Figure 3a also shows the process of remapping multiple target points into the image space through the inverse Hough transform to obtain multiple boundary lines existing in the target image, which is used in the model reasoning stage.
  • BCELoss Binary CrossEntropy Loss
  • the vanishing point detection model based on Hough transform includes a second feature extraction network
  • the second feature extraction network can be any network capable of feature extraction, for example, a backbone network in image classification, for example, UNet, a stacked hourglass model, etc., wherein UNet is a variant of Fully Convolutional Networks (FCN), and its network structure is symmetrical and shaped like the English letter U, so it is called UNet; the Stacked Hourglass model is a network structure that uses multi-scale features to recognize postures.
  • FCN Fully Convolutional Networks
  • the target image can be input into the vanishing point detection model based on Hough transform, and the second feature extraction network is used to extract features from the target image to obtain a fourth target feature map; as shown in FIG3b, the target image is [512x512]x3 as an example, wherein [512x512] is the scale of the target image, and 3 is the number of channels; the fourth target feature map is [128x128]x3 as an example for illustration, but it is not limited thereto.
  • the fourth target feature map can be mapped to the Gaussian spherical space, for example, the straight line in the target image is projected to the Gaussian spherical space with the center of the target camera as the sphere center.
  • the vanishing point detection model based on the Hough transform also includes a Hough transform network based on the Gaussian sphere.
  • the fourth target feature map is sent to the Hough transform network based on the Gaussian sphere, in which the Hough transform based on the Gaussian sphere is used to perform a Hough transform on the fourth target feature map to map the straight lines in the target image to points in the Gaussian spherical space; then, at least two points whose probability values and angle values meet the requirements are selected in the Gaussian spherical space as the vanishing points in at least two orthogonal main directions, and remapped to the image space to obtain the vanishing points in at least two orthogonal main directions in the target image.
  • the Gaussian sphere space includes multiple points, each point is obtained by a straight line in the target image, and each point has two attributes: brightness value and angle value.
  • the brightness value indicates the probability value of the corresponding point being a point obtained by direct intersection (i.e., a vanishing point).
  • the angle value indicates the angle between the straight line corresponding to the point and the x-axis in the image space. According to the angle value, it can be determined whether the straight lines corresponding to each point on the Gaussian sphere are perpendicular to each other. Based on this, at least two points whose probability values and angle values meet the requirements can be selected from the Gaussian sphere space as vanishing points in at least two orthogonal main directions.
  • the Gaussian sphere-based Hough transform network includes a Hough transform module and a Gaussian sphere transform module. Based on this, after extracting the features of the target image to obtain the fourth target feature map, the fourth target feature map can be transformed into the Hough space and the Gaussian sphere space in sequence.
  • the straight line in the target image will be mapped to a point (referred to as the Hough point), but it is impossible to directly distinguish whether the Hough point is a point where multiple straight lines intersect; further, the Hough point is mapped to the Gaussian sphere space.
  • the response value (i.e., brightness value) of each point (referred to as the Gaussian point) will be different due to the different number of straight lines corresponding to the Gaussian point.
  • the response value of the Gaussian point formed by one straight line is smaller than the response value of the Gaussian point formed by the intersection of multiple straight lines. Therefore, several Gaussian points with the strongest response values can be selected as the vanishing points.
  • the fourth target feature map is input into the Hough transform module, and within the module, the fourth target feature map is subjected to a Hough transform to obtain a fifth target feature map in the Hough space based on polar coordinates, as shown in FIG3b, and the dimension of the fifth target feature map is [184x180]x128 as an example for illustration, but is not limited thereto, wherein 184 is the maximum value in the distance dimension represented by ⁇ , 180 is 180° in the ⁇ angle dimension, and 128 is the number of channels.
  • HT represents the Hough space
  • ( ⁇ , ⁇ ) is the coordinate in the Hough space.
  • the fifth target feature map is subjected to a Hough convolution in the Hough space to obtain a sixth target feature map, wherein the dimension of the sixth target feature map is unchanged relative to the fifth target feature map, and in FIG3b, the Hough convolution is represented as HT Conv.
  • the sixth target feature map is input into the Gaussian spherical transformation module. Inside the module, the sixth target feature map is subjected to a spherical Gaussian spherical transformation to obtain a seventh target feature map in the Gaussian spherical space.
  • ( ⁇ , ⁇ ) are variables in the Gaussian spherical space.
  • the dimension of the seventh target feature map is [32768]x128 as an example for illustration. 128 is the number of channels. [32768] indicates that the Gaussian sphere is discretized into 32768 points. Through the spherical Gaussian spherical transformation, it can be determined that the points in the Hough space correspond to the discrete points in the Gaussian spherical space.
  • the seventh target feature map is subjected to a spherical convolution in the Gaussian spherical space to obtain an eighth target feature map.
  • the eighth target feature map includes a plurality of Gaussian points, each of which corresponds to a straight line in the target image.
  • the spherical convolution is Shperical Conv
  • the dimension of the eighth target feature map is [32768]x128 as an example for illustration, but it is not limited thereto.
  • a point whose probability value meets the set requirements is selected from multiple Gaussian points as a vanishing point, for example, a point whose probability value exceeds a set probability threshold is used as a vanishing point, and the probability threshold may be 80%, 90% or 95%, etc.; at least two vanishing points whose angles are greater than the set angle are selected from the vanishing points as the vanishing points in at least two orthogonal main directions. It is explained here that two or three vanishing points with the largest angles can be selected from the vanishing points as the vanishing points in at least two orthogonal main directions. In the Gaussian sphere shown in FIG. 3b, the selection of three vanishing points is used as an example for illustration.
  • the above-mentioned method of determining the intrinsic camera parameters of the target camera and the gravity direction in the camera coordinate system based on the vanishing points in at least two orthogonal main directions includes: when the vanishing points in at least two orthogonal main directions include at least two finite vanishing points, selecting two target vanishing points from the at least two finite vanishing points; determining the intrinsic camera parameters of the target camera based on the constraint relationship between the two target vanishing points and the intrinsic camera parameters; and converting the vanishing points in the gravity direction in at least two orthogonal main directions to the camera coordinate system based on the intrinsic camera parameters to obtain the gravity direction in the camera coordinate system.
  • vanishing points may intersect at infinity or may not intersect at infinity.
  • a certain definition is given to infinity.
  • the vanishing point is considered to intersect at infinity.
  • the vanishing point that intersects at infinity is called an infinite vanishing point, and the vanishing point that does not intersect at infinity is called a finite vanishing point.
  • the vanishing points on at least two orthogonal main directions in the embodiment of the present application may also include infinite vanishing points and finite vanishing points.
  • the vanishing points on at least two orthogonal main directions include at least two finite vanishing points
  • two vanishing points are selected from the at least two finite vanishing points as target vanishing points.
  • the implementation method of selecting two target vanishing points from at least two finite vanishing points is not limited.
  • two vanishing points can be randomly selected from at least two finite vanishing points as target vanishing points; for another example, two vanishing points closest to the optical center of the camera can be selected from at least two finite vanishing points as target vanishing points.
  • the two target vanishing points may include a vanishing point located in the direction of gravity, or may not include a vanishing point located in the direction of gravity, and there is no limitation on this.
  • the camera intrinsic parameters of the target camera can be determined according to the constraint relationship between the two target vanishing points and the camera intrinsic parameters.
  • the straight line directions corresponding to the two target vanishing points in the three-dimensional space are perpendicular to each other.
  • VP 1 represents the three-dimensional vector of the first target vanishing point.
  • the three-dimensional vector is formed by adding a dimension after the image coordinates of the first target vanishing point. For example, 1 can be added after the image coordinates to form the three-dimensional vector of the first target vanishing point.
  • K represents the intrinsic parameters of the target camera, which is a 3X3 matrix including the focal length and the principal point position.
  • the principal point position is the position of the camera optical center. For example, the position with image coordinates (0.5, 0.5) can be taken as the camera optical center position, but it is not limited to this. Among them, the focal length is to be solved and the principal point position is known.
  • K -1 represents the inverse of the camera intrinsic parameters.
  • VP 2 is the three-dimensional vector of the second target vanishing point.
  • the three-dimensional vector is formed by adding a dimension after the image coordinates of the second target vanishing point. For example, 1 can be added after the image coordinates to form the three-dimensional vector of the second target vanishing point.
  • the vanishing point constrains the perspective relationship of the corresponding main direction.
  • the normal direction of a physical main structure such as a wall
  • Ni is the quantity to be determined
  • the vanishing point corresponding to the physical main structure is VPi
  • it is necessary to satisfy the constraint that the normal direction of the physical main structure and the main direction to which the parallel straight lines intersecting to obtain the vanishing point belong are parallel to each other, that is, Ni ⁇ (K -1 * VPi ) 1.
  • the reference physical main structure contained in the first spatial object can be determined.
  • the reference physical main structure can be the ground or the ceiling, depending on the physical main structure contained in the target image. For example, if the target image contains the ground, the ground is used as the reference physical main structure. If the target image does not contain the ground but contains the ceiling, the ceiling is used as the reference physical main structure. Body structure.
  • a reference plane of the three-dimensional space is constructed in the camera coordinate system, and the reference plane corresponds to the reference physical main structure; multiple reference intersection lines intersecting with the reference physical main structure are identified from multiple intersection lines.
  • the multiple reference intersection lines can be multiple ground lines obtained by the intersection of multiple walls and the ground. If the reference physical main structure is the ceiling, the multiple reference intersection lines can be multiple ceiling lines obtained by the intersection of multiple walls and the ceiling; according to the intersection positions between the multiple reference intersection lines, a reference model main structure corresponding to the reference physical main structure is constructed on the reference plane; according to the preset height from the camera optical center to the reference plane and the intersection positions between the multiple reference intersection lines, other model main structures corresponding to other physical main structures are constructed on the reference model main structure to obtain an initial three-dimensional model corresponding to the first spatial object, wherein, if the reference physical main structure is the ground, the other physical main structures can be walls, ceilings, etc. As shown in FIG.
  • FIG. 3e an exemplary display of the reference model main structure and other model main structures is shown.
  • the main structure of the reference model refers to the ground in the initial three-dimensional model
  • the main structures of other models refer to the walls in the initial three-dimensional model.
  • an implementation of constructing a reference model main structure corresponding to the reference physical main structure on a reference plane based on the intersection positions between multiple reference boundary lines includes: selecting valid reference boundary lines from multiple reference boundary lines, the number of valid reference boundary lines is multiple, and sorting the valid reference boundary lines according to the angles between the valid reference boundary lines and the x-axis in the image coordinate system to obtain the adjacent relationship between the valid reference boundary lines; determining the intersection positions between the valid reference boundary lines based on the adjacent relationship between the valid reference boundary lines, and dividing the intersection positions into a first intersection position that intersects with the boundary of the target image and a second intersection position that does not intersect with the boundary of the target image, and the first intersection position and the second intersection position are exemplarily shown in Figure 3c, but are not limited to this; based on the first intersection position and the second intersection position, drawing a model boundary corresponding to the reference physical main structure on the reference plane to obtain the reference model main structure, as shown in Figure 3d.
  • the grid area is the reference plane
  • the white area on the grid area is the reference model main structure formed by the model boundary corresponding to the reference physical main structure, which is connected to the intersection position shown in Figure 3c.
  • the reference model main structure in Figure 3d is the ground formed by the white area.
  • the embodiment of the present application does not limit the method of selecting a valid reference boundary line from multiple reference boundary lines, which is described below by way of example.
  • Example B1 Select a first baseline from multiple baselines based on the angles between the multiple baselines and the x-axis in the image coordinate system; for example, select a baseline from multiple baselines whose angle with the x-axis in the image coordinate system is less than a set angle threshold as the first baseline, and the angle threshold may be 10 degrees, 15 degrees, or 20 degrees, etc.
  • the baseline with the smallest angle may be selected from the baselines with an angle less than the angle threshold as the first baseline, or a baseline may be randomly selected from the baselines with an angle less than the angle threshold as the first baseline; then, based on the first baseline, remove some falsely detected baselines, specifically, remove the baselines with an angle less than the set angle threshold based on the angles between other baselines and the first baseline, and the angle threshold may be 3 degrees, 5 degrees, or 10 degrees, etc.; the baselines that are not removed and the first baseline are used as valid baselines.
  • Example B2 Selecting a first reference boundary line from multiple reference boundary lines based on their lengths;
  • a baseline intersection line whose length is greater than a set length threshold is selected from the baseline intersection lines as the first baseline intersection line.
  • the length threshold is not limited and depends on the size of the target image.
  • some falsely detected baseline intersection lines are eliminated based on the first baseline intersection line. Specifically, based on the angle between other baseline intersection lines and the first baseline intersection line, the baseline intersection lines whose angles are less than the set angle threshold are eliminated, and the baseline intersection lines that are not eliminated and the first baseline intersection line are used as valid baseline intersection lines.
  • Example B3 Select a first baseline from multiple baselines based on the angles between the multiple baselines and the x-axis in the image coordinate system and the lengths of the multiple baselines; for example, select a candidate baseline from multiple baselines whose angle with the x-axis in the image coordinate system is less than a set angle threshold, and then select a baseline from the candidate baselines as the first baseline; based on the angles between other baselines and the first baseline, eliminate baselines whose angles are less than the set angle threshold, and use the baselines that are not eliminated and the first baseline as valid baselines.
  • other model main structures corresponding to other physical main structures are constructed on the reference model main structure according to the preset height from the camera optical center to the reference plane and the intersection position between the multiple reference intersection lines, so as to obtain an implementation of the initial three-dimensional model corresponding to the first spatial object, including: determining the adjacent model boundaries on the reference model main structure intersecting at the second intersection position according to the second intersection position; determining the initial height of the other model main structures according to the preset height from the camera optical center to the reference plane and the preset scaling ratio; and constructing other model main structures on the adjacent model boundaries intersecting at the second intersection position according to the initial height of the other model main structures, so as to obtain the initial three-dimensional model corresponding to the first spatial object.
  • constructing other model main structures on the adjacent model boundaries intersecting at the second intersection position specifically includes constructing a wall on the adjacent model boundaries intersecting at the second intersection position, and further, supplementing a ceiling above the adjacent wall to obtain the initial three-dimensional model corresponding to the first spatial object.
  • the above-mentioned optimization of the initial three-dimensional model according to the constraint relationship between multiple boundary lines, the constraint relationship between the vanishing points in at least two orthogonal main directions, and the internal parameters of the camera to obtain a target three-dimensional model includes: constructing an optimization function with position parameters and/or height parameters of each model main structure as optimization variables according to the constraint relationship between multiple boundary lines, the constraint relationship between the vanishing points in at least two orthogonal main directions, and the internal parameters of the camera; solving the optimization function using a least squares algorithm to obtain optimized position parameters and/or height parameters of each model main structure, and adjusting the position of each model main structure according to the optimized position parameters and/or height parameters to obtain the target three-dimensional model.
  • the reference model main structure and other model main structures may have position parameters and/or height parameters, wherein the position parameters include the normal vector of the model main structure and the distance from the model main structure to the optical center of the camera, which two pieces of information can uniquely determine the position of the model main structure in the initial three-dimensional model and its relative position with other model main structures;
  • the height parameter is the height information from the model main structure to the reference plane, for example, the height information of the ceiling, the height information of the wall and the height information of the ground, wherein if the ground is taken as the reference plane, the height information of the ground is 0.
  • At least one of the first type of optimization items, the second type of optimization items, and the third type of optimization items can be constructed according to the constraint relationship between the multiple boundary lines, the constraint relationship between the vanishing points in at least two orthogonal main directions, and the camera internal parameters, and an optimization function can be generated according to the at least one optimization item.
  • the optimization function can be generated according to the first type of optimization items, the second type of optimization items, and the third type of optimization items at the same time.
  • a reference normal vector of the model main structure in the camera coordinate system is generated according to the camera internal parameters and the vanishing point perpendicular to the model main structure.
  • the reference normal vector and the normal vector of the model main structure should be parallel. Therefore, the first type of optimization item can be constructed with the dot product of the normal vector of the model main structure and the reference normal vector as 1 as the optimization target.
  • the normal vectors of any adjacent model main structure should be vertical, for example, the ground and the wall are vertical, the wall and the ceiling are vertical, and the adjacent walls are vertical. Therefore, the second type of optimization item can be constructed by taking the dot product of the normal vectors between any adjacent model main structures as 0 as the optimization target.
  • the boundary line of any adjacent model main structure in the image coordinate system is generated according to the camera internal parameters, the normal vector of any adjacent model main structure and the distance from any adjacent model main structure to the camera optical center; in addition, the boundary line between the corresponding adjacent physical main structures in the target image is detected based on the Hough transform.
  • the boundary line of any adjacent model main structure in the image coordinate system (the boundary line here can be regarded as the projection of the boundary line between any adjacent model main structures in the initial three-dimensional model in the image space) and the boundary line between the corresponding adjacent physical main structures detected from the target image (the boundary line here can be regarded as the projection of the boundary line between the adjacent physical main structures in the first spatial object in the image space) should be the same. Therefore, the third type of optimization item can be constructed with the optimization goal that the boundary line of any adjacent model main structure in the image coordinate system is the same as the boundary line between the corresponding adjacent physical main structures in the target image.
  • the following is an example of the construction process of the above optimization items by combining the plane equations corresponding to the main structure of the model.
  • the main structure of the model to be optimized is the wall, ceiling and ground.
  • the wall equation to be optimized is ⁇ n i ,d i ⁇
  • the normal vector of the i-th wall estimated by the vanishing point is The boundary line between the i-th wall and the ground detected from the image is recorded as The boundary line between the i-th wall and the j-th wall detected from the image is recorded as The boundary line between the i-th wall and the ceiling detected from the image is recorded as Among them,
  • ni is the normal vector of the ith wall, which is the quantity to be optimized and has an initial value (i.e., the value determined in the initial 3D model)
  • d i is the distance from the ith wall to the optical center of the camera
  • the height of the ceiling d ceiling is the value assumed when constructing the initial 3D model and
  • T represents the set of walls that appear in the image.
  • the least squares algorithm is used to solve the optimization function, with the minimum sum of all items in the optimization function as the optimization goal, to obtain the position parameters of the wall (such as the wall equation) and the ceiling
  • the height parameter of the board contains six items, among which the first item belongs to the first type of optimization item, the second and third items belong to the second type of optimization items, and the fourth, fifth and sixth items belong to the third type of optimization items. The following is a detailed description:
  • the first term is the normal vector n i of the i-th wall and the normal vector of the wall obtained from the vanishing point.
  • the directions of the two should be consistent, the dot product of the two should be 1, and after subtracting 1, it should be 0, where ni is the amount to be optimized.
  • ni and nj belong to the normal vectors of adjacent walls.
  • the normal vectors of adjacent walls are perpendicular, and the dot product of their normal vectors is 0.
  • Both ni and nj are quantities to be optimized.
  • ni and nfloor represent the wall and the ground respectively.
  • the wall and the ground are perpendicular to each other.
  • the dot product of ni and nfloor is 0.
  • nfloor is known, that is, the direction of gravity in the camera coordinate system, and ni is the quantity to be optimized.
  • the fourth item represents the boundary line between the i-th wall and the ground, Indicates that the boundary line is projected into the image coordinate system, and the boundary line projected into the image coordinate system is the same as the boundary line detected from the image The difference is 0, It is known that the boundary between the wall and the ground in the target image is detected based on Hough transform.
  • the fifth item is the intersection of the ceiling line and the ceiling line detected from the image. The distance between them is 0.
  • the sixth item is the boundary line between the adjacent i-th wall and the j-th wall, and the boundary line between the i-th wall and the j-th wall detected from the image. The difference between them is 0.
  • the target three-dimensional model corresponding to the first spatial object can be obtained.
  • the 3D structural information is not directly detected from the target image, but the 2D visual information such as the boundary line and the vanishing point is directly detected from the target image, and then, the 3D structural information is generated using these 2D visual information and the Manhattan constraint and the vanishing point constraint of the indoor scene.
  • the embodiment of the present application has higher stability and stronger interpretability, and is more robust for the situation where the user's shooting angle is tilted and the scene is messy.
  • the boundary lines such as the wall line, the ground line and the ceiling line are directly detected from the target image, and these boundary lines are used as constraints to optimize the three-dimensional equation of the wall.
  • the target three-dimensional model generated at last has a higher degree of fit with the actual wall, ground and ceiling dividing lines, and the model quality is higher.
  • other real scene information such as texture in the 2D image can also be reflected in the target three-dimensional model.
  • the target three-dimensional model can be displayed to the user so that the user can view or understand the three-dimensional structure of the first spatial object.
  • the target three-dimensional model is applied to an online home improvement scene for viewing the matching effect of home improvement objects (such as paintings, sofas, etc.) and the real environment of the first spatial object, and selecting a home improvement plan based on the matching effect.
  • goods are purchased based on the target three-dimensional model to view the matching effect of the goods to be purchased and the real environment of the first spatial object, and goods are purchased based on the matching effect.
  • the following embodiments will be described using online home improvement scenes and online shopping scenes as examples.
  • FIG4a is a flow chart of an online home improvement method provided by an exemplary embodiment of the present application. As shown in FIG4a , the method includes:
  • a target image corresponding to a first spatial object is obtained, where the first spatial object is at least a portion of a spatial object in a target physical space;
  • the target three-dimensional model corresponding to the object is obtained to obtain the target three-dimensional model of the fused target home improvement object;
  • the online home improvement method can be implemented on a terminal device, or can be implemented by the terminal device and the server device in cooperation, which is not limited. The following description is taken as an example of the implementation by the terminal device and the server device in cooperation.
  • the user can open the home improvement App (Application) on the terminal device and enter the online home improvement page; on the online home improvement page, there are a variety of home improvement methods, such as home improvement methods based on model rooms, home improvement methods based on real background images, etc.; users can choose home improvement methods based on real background images, and in response to the user's selection of home improvement methods based on real background images, the addition page of the real background image can be displayed.
  • the user can collect the target image of the first spatial object through the terminal device, or directly select the target image of the first spatial object from the gallery; the home improvement App responds to the image acquisition operation or the image selection operation, and can obtain the target image corresponding to the first spatial object.
  • the first spatial object can be at least part of the spatial object in the target physical space.
  • the target physical space is a house to be renovated by the user, which includes a bedroom, a living room, a kitchen or a bathroom, etc.
  • the first spatial object can be a bedroom and a living room, or it can be a partial space of the living room. There is no limitation on this. For details, please refer to the aforementioned embodiment, which will not be repeated here.
  • the target image can be uploaded to the server device.
  • the server device uses the three-dimensional scene reconstruction method described in the aforementioned method embodiment to generate a target three-dimensional model corresponding to the first spatial object.
  • the generation method of the target three-dimensional model please refer to the aforementioned embodiment, which will not be repeated here. It is explained here that the process of generating the target three-dimensional model is imperceptible to the user.
  • Figure 4b a schematic diagram of the target image and the process from the target image to the generation of the target three-dimensional model is shown in an embodiment of the present application.
  • the home improvement app can display a variety of home improvement objects in the associated area of the target image.
  • the home improvement objects can be home appliances, furniture, or decorative items.
  • the display of a sofa, a coffee table, a carpet, a single chair, and a painting in the delivery area of the target image is used as an example.
  • the user can select a target home improvement object from a variety of home improvement objects and place the target home improvement object on the target image. For example, the image of the selected target home improvement object can be dragged to the corresponding position on the target image and then released.
  • the home improvement App can respond to the placement operation of the target home improvement object on the target image, and provide the identification information of the target home improvement object (such as name, product ID or image, etc.) and the position range information of the target home improvement object on the target image to the server device; the server device obtains the three-dimensional model corresponding to the target home improvement according to the identification information of the target home improvement object, and fuses the three-dimensional model of the target home improvement object into the target three-dimensional model according to the position range information of the target home improvement object on the target image, and projects the target three-dimensional model of the fused target home improvement object onto the target image to obtain a home improvement rendering containing the target home improvement object, and provides the home improvement rendering to the terminal device; the terminal device displays the home improvement rendering.
  • the identification information of the target home improvement object such as name, product ID or image, etc.
  • the server device obtains the three-dimensional model corresponding to the target home improvement according to the identification information of the target home improvement object, and fuses the three-dimensional model of the target home improvement object
  • Figure 4c taking the example of a user selecting a painting and placing the painting in the center of the main wall, the perspective view of the home improvement rendering shown in Figure 4c The deformation is correct, and the moving area of the painting can be limited to the wall part, with high accuracy.
  • Figure 4d taking the example of the user selecting a sofa and a carpet in turn, placing the sofa on the midline of the main wall and the ground, and laying the carpet on the ground, the perspective relationship in the home decoration rendering shown in Figure 4d is correct and the placement position is reasonable.
  • the terminal device and the server device cooperate with each other to implement the online home improvement process as an example, but it is not limited to this. Of course, it can also be implemented independently by the terminal device.
  • the home improvement App can use the three-dimensional scene reconstruction method described in the aforementioned method embodiment to generate a target three-dimensional model corresponding to the first spatial object; at the same time, the home improvement App can display a variety of home improvement objects in the associated area of the target image, and the home improvement App can respond to the placement operation of the target home improvement object on the target image, obtain the identification information of the target home improvement object and the position range information of the target home improvement object on the target image; according to the identification information of the target home improvement object, obtain the three-dimensional model corresponding to the target home improvement, and according to the position range information of the target home improvement object on the target image, fuse the three-dimensional model of the target home improvement object into the target three-dimensional model, and project the target three-dimensional model of the fused
  • target three-dimensional model can be constructed when the target image is uploaded for the first time, and the pre-constructed target three-dimensional model can be directly used in subsequent online home improvement.
  • FIG4e is a flow chart of a commodity selection method provided by an exemplary embodiment of the present application. As shown in FIG4e , the method includes:
  • the commodity selection method can be implemented on a terminal device, or can be implemented by the terminal device and the server device in cooperation, and this is not limited.
  • the following description is made by taking the commodity selection method implemented on a terminal device as an example, but it is not limited to this.
  • the user in a shopping scene, when a user buys some furniture, home appliances or decorative products, in order to facilitate the user to choose products that are more suitable for the actual environment, the user can add the selected products to the target image corresponding to the first spatial object and view the matching effect after adding.
  • the first spatial object can be any area with a spatial concept, such as a home environment, a shopping mall environment, an office scene, etc.
  • the product can be any item that needs to be placed in the first spatial object.
  • the product in a home environment, can be a wardrobe, a painting or a table and chair, etc.
  • the product in a shopping mall environment, the product can be a locker, a shelf or a shelf, etc.
  • a product page can be displayed on the terminal device.
  • a variety of products are displayed on the product page.
  • the user can By selecting a target product on a product page, the terminal device can respond to the selection operation on the product page and determine the selected target product, and the target product corresponds to a product three-dimensional model.
  • a collocation effect viewing control can be added to the shopping cart page, the order page or the product details page, and the user can initiate an operation to view the collocation effect of the product through the control.
  • the e-commerce App can allow the user to select the target image corresponding to the first spatial object; the user can collect the target image of the first spatial object through the terminal device, or directly select the target image of the first spatial object from the gallery; after the e-commerce App obtains the target image, the three-dimensional scene reconstruction method provided in the above embodiment can be used to construct a target three-dimensional model corresponding to the first spatial object, and the three-dimensional model of the target product selected by the user is added to the target three-dimensional model corresponding to the first spatial object to obtain a target three-dimensional model of the fused target product; the target three-dimensional model of the fused target product is projected onto the target image to obtain a collocation effect diagram of the target product and the first spatial object, and the collocation effect diagram is displayed.
  • the user can determine whether to purchase the target product through the collocation effect diagram. Further optionally, the collocation effect diagram can be automatically scored and a collocation score can be given according to a preset collocation rule or strategy to assist the user in determining whether to purchase the target product, which is conducive to improving the user's shopping experience and increasing the probability of the user successfully purchasing the desired product, thereby reducing the probability of return and exchange.
  • the execution subject of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices.
  • the execution subject of steps 101 to 103 can be device A; for another example, the execution subject of steps 101 and 102 can be device A, and the execution subject of step 103 can be device B; and so on.
  • FIG5 is a schematic diagram of the structure of a three-dimensional scene reconstruction device provided by an exemplary embodiment of the present application. As shown in FIG5 , the device includes: an acquisition module 51 , a detection module 52 , a determination module 53 , a reconstruction module 54 and an optimization module 55 .
  • An acquisition module 51 is used to acquire a target image corresponding to a first spatial object, where the first spatial object is at least a part of a spatial object in a target physical space;
  • a detection module 52 configured to detect a plurality of boundary lines existing in the target image and vanishing points in at least two orthogonal main directions, wherein the boundary lines are intersection lines between adjacent physical main structures in the first spatial object;
  • a determination module 53 configured to determine the camera intrinsic parameters of a target camera and the gravity direction in a camera coordinate system according to the vanishing points in at least two orthogonal principal directions, wherein the target camera refers to a camera used to capture a target image;
  • a reconstruction module 54 configured to reconstruct an initial three-dimensional model corresponding to the first spatial object according to the plurality of boundary lines and the gravity direction in the camera coordinate system, wherein the initial three-dimensional model includes adjacent model main structures corresponding to the adjacent physical main structures;
  • the optimization module 55 is used to optimize the initial three-dimensional model according to the constraint relationship between the multiple boundary lines, the constraint relationship between the vanishing points in at least two orthogonal main directions and the camera internal parameters to obtain the target three-dimensional model.
  • the detection module 52 is specifically used to: input the target image into a boundary line detection model based on Hough transform to perform boundary line detection to obtain multiple boundary lines existing in the target image; input the target image into a vanishing point detection model based on Hough transform to perform vanishing point detection to obtain vanishing points in at least two orthogonal main directions in the target image.
  • the detection module 52 is specifically used to: input the target image into a boundary detection model based on Hough transform, in which a first feature extraction network that integrates jump connections and multi-scale features is used to extract features of the target image to obtain first target feature maps of multiple scales; perform Hough transform on the first target feature maps of multiple scales using a Hough transform based on polar coordinates to map straight lines in the target image to points in the Hough space; select multiple target points in the Hough space that match the boundary features, and remap the multiple target points to the image space to obtain multiple boundary lines in the target image.
  • Hough transform in which a first feature extraction network that integrates jump connections and multi-scale features is used to extract features of the target image to obtain first target feature maps of multiple scales
  • the detection module 52 is specifically used to: extract features from the target image to obtain a first intermediate feature map of the largest scale, perform N downsampling processes on the first intermediate feature map of the largest scale to obtain multiple first intermediate feature maps of other scales; use the first intermediate feature map of the smallest scale as the first target feature map of the smallest scale, perform N upsampling processes on the first target feature map of the smallest scale, and perform jump connections with the first intermediate feature map of the same scale in each upsampling process to obtain first target feature maps of other scales.
  • the detection module 52 is specifically used to: perform Hough transform on the first target feature map of multiple scales to obtain second target feature maps of multiple scales in the Hough space based on polar coordinates; perform scale transform on the second target feature map of multiple scales to obtain multiple feature maps of the same scale, and splice the multiple feature maps of the same scale to obtain a third target feature map; perform convolution dimensionality reduction on the third target feature map to obtain a two-dimensional image in the Hough space, wherein the two-dimensional image contains multiple points, each point corresponding to a straight line in the target image.
  • the detection module 52 is specifically used to: input the target image into a vanishing point detection model based on Hough transform, in which the second feature extraction network is used to extract features of the target image to obtain a fourth target feature map; perform Hough transform on the fourth target feature map using Hough transform based on Gaussian sphere to map the straight lines existing in the target image to points in Gaussian sphere space; select vanishing points in at least two orthogonal main directions whose probability values and angle values meet the requirements in Gaussian sphere space, and remap them into the image space to obtain vanishing points in at least two orthogonal main directions in the target image.
  • the detection module 52 is specifically used to: perform Hough transform on the fourth target feature map to obtain a fifth target feature map in a Hough space based on polar coordinates; perform Hough convolution on the fifth target feature map in the Hough space to obtain a sixth target feature map; perform spherical Gaussian spherical transform on the sixth target feature map to obtain a seventh target feature map in a Gaussian spherical space; perform spherical convolution on the seventh target feature map in the Gaussian spherical space to obtain an eighth target feature map, the eighth target feature map including multiple points, each point corresponding to a straight line in the target image.
  • the detection module 52 is specifically used to: select points whose probability values meet the set requirements from multiple points as vanishing points based on the probability values of each point in the eighth target feature map; select at least two vanishing points whose angles are greater than the set angle from the vanishing points as vanishing points in at least two orthogonal main directions.
  • the determination module 53 is specifically configured to: include the vanishing points in at least two orthogonal main directions
  • two target vanishing points closest to the optical center of the camera are selected from the at least two finite vanishing points;
  • the intrinsic parameters of the target camera are determined according to the constraint relationship between the two target vanishing points and the intrinsic parameters of the camera; according to the intrinsic parameters of the camera, the vanishing points in the gravity direction in at least two orthogonal main directions are converted to the camera coordinate system to obtain the gravity direction in the camera coordinate system.
  • the reconstruction module 54 is specifically used to: construct a reference plane of the three-dimensional space in the camera coordinate system according to the gravity direction in the camera coordinate system, and the reference plane corresponds to the reference physical main structure contained in the first spatial object; identify multiple reference boundary lines that intersect with the reference physical main structure from multiple boundary lines, and construct a reference model main structure corresponding to the reference physical main structure on the reference plane according to the intersection positions between the multiple reference boundary lines; construct other model main structures corresponding to other physical main structures on the reference model main structure according to the preset height from the camera optical center to the reference plane and the intersection positions between the multiple reference boundary lines, so as to obtain an initial three-dimensional model corresponding to the first spatial object.
  • the reconstruction module 54 is specifically used to: select a valid benchmark boundary line from multiple benchmark boundary lines, sort the valid benchmark boundary lines according to the angle between the valid benchmark boundary lines and the x-axis in the image coordinate system to obtain the adjacent relationship between the valid benchmark boundary lines; determine the intersection positions between the valid benchmark boundary lines according to the adjacent relationship between the valid benchmark boundary lines, and divide the intersection positions into a first intersection position that intersects with the boundary of the target image and a second intersection position that does not intersect with the boundary of the target image; draw a model boundary corresponding to the benchmark physical main structure on the benchmark plane according to the first intersection position and the second intersection position to obtain the benchmark model main structure.
  • the reconstruction module 54 is specifically used to: select a first baseline intersection line from multiple baseline intersection lines based on the angles between the multiple baseline intersection lines and the x-axis in the image coordinate system and/or the lengths of the multiple baseline intersection lines; eliminate baseline intersection lines with angles less than a set angle threshold based on the angles between other baseline intersection lines and the first baseline intersection line, and use the baseline intersection lines that are not eliminated and the first baseline intersection line as valid baseline intersection lines.
  • the reconstruction module 54 is specifically used to: determine the adjacent model boundaries on the main structure of the reference model that intersect at the second intersection position according to the second intersection position; determine the initial heights of other model main structures according to a preset height from the camera optical center to the reference plane and a preset scaling ratio; and construct other model main structures on the adjacent model boundaries that intersect at the second intersection position according to the initial height to obtain an initial three-dimensional model corresponding to the first spatial object.
  • the optimization module 55 is specifically used to: construct an optimization function with position parameters and/or height parameters of each model main structure as optimization variables according to the constraint relationship between multiple intersection lines, the constraint relationship between vanishing points in at least two orthogonal main directions, and internal camera parameters, the position parameters including the normal vector of the model main structure and the distance to the optical center of the camera; solve the optimization function using a least squares algorithm to obtain the optimized position parameters and/or height parameters of each model main structure, and adjust the position of each model main structure according to the optimized position parameters and/or height parameters to obtain the target three-dimensional model.
  • the optimization module 55 is specifically used to: for each model main structure, generate a reference normal vector of the model main structure in the camera coordinate system according to the camera internal parameters and the vanishing point perpendicular to the model main structure, and construct a first type of optimization item with the dot product of the normal vector of the model main structure and the reference normal vector being 1 as the optimization target; for any adjacent model main structure, the dot product of the normal vectors between any adjacent model main structures being 0 as the optimization target;
  • the second type of optimization items are constructed as the goal; for any adjacent model main structure, the boundary line of any adjacent model main structure in the image coordinate system is generated according to the camera internal parameters and the normal vector of any adjacent model main structure and the distance to the camera optical center;
  • the third type of optimization items are constructed with the boundary line of any adjacent model main structure in the image coordinate system being the same as the boundary line between the corresponding adjacent physical main structures in the target image as the optimization goal; the optimization function is generated according to the first type of optimization items, the second type of optimization items and the third type of
  • the embodiment of the present application also provides an online home improvement device, which includes: an acquisition module, a fusion module and a projection module.
  • the acquisition module is used to respond to the image upload operation and acquire the target image corresponding to the first spatial object, and the first spatial object is at least a part of the spatial object in the target physical space.
  • the fusion module is used to respond to the placement operation of the target home improvement object on the target image and fuse the target home improvement object into the target three-dimensional model corresponding to the first spatial object to obtain a target three-dimensional model of the fused target home improvement object.
  • the projection module is used to project the target three-dimensional model of the fused target home improvement object onto the target image to obtain a home improvement rendering containing the target home improvement object; wherein the target three-dimensional model is constructed according to the steps in the three-dimensional scene reconstruction method provided in the embodiment of the present application.
  • the embodiment of the present application also provides a product selection device, which includes: a determination module, a selection module, an addition module and a projection module.
  • the determination module is used to respond to the selection operation on the product page to determine the selected target product, and the target product has a three-dimensional model of the product.
  • the selection module is used to respond to the matching effect viewing operation and select the target image corresponding to the first spatial object to be matched with the target product.
  • the adding module is used to add the three-dimensional model of the product to the target three-dimensional model corresponding to the first spatial object to obtain a target three-dimensional model of the fused target product.
  • the fusion module is used to project the target three-dimensional model of the fused target product onto the target image to obtain a matching effect diagram of the target product and the first spatial object; wherein the target three-dimensional model is constructed according to the steps in the three-dimensional scene reconstruction method provided in the embodiment of the present application.
  • Fig. 6 is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present application. As shown in Fig. 6 , the device includes: a memory 64 and a processor 65 .
  • the memory 64 is used to store computer programs and may be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application program or method operating on the computing platform.
  • the memory 64 can be implemented by any type of volatile or nonvolatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
  • SRAM static random access memory
  • EEPROM electrically erasable programmable read-only memory
  • EPROM erasable programmable read-only memory
  • PROM programmable read-only memory
  • ROM read-only memory
  • magnetic storage flash memory
  • flash memory magnetic disk or optical disk.
  • the processor 65 is coupled to the memory 64 and is used to execute the computer program in the memory 64, so as to: obtain a target image corresponding to a first spatial object, where the first spatial object is at least a part of a spatial object in a target physical space; detect multiple boundary lines and vanishing points in at least two orthogonal main directions in the target image, where the boundary line is an intersection line between adjacent physical main structures in the first spatial object; and detect the vanishing points in at least two orthogonal main directions according to the vanishing points in the at least two orthogonal main directions.
  • the target camera refers to the camera used to shoot the target image
  • the target camera refers to the camera used to shoot the target image
  • reconstruct the initial three-dimensional model corresponding to the first spatial object the initial three-dimensional model includes adjacent model main structures corresponding to the adjacent physical main structures
  • the constraint relationship between the multiple boundary lines the constraint relationship between the vanishing points in at least two orthogonal main directions and the camera intrinsic parameters, optimize the initial three-dimensional model to obtain the target three-dimensional model.
  • the processor 65 when the processor 65 detects multiple boundary lines and vanishing points in at least two orthogonal main directions in the target image based on the Hough transform, it is specifically used to: input the target image into a boundary line detection model based on the Hough transform to perform boundary line detection to obtain the multiple boundary lines in the target image; input the target image into a vanishing point detection model based on the Hough transform to perform vanishing point detection to obtain the vanishing points in at least two orthogonal main directions in the target image.
  • the processor 65 when the processor 65 inputs the target image into a junction detection model based on Hough transform to detect the junction to obtain multiple junctions in the target image, it is specifically used to: input the target image into a junction detection model based on Hough transform, in which a first feature extraction network that integrates jump connections and multi-scale features is used to extract features of the target image to obtain first target feature maps of multiple scales; perform Hough transform on the first target feature maps of multiple scales using a Hough transform based on polar coordinates to map the straight lines in the target image to points in the Hough space; select multiple target points in the Hough space that match the junction features, and remap the multiple target points to the image space to obtain multiple junctions in the target image.
  • Hough transform in which a first feature extraction network that integrates jump connections and multi-scale features is used to extract features of the target image to obtain first target feature maps of multiple scales
  • the processor 65 uses the first feature extraction network that integrates jump connections and multi-scale features to extract features of the target image to obtain first target feature maps of multiple scales, it is specifically used to: extract features from the target image to obtain a first intermediate feature map of the largest scale, downsample the first intermediate feature map of the largest scale N times to obtain first intermediate feature maps of other scales; use the first intermediate feature map of the smallest scale as the first target feature map of the smallest scale, upsample the first target feature map of the smallest scale N times, and perform jump connections with the first intermediate feature map of the same scale in each upsampling process to obtain first target feature maps of other scales.
  • the processor 65 uses the polar coordinate-based Hough transform to perform Hough transform on the first target feature maps of multiple scales to map the straight lines in the target image to points in the Hough space, it is specifically used to: perform Hough transform on the first target feature maps of multiple scales to obtain second target feature maps of multiple scales in the Hough space based on polar coordinates; perform scale transformation on the second target feature maps of multiple scales to obtain multiple feature maps of the same scale, and splice the multiple feature maps of the same scale to obtain a third target feature map; perform convolution dimensionality reduction on the third target feature map to obtain a two-dimensional image in the Hough space, wherein the two-dimensional image contains multiple points, each point corresponding to a straight line in the target image.
  • the processor 65 when the processor 65 inputs the target image into the vanishing point detection model based on Hough transform to perform vanishing point detection to obtain vanishing points in at least two orthogonal main directions in the target image, the processor 65 is specifically used to: input the target image into the vanishing point detection model based on Hough transform, in which the second feature extraction is used The network extracts features of the target image to obtain a fourth target feature map; the fourth target feature map is subjected to Hough transform based on Gaussian sphere to map straight lines in the target image to points in Gaussian sphere space; in Gaussian sphere space, vanishing points in at least two orthogonal main directions whose probability values and angle values meet requirements are selected, and remapped into the image space to obtain vanishing points in at least two orthogonal main directions in the target image.
  • the processor 65 when the processor 65 performs Hough transform on the fourth target feature map using the Hough transform based on the Gaussian sphere to map the straight lines existing in the target image to points in the Gaussian sphere space, it is specifically used to: perform Hough transform on the fourth target feature map to obtain a fifth target feature map in the Hough space based on polar coordinates; perform Hough convolution on the fifth target feature map in the Hough space to obtain a sixth target feature map; perform spherical Gaussian spherical transform on the sixth target feature map to obtain a seventh target feature map in the Gaussian sphere space; perform spherical convolution on the seventh target feature map in the Gaussian sphere space to obtain an eighth target feature map, the eighth target feature map including multiple points, each point corresponding to a straight line existing in the target image.
  • the processor 65 when the processor 65 selects vanishing points in at least two orthogonal main directions whose probability values and angle values meet the requirements in the Gaussian sphere space, it is specifically used to: select points whose probability values meet the set requirements from multiple points as vanishing points based on the probability values of each point in the eighth target feature map; select at least two vanishing points whose angles are greater than the set angle from the vanishing points as vanishing points in at least two orthogonal main directions.
  • the processor 65 determines the intrinsic camera parameters of the target camera and the gravity direction in the camera coordinate system based on the vanishing points in at least two orthogonal principal directions
  • the processor 65 is specifically used to: when the vanishing points in at least two orthogonal principal directions include at least two finite vanishing points, select two target vanishing points closest to the optical center of the camera from the at least two finite vanishing points; determine the intrinsic camera parameters of the target camera based on the constraint relationship between the two target vanishing points and the intrinsic camera parameters; and convert the vanishing points in the gravity direction in at least two orthogonal principal directions into the camera coordinate system based on the intrinsic camera parameters to obtain the gravity direction in the camera coordinate system.
  • the processor 65 when the processor 65 reconstructs the initial three-dimensional model corresponding to the first spatial object based on multiple boundary lines and the gravity direction in the camera coordinate system, it is specifically used to: construct a reference plane of the three-dimensional space in the camera coordinate system according to the gravity direction in the camera coordinate system, and the reference plane corresponds to the reference physical main structure contained in the first spatial object; identify multiple reference boundary lines that intersect with the reference physical main structure from multiple boundary lines, and construct a reference model main structure corresponding to the reference physical main structure on the reference plane according to the intersection positions between the multiple reference boundary lines; construct other model main structures corresponding to other physical main structures on the reference model main structure according to the preset height from the camera optical center to the reference plane and the intersection positions between the multiple reference boundary lines, so as to obtain the initial three-dimensional model corresponding to the first spatial object.
  • the processor 65 when the processor 65 constructs a reference model main structure corresponding to the reference physical main structure on the reference plane according to the intersection positions between the plurality of reference boundary lines, the processor 65 is specifically used to: select a valid reference boundary line from the plurality of reference boundary lines, sort the valid reference boundary lines according to the angles between the valid reference boundary lines and the x-axis in the image coordinate system to obtain the adjacent relationship between the valid reference boundary lines; determine the intersection positions between the valid reference boundary lines according to the adjacent relationship between the valid reference boundary lines, and divide the intersection positions into a first intersection position that intersects with the boundary of the target image and a second intersection position that does not intersect with the boundary of the target image; and At the first intersection point position and the second intersection point position, a model boundary corresponding to the reference physical main structure is drawn on the reference plane to obtain the reference model main structure.
  • the processor 65 when the processor 65 selects a valid baseline intersection line from multiple baseline intersection lines, it is specifically used to: select a first baseline intersection line from multiple baseline intersection lines based on the angles between the multiple baseline intersection lines and the x-axis in the image coordinate system and/or the lengths of the multiple baseline intersection lines; based on the angles between other baseline intersection lines and the first baseline intersection line, eliminate baseline intersection lines with angles less than a set angle threshold, and use the baseline intersection lines that are not eliminated and the first baseline intersection line as valid baseline intersection lines.
  • the processor 65 when the processor 65 constructs other model main structures corresponding to other physical main structures on the reference model main structure according to the preset height from the camera optical center to the reference plane and the intersection position between multiple reference intersection lines to obtain the initial three-dimensional model corresponding to the first spatial object, it is specifically used to: determine the adjacent model boundaries on the reference model main structure that intersect at the second intersection position according to the second intersection position; determine the initial height of the other model main structures according to the preset height from the camera optical center to the reference plane and the preset scaling ratio; and construct other model main structures on the adjacent model boundaries that intersect at the second intersection position according to the initial height to obtain the initial three-dimensional model corresponding to the first spatial object.
  • the processor 65 when the processor 65 optimizes the initial three-dimensional model according to the constraint relationship between multiple boundary lines, the constraint relationship between the vanishing points in at least two orthogonal main directions, and the internal parameters of the camera to obtain the target three-dimensional model, it is specifically used to: construct an optimization function with position parameters and/or height parameters of each model main structure as optimization variables according to the constraint relationship between multiple boundary lines, the constraint relationship between the vanishing points in at least two orthogonal main directions, and the internal parameters of the camera, the position parameters including the normal vector of the model main structure and the distance to the optical center of the camera; solve the optimization function using a least squares algorithm to obtain the optimized position parameters and/or height parameters of each model main structure, and adjust the position of each model main structure according to the optimized position parameters and/or height parameters to obtain the target three-dimensional model.
  • the processor 65 when the processor 65 constructs an optimization function with position parameters and/or height parameters of each model main structure as optimization variables according to the constraint relationship between multiple intersection lines, the constraint relationship between the vanishing points in at least two orthogonal main directions, and the camera intrinsic parameters, the processor 65 is specifically used to: for each model main structure, generate a reference normal vector of the model main structure in the camera coordinate system according to the camera intrinsic parameters and the vanishing point perpendicular to the model main structure, and construct a first type of optimization item with the dot product of the normal vector of the model main structure and the reference normal vector being 1 as the optimization target; for any adjacent model main structure, construct a second type of optimization item with the dot product of the normal vectors between any adjacent model main structures being 0 as the optimization target; for any adjacent model main structure, generate the intersection line of any adjacent model main structure in the image coordinate system according to the camera intrinsic parameters and the normal vector of any adjacent model main structure and the distance to the camera optical center; construct a third type of optimization item with the intersection line of any adjacent model
  • the electronic device also includes: a communication component 66, a display 67, a power component 68, an audio component 69 and other components.
  • FIG6 only schematically shows some components, which does not mean that the electronic device only includes FIG6 Components shown.
  • the embodiment of the present application also provides an electronic device, the structure of which is the same or similar to that of the electronic device shown in FIG6 , and the structure of the electronic device shown in FIG6 can be specifically referred to.
  • the difference between the electronic device provided in the present embodiment and the electronic device in the embodiment shown in FIG6 is mainly that: the functions implemented by the processor executing the computer program stored in the memory are different.
  • the electronic device For the electronic device provided in the present embodiment, its processor executes the computer program stored in the memory, which can be used to: respond to the image upload operation, obtain the target image corresponding to the first spatial object, and the first spatial object is at least part of the spatial object in the target physical space; respond to the placement operation of the target home improvement object on the target image, fuse the target home improvement object into the target three-dimensional model corresponding to the first spatial object, so as to obtain the target three-dimensional model of the fused target home improvement object; project the target three-dimensional model of the fused target home improvement object onto the target image, so as to obtain the home improvement effect diagram containing the target home improvement object; wherein, the target three-dimensional model is constructed according to the steps in the three-dimensional scene reconstruction method provided in the embodiment of the present application.
  • the embodiment of the present application also provides an electronic device, the structure of which is the same or similar to that of the electronic device shown in FIG6 , and the structure of the electronic device shown in FIG6 can be specifically referred to.
  • the difference between the electronic device provided in the present embodiment and the electronic device in the embodiment shown in FIG6 is mainly that: the functions implemented by the computer program stored in the memory executed by the processor are different.
  • the electronic device For the electronic device provided in the present embodiment, its processor executes the computer program stored in the memory, which can be used to: respond to the selection operation on the product page, determine the selected target product, and the target product has a three-dimensional product model; respond to the matching effect viewing operation, select the target image corresponding to the first spatial object to be matched with the target product; add the three-dimensional product model to the target three-dimensional model corresponding to the first spatial object to obtain the target three-dimensional model of the fused target product; project the target three-dimensional model of the fused target product onto the target image to obtain the matching effect diagram of the target product and the first spatial object; wherein the target three-dimensional model is constructed according to the steps in the three-dimensional scene reconstruction method provided in the embodiment of the present application.
  • an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed, can implement the steps in the method shown in FIG. 1b , FIG. 4a and FIG. 4e above.
  • the communication component in Figure 6 above is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices.
  • the device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G/LTE, 5G and other mobile communication networks, or a combination thereof.
  • the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.
  • the communication component also includes a near field communication (NFC) module to facilitate short-range communication.
  • the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
  • RFID radio frequency identification
  • IrDA infrared data association
  • UWB ultra-wideband
  • Bluetooth Bluetooth
  • the display in FIG. 6 includes a screen, and the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from a user.
  • the touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundary of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.
  • the power supply assembly in FIG6 provides power to various components of the device where the power supply assembly is located.
  • the power supply assembly may include The system includes a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply components are located.
  • the audio component in Figure 6 above can be configured to output and/or input audio signals.
  • the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal.
  • the received audio signal can be further stored in a memory or sent via a communication component.
  • the audio component also includes a speaker for outputting an audio signal.
  • the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
  • a computer-usable storage media including but not limited to disk storage, CD-ROM, optical storage, etc.
  • These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and/or one or more boxes in the block diagram.
  • These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and/or one or more boxes in the block diagram.
  • a computing device includes one or more processors (CPU), input/output interfaces, network interfaces, and memory.
  • processors CPU
  • input/output interfaces network interfaces
  • memory volatile and non-volatile memory
  • Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and/or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of a computer-readable medium.
  • RAM random access memory
  • ROM read-only memory
  • flash RAM flash memory
  • Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), Dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
  • PRAM phase change memory
  • SRAM static random access memory
  • DRAM Dynamic random access memory
  • RAM random access memory
  • ROM read-only memory

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Computer Graphics (AREA)
  • Software Systems (AREA)
  • Architecture (AREA)
  • Computer Hardware Design (AREA)
  • General Engineering & Computer Science (AREA)
  • Geometry (AREA)
  • Image Analysis (AREA)

Abstract

本申请实施例提供一种三维场景重建、在线家装与商品获取方法、设备及介质。在本申请实施例中,以第一空间对象对应的图像为基础,通过检测图像中存在的交界线和消影点等2D视觉信息,基于这些2D视觉信息,结合交界线之间以及消影点之间存在的约束关系,得到第一空间对应的3D结构信息,即目标三维模型,在该目标三维模型中可体现第一空间对象对应的真实场景信息,有利于提高基于该三维模型的应用效果,例如在家装场景中,可以提高基于该三维模型的家装搭配效果。而且,三维模型的稳定性更高,可解释性更强,效果更具鲁棒性;同时,结合交界线之间和消影点之间的约束关系,所生成的三维模型中模型主体结构之间的贴合程度更高。

Description

三维场景重建、在线家装与商品获取方法、设备及介质
本申请要求于2022年12月14日提交中国专利局、申请号为202211599959.2、申请名称为“三维场景重建、在线家装与商品获取方法、设备及介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及三维重建技术领域,尤其涉及一种三维场景重建、在线家装与商品获取方法、设备及介质。
背景技术
随着互联网应用的发展,用户可以在线进行各种操作,例如选购商品、在线家装等。以在线家装为例,用户可以通过家装类应用(APP)选择与房屋户型相似度最高的已有3D样板间,在3D样板间中放置各种家具或装饰物,以查看家装效果,进而根据家装效果来选购家具进行线下实际家装。
但是,现有3D样板间是通过使用3D制图技术制作出立体的房屋三维模型,缺少用户真实的房屋场景信息,面对需要与真实房屋环境进行搭配的家装需求,家装搭配效果不是很理想。为此,亟需一种能够在房屋三维模型中体现真实房屋场景信息的解决方案,用以提高在线家装的搭配效果。
发明内容
本申请的多个方面提供一种三维场景重建、在线家装与商品获取方法、设备及介质,用以在三维场景模型中体现真实三维场景信息,提高基于三维场景模型的应用效果,例如家装搭配效果。
本申请实施例提供一种三维场景重建方法,包括:获取第一空间对象对应的目标图像,所述第一空间对象是目标物理空间中的至少部分空间对象;检测所述目标图像中存在的多条交界线以及至少两个正交主方向上的消影点,所述交界线是所述第一空间对象中相邻物理主体结构之间的交线;根据所述至少两个正交主方向上的消影点,确定目标相机的相机内参数和相机坐标系下的重力方向,所述目标相机是指拍摄所述目标图像使用的相机;根据所述多条交界线和所述相机坐标系下的重力方向,重建所述第一空间对象对应的初始三维模型,所述初始三维模型中包括与所述相邻物理主体结构对应的相邻模型主体结构;根据所述多条交界线之间的约束关系、所述至少两个正交主方向上的消影点之间的约束关系以及所述相机内参数,对所述初始三维模型进行优化,以得到目标三维模型。
本申请实施例还提供一种在线家装方法,包括:响应图像上传操作,获取第一空 间对象对应的目标图像,所述第一空间对象是目标物理空间中的至少部分空间对象;响应目标家装对象在所述目标图像上的放置操作,将所述目标家装对象融合到所述第一空间对象对应的目标三维模型中,以得到融合所述目标家装对象的目标三维模型;将融合所述目标家装对象的目标三维模型投影至所述目标图像上,以得到包含所述目标家装对象的家装效果图;其中,所述目标三维模型是根据本申请实施例提供的三维场景重建方法中的步骤构建的。
本申请实施例还提供一种商品选择方法,包括:响应商品页面上的选择操作,确定被选择的目标商品,所述目标商品具有商品三维模型;响应搭配效果查看操作,选择待与所述目标商品进行搭配的第一空间对象对应的目标图像;将所述商品三维模型添加至所述第一空间对象对应的目标三维模型中,以得到融合所述目标商品的目标三维模型;将融合所述目标商品的目标三维模型投影至所述目标图像上,以得到所述目标商品与所述第一空间对象的搭配效果图;其中,所述目标三维模型是根据本申请实施例提供的三维场景重建方法中的步骤构建的。
本申请实施例还提供一种电子设备,包括:存储器和处理器;所述存储器,用于存储计算机程序;所述处理器与所述存储器耦合,用于执行所述计算机程序,以用于执行本申请实施例提供的三维场景重建方法、在线家装方法或商品选择方法中的步骤。
本申请实施例还提供一种存储有计算机程序的计算机可读存储介质,当所述计算机程序被处理器执行时,致使所述处理器能够实现本申请实施例提供的三维场景重建方法、在线家装方法或商品选择方法中的步骤。
在本申请实施例中,以第一空间对象对应的图像为基础,通过检测图像中存在的交界线和消影点等2D视觉信息,基于这些2D视觉信息,结合交界线之间以及消影点之间存在的约束关系,得到第一空间对应的3D结构信息,即目标三维模型,在该目标三维模型中可体现第一空间对象对应的真实场景信息,有利于提高基于该三维模型的应用效果,例如在家装场景中,可以提高基于该三维模型的家装搭配效果。
进一步,在本申请实施例中,直接从图像中检测交界线和消影点等2D视觉信息,之后再利用这些2D视觉信息和相关的约束关系生成3D结构信息,而不是直接从图像中回归出3D结构信息,因此,本申请实施例生成的三维模型的稳定性更高,可解释性更强,对于拍摄角度倾斜大、场景杂乱的情况,效果更具鲁棒性;同时,结合交界线之间和消影点之间的约束关系,所生成的三维模型中模型主体结构之间的贴合程度更高。
附图说明
此处所说明的附图用来提供对本申请的进一步理解,构成本申请的一部分,本申请的示意性实施例及其说明用于解释本申请,并不构成对本申请的不当限定。在附图中:
图1a为本申请示例性实施例提供的一种三维场景重建系统的结构示意图;
图1b为本申请示例性实施例提供的一种三维场景重建方法的流程示意图;
图2a为本申请示例性实施例提供的一种第一空间对象中的交界线的示意图;
图2b为本申请示例性实施例提供的一种消影点的结构示意图;
图3a为本申请示例性实施例提供的一种基于霍夫变换的交界线检测模型的结构示意图;
图3b为本申请示例性实施例提供的一种基于霍夫变换的消影点检测模型的结构示意图;
图3c为本申请示例性实施例提供的一种交界线之间的交点位置的示意图;
图3d为本申请示例性实施例提供的一种基准模型主体结构的结构示意图;
图3e为本申请示例性实施例提供的一种其它模型主体结构的结构示意图;
图4a为本申请示例性实施例提供的一种在线家装方法的流程示意图;
图4b为本申请示例性实施例提供的一种三维场景重建的过程示意图;
图4c为本申请示例性实施例提供的一种基于目标三维模型的在线家装效果示意图;
图4d为本申请示例性实施例提供的另一种基于目标三维模型的在线家装效果示意图;
图4e为本申请示例性实施例提供的一种商品选择方法的流程示意图;
图5为本申请示例性实施例提供的一种三维场景重建装置的结构示意图;
图6为本申请示例性实施例提供的一种电子设备的结构示意图。
具体实施方式
为使本申请的目的、技术方案和优点更加清楚,下面将结合本申请具体实施例及相应的附图对本申请技术方案进行清楚、完整地描述。显然,所描述的实施例仅是本申请一部分实施例,而不是全部的实施例。基于本申请中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本申请保护的范围。
针对现有在线家装场景中房屋三维模型缺少用户真实的房屋场景信息导致家装搭配效果较差的技术问题,本申请实施例提供一种三维场景重建方法,在该方法中,以第一空间对象对应的图像为基础,通过检测图像中存在的交界线和消影点等2D视觉信息,基于这些2D视觉信息,结合交界线之间以及消影点之间存在的约束关系,得到第一空间对应的3D结构信息,即目标三维模型,在该目标三维模型中可体现第一空间对象对应的真实场景信息,有利于提高基于该三维模型的应用效果,例如在家装场景中,可以提高基于该三维模型的家装搭配效果,又例如在在线购物场景中,基于该三维模型与商品的搭配效果,可提高购物体验,降低退换货概率。
以下结合附图,详细说明本申请各实施例提供的技术方案。
图1a为本申请示例性实施例提供的一种三维场景重建系统的结构示意图。如图1a所示,该系统包括:终端设备10和服务端设备20。其中,终端设备10和服务端设备20通信连接。
在本实施例中,终端设备10可以是手机、笔记本电脑或台式电脑等,服务端设备20 可以是物理服务器,云服务器或服务器阵列等,在图1a中以终端设备101是智能手机,服务端设备20是物理服务器为例进行图示,但并不限于此。
其中,终端设备10可以获取第一空间对象对应的目标图像,例如,终端设备自带相机,终端设备可以通过自带的相机在第一空间对象中采集目标图像;又例如,通过独立于终端设备10的相机采集第一空间对象对应的目标图像,由独立于终端设备的相机将采集到的目标图像提供给终端设备10。
在本实施例中,终端设备10可以将获取到的第一空间对象对应的目标图像提供给服务端设备20,由服务端设备20生成第一空间对象对应的目标三维模型。其中,服务端设备20生成第一空间对象对应的目标三维模型的过程包括:检测目标图像中存在的多条交界线以及至少两个正交主方向上的消影点,交界线是第一空间对象中相邻物理主体结构之间的交线;根据至少两个正交主方向上的消影点,确定目标相机的相机内参数和相机坐标系下的重力方向,目标相机是指拍摄目标图像使用的相机;根据多条交界线和相机坐标系下的重力方向,重建第一空间对象对应的初始三维模型,初始三维模型中包括与相邻物理主体结构对应的相邻模型主体结构;根据多条交界线之间的约束关系、至少两个正交主方向上的消影点之间的约束关系以及相机内参数,对初始三维模型进行优化,以得到目标三维模型。关于各操作的详细内容可参见后续实施例,在此暂不详述。
在目标三维模型的基础上,用户还可以在终端设备10上针对目标图像添加家装所需的目标物品,例如沙发、挂画、办公椅或办公桌等物品;终端设备10可以获取目标物品在目标图像中的位置范围信息,将目标物品的信息以及目标物品在目标图像中的位置范围信息提供给服务端设备20;服务端设备20根据目标物品在目标图像中的位置范围信息,将该目标物品的三维模型融合在第一空间对象对应的目标三维模型中,并将融合目标物品后的目标三维模型投影至目标图像上,得到包含目标物品的目标图像,将包含目标物品的目标图像返回给终端设备10,由终端设备10展示给用户,以实现在线家装搭配效果的展示。
在此说明,本申请实施例提供的三维场景重建方法不仅可以应用于图1a所示的三维场景重建系统,由终端设备和服务端设备相互配合完成,也可以由终端设备独立实施,在终端设备独立实施的情况下,原本由服务端设备执行的动作可以本实施例中的终端设备执行,其它内容与图1a所示系统中的相同或相似,在此不再赘述。关于三维场景重建的过程可参见下述方法实施例中的描述。
图1b为本申请示例性实施例提供的一种三维场景重建方法的流程示意图。如图1b所示,该方法包括:
101、获取第一空间对象对应的目标图像,第一空间对象是目标物理空间中的至少部分空间对象;
102、检测目标图像中存在的多条交界线以及至少两个正交主方向上的消影点,交界线是第一空间对象中相邻物理主体结构之间的交线;
103、根据至少两个正交主方向上的消影点,确定目标相机的相机内参数和相机坐标系下的重力方向,目标相机是指拍摄目标图像使用的相机;
104、根据多条交界线和相机坐标系下的重力方向,重建第一空间对象对应的初始三维模型,初始三维模型中包括与相邻物理主体结构对应的相邻模型主体结构;
105、根据多条交界线之间的约束关系、至少两个正交主方向上的消影点之间的约束关系以及相机内参数,对初始三维模型进行优化,以得到目标三维模型。
在本实施例中,目标物理空间指的是特定场景下的空间区域,例如,目标物理空间可以是商场、超市、机场或房屋等各种具有空间概念的区域。目标物理空间中包含至少一个空间对象,换句话说,至少一个空间对象组成了目标物理空间。例如,目标物理空间可以是物理房屋,物理房屋中包括的多个空间对象,例如厨房、卧室、客厅和卫生间等。又例如,目标物理空间可以是商场,商场中包括多个楼层,不同楼层中具有不同商户,每个商户可视为一个空间对象。其中,为了便于区分和描述,将目标物理空间中的至少部分空间对象称为第一空间对象,例如,第一空间对象可以是目标物理空间中的一个空间对象,也可以是目标物理空间中的多个空间对象,如,2个或3个等;又例如,第一空间对象可以是某个空间对象中的局部空间。以目标物理空间是物理房屋为例,第一空间对象可以是整个物理房屋,也可以是其中的主卧、次卧或客厅等某个独立的空间对象,还可以是主卧、次卧或客厅等某个独立空间对象中的局部空间,还可以是同时包含多个独立空间对象中的局部空间,如同时包含部分客厅空间和部分阳台空间。
在本实施例中,可以获取第一空间对象对应的目标图像,例如,通过相机采集第一空间对象中的图像,从相机获取第一空间对象对应的目标图像,或者,终端设备上安装有相机,通过终端设备上的相机采集第一空间对象对应的目标图像。其中,目标图像是第一空间对象的环境图像,包含第一空间对象的真实环境信息。
在本实施例中,目标图像中存在多条交界线,交界线是第一空间对象中相邻物理主体结构之间的交线,其中,物理主体结构可以包含但不限于:墙面、天花板或地面等。其中,可以将墙面之间的交线称为墙线,将墙面和地面之间的交线称为地线,将墙面和天花板的之间的交线称为天花板线。如图2a所示,以第一空间对象是客厅的局部空间为例进行图示,第一空间对象中包括:天花板、地面和墙面,提取的地线为C1、提取的墙线为C2、提取的天花板线为C3。图2a中的地线C1、墙线C2以及天花板线C3均属于本申请实施例中的交界线的示例,但并不限于此。
在本实施例中,第一空间对象可以符合曼哈顿结构,其中,曼哈顿结构包含三个相互正交的主方向,例如,三个正交主方向分别是重力方向、主体墙面正对的方向以及侧边墙面正对的方向。其中,重力方向可以认为是与地面或者天花板垂直的方向;主体墙面是指与目标相机的光轴夹角最小的墙面,目标相机是拍摄目标图像使用的相机,目标相机的光轴是指通过镜头中心的线,主体墙面正对的方向就是垂直于主体墙面的方向,即主体墙面的法向量的方向;侧边墙面是相对于主体墙面而言的,主要是指与主体墙面相邻的墙面, 侧边墙面正对的方向是指垂直于侧边墙面的方向,通常来讲,在目标图像中可以有多个侧边墙面,例如,2个,多个侧边墙面是相对的或平行的,多个侧边墙面对应的方向是一样的,即侧边墙面的法向量方向。在此说明,目标图像中可能同时包含与三个正交主方向对应的物理主体结构,例如同时包含主体墙面、侧边墙面以及地面或天花板;也可能仅包含与两个正交主方向对应的物理主体结构,例如仅包含主体墙面和地面,或者仅包含侧边墙面和天花板,或者仅包含主体墙面、天花板和地面等。至于目标图像中是包含与三个正交主方向对应的物理主体结构,还是包含与两个正交主方向对应的物理主体结构,具体可视目标图像的拍摄视角而定。简单来说,目标图像中至少存在与两个正交主方向对应的物理主体结构。
在本实施例中,三维空间中的平行直线经过透视变化后在该三维空间对应的二维图像中相交于一点,该点即为消影点(Vanishing Points,VP),消影点的方向就代表相交得到该消影点的三维空间中的平行直线的方向。其中,在目标图像中存在多个消影点,这些消影点中包括位于至少两个正交主方向上的消影点;对于位于某个正交主方向上的消影点,相交得到该消影点的三维空间中的平行直线可以垂直于三维空间中的与该正交主方向对应的物理主体结构,而三维空间中的物理主体结构符合曼哈顿结构,每个物理主体结构对应有主方向,也就是说,目标图像中位于某个正交主方向上的消影点的方向对应于三维空间中的一个主方向,即位于某个正交主方向上的消影点的方向代表某一物理主体结构的法向量的方向。在目标图像中至少存在与两个正交主方向对应的物理主体结构的情况下,也就意味着目标图像中包含至少两个正交主方向上的消影点。如图2b所示,其中线1与右侧墙面(即图2b中挂画的墙面)平行,线1相交得到的消影点D1的方向代表左侧墙面(即图2b中有窗户的墙面)的法向量方向;线2与左侧墙面(即图2b中有窗户的墙面)平行,线2相交得到的消影点D2的方向代表右侧墙面(即图2b中挂画的墙面)的法向量方向;线3与地面(或天花板)平行,线3相交得到的消影点D3的方向代表地面(或天花板)的法向量方向,即与重力方向平行的方向。
在本实施例中,可以检测目标图像中存在的多条交界线以及至少两个正交主方向上的消影点。其中,在检测交界线和消影点时,无需对目标图像进行分割,而是直接在目标图像的基础上进行检测,不会受到图像分割带来的干扰,因此检测到的交界线和消影点的准确度和精度均较高。
在本实施例中,目标图像中消影点的位置由目标相机的内参数和消影点对应的3D方向(即消影点所在的主方向)决定,目标相机的内参数是与目标相机自身特性相关的参数,例如,目标相机的焦距、像素大小以及主点位置(即相机光心位置)等。反过来讲,在已知目标图像上至少两个正交主方向上的消影点的情况下,可以根据至少两个正交主方向上的消影点,确定目标相机的相机内参数以及相机坐标系下的重力方向。具体地,可以根据至少两个正交主方向上的消影点,确定目标相机的相机内参数,根据相机内参数和重力方向上的消影点,确定相机坐标系下的重力方向。
在得到相机坐标系下的重力方向和目标图像中存在的多条交界线的基础上,可以根据目标图像中存在的多条交界线以及相机坐标系下的重力方向,重建第一空间对象对应的初始三维模型,初始三维模型中包括与相邻物理主体结构对应的相邻模型主体结构。物理主体结构是构成第一空间对象的主体结构,模型主体结构是构成初始三维模型的主体结构;模型主体结构是物理主体结构在初始三维模型中的体现。其中,第一空间对象对应的初始三维模型是根据各个物理主体结构的平面方程构建的,例如可以根据房屋中包含的各个墙面、地面或天花板的平面方程构建出对应的房屋三维模型。其中,物理主体结构的平面方程可以唯一性的表示该物理主体结构,可选地,可以通过该物理主体结构的法向量和到相机光心的距离进行表示,但并不限于此。在得到初始三维模型之后,该初始三维模型中的模型主体结构也具有平面方程,理论上模型主体结构的平面方程与其对应的物理主体结构的平面方向应该相同。
其中,根据目标图像中存在的多条交界线以及相机坐标系下的重力方向,重建第一空间对象对应的初始三维模型的大体原理为:根据相机坐标系下的重力方向,构建三维空间中的地平面,该地平面与相机坐标系下的重力方向是垂直的;基于目标图像中存在的多条交界线,可以确定第一空间对象在地平面上形成的边界以及边界上存在的模型主体结构;进一步可以假定相机光心与地面的高度,即相机高度,基于假定的相机高度和预先假设的缩放比例,确定模型主体结构(如天花板)的高度,从而得到第一空间对象对应的初始三维模型。
在本实施例中,第一空间对象对应的初始三维模型是基于目标图像中的交界线进行构建的,当交界线存在误差时,会导致初始三维模型中的透视关系不正确,如,相邻墙面之间不存在垂直关系,进一步,假定相机高度和缩放比例,也会导致生成的墙面高度并不准确。基于此,在本实施例中,在得到初始三维模型之后,还会根据多条交界线之间的约束关系,以及至少两个正交主方向上的消影点之间的约束关系,结合相机内参数,对初始三维模型进行优化,以得到目标三维模型。其中,多条交界线之间的约束关系为:形成交界线的相邻物理主体结构之间相互垂直,该约束关系也可称为曼哈顿约束,即第一空间对象中存在相邻关系的墙面、地面以及天花板之间相互垂直。至少两个正交主方向上的消影点之间的约束关系为:任意两个正交主方向上消影点对应的物理主体结构之间相互垂直,同一消影点对应的物理主体结构的法向量方向与该消影点表示的平行直线的方向相互平行。基于这些约束关系,可以对每个初始三维模型中的模型主体结构对应的法向量方向和到相机光心的距离进行优化,以得到位置关系、高度等均与物理主体结构高度吻合的模型主体结构。
在本申请实施例中,以第一空间对象对应的图像为基础,通过检测图像中存在的交界线和消影点等2D视觉信息,基于这些2D视觉信息,结合交界线之间以及消影点之间存在的约束关系,得到第一空间对应的3D结构信息,即目标三维模型,在该目标三维模型中可体现第一空间对象对应的真实场景信息,有利于提高基于该三维模型的应用效果,例如 在家装场景中,可以提高基于该三维模型的家装搭配效果。进一步,在本实施例中,直接从图像中检测交界线和消影点等2D视觉信息,之后再利用这些2D视觉信息和相关的约束关系生成3D结构信息,而不是直接从图像中回归出3D结构信息,因此,本实施例生成的目标三维模型的稳定性更高,可解释性更强,对于拍摄角度倾斜大、场景杂乱的情况,效果更具鲁棒性;同时,结合交界线之间和消影点之间的约束关系,所生成的目标三维模型中模型主体结构之间的贴合程度更高。
在本申请一可选实施例中,在上述检测交界线和消影点时,可以基于霍夫变换检测目标图像中存在的多条交界线以及至少两个正交主方向上的消影点。其中,霍夫变换是一种特征提取(feature extraction)技术,被广泛应用在图像分析(image analysis)、计算机视觉(computer vision)以及数位影像处理(digital image processing)等领域,在本实施例中,利用霍夫变换可以检测目标图像中存在的多条交界线,如地线、墙线以及天花板线等,同时,基于霍夫变换还可以检测目标图像中包含的至少两个正交主方向上的消影点。
进一步可选地,在基于霍夫变换检测目标图像中存在的多条交界线以及至少两个正交主方向上的消影点时,可以基于霍夫变换的深度神经网络模型检测目标图像中存在的多条交界线以及至少两个正交主方向上的消影点。进一步可选地,基于霍夫变换的深度神经网络模型包括:基于霍夫变换的交界线检测模型和基于霍夫变换的消影点检测模型。基于霍夫变换的交界线检测模型可以是任何能够从图像中进行交界线的检测的神经网络模型,基于霍夫变换的消影点检测模型可以是任何能够从图像中进行消影点检测的神经网络模型。基于此,一方面可以将目标图像输入基于霍夫变换的交界线检测模型进行交界线的检测,以获取目标图像中存在的多条交界线;一方面可以将目标图像输入基于霍夫变换的消影点检测模型进行消影点检测,以获取目标图像中至少两个正交主方向上的消影点。下面分别针对进行交界线检测和消影点检测的神经网络模型的架构和检测过程进行示例性说明。
基于霍夫变换的交界线检测:
在一可选实施例中,如图3a所示,基于霍夫变换的交界线检测模型中包含:第一特征提取网络,第一特征提取网络是融合有跳跃连接(Skip Connections)和多尺度特征的神经网络。其中,多尺度特征是指针对目标图像提取多个不同尺度的特征图的技术,不同尺度的特征图包含不同的特征信息,特征图的尺寸越小,深度越大,属于深层网络特征;反之,特征图的尺寸越大,深度越小,属于浅层网络特征。深层网络特征的感受野(Receptive Field)比较大,特征图的分辨率低,语义信息表征能力强,几何细节信息的表征能力较弱;浅层网络特征的感受野比较小,特征图的分辨率高,语义信息表征能力弱,几何细节信息表征能力强。其中,感受野是卷积神经网络每一层输出的特征图(feature map)上的像素点在输入图片上映射的区域大小。其中,跳跃连接是一种简单而有效的融合深层网络特征和浅层网络特征的操作,跳跃连接可以将浅层网络特征与同尺度的深层网络特征进行元素级的相加。基于此,可以将目标图像输入基于霍夫变换的交界线检测模型中,利用融合跳跃连接和多尺度特征的第一特征提取网络对目标图像进行特征提取,得到多个尺度的特征图。 为了便于描述和区分,将这里的第一特征提取网络最终输出的多个尺度的特征图称为多个尺度的第一目标特征图。
其中,多个尺度的第一目标特征图属于图像像素空间中的特征,为了便于进行交界线提取,可以将图像像素空间中的特征转换为霍夫空间中的特征,对霍夫空间中的特征进行二分类,即区分霍夫空间中的特征是否为图像像素空间中的交界线对应的特征。基于此,如图3a所示,基于霍夫变换的交界线检测模型还包括:霍夫变换网络。基于该霍夫变换网络,可以利用基于极坐标的霍夫变换对多个尺度的第一目标特征图进行霍夫变换,以将目标图像中存在的直线映射为霍夫空间中的点,在霍夫空间中对转换得到的点进行二分类,例如,区分霍夫空间中哪些点是图像像素空间中的交界线形成的点,哪些点不是图像像素空间中的交界线形成的点。
其中,对任一尺度的第一目标特征图而言,都属于图像空间,即其采用的是像素坐标系,像素坐标系是以目标图像中心为原点,水平方向是x轴,垂直方向是y轴的坐标系,其单位长度是像素。与像素坐标系不同,霍夫空间使用的是极坐标系,在极坐标系下,目标图像中的一条线可以通过(ρ,θ)进行表示的,ρ表示该条直线的末端到像素坐标系的原点的距离,θ表示该条直线与像素坐标系的x轴的夹角,霍夫空间下的第一目标特征图的坐标为(ρ,θ),因此通过极坐标表示目标图像的一条线就会变成一个点。其中,目标图像中存在的直线都会被映射为霍夫空间中的点,这些点包括由交界线映射而成的点,即与交界线特征匹配的点,也包括由非交界线映射而成的点,即不与交界线特征匹配的点,其中,为了便于区分和描述,将霍夫空间中与交界线特征匹配的点称为目标点。基于此,可以在霍夫空间中选择与交界线特征匹配的多个目标点,并将多个目标点重新映射到图像空间中,以得到目标图像中存在的多条交界线。如图3a所示,可以通过反向霍夫变换(Reverse Hough transform,RHT),将多个目标点重新映射到图像空间中,以得到目标图像中存在的多条交界线。
进一步可选地,如图3a所示,第一目标特征提取网络包括多个下采样模块、多个上采样模块以及跳跃连接模块。在一可选实施方式中,对目标图像进行特征提取,得到初始特征图,将该初始特征图作为最大尺度的第一中间特征图,利用下采样模块对最大尺度的第一中间特征图进行N次下采样(undersampling)处理,得到其它尺度的第一中间特征图,其中,N为正整数,例如,N可以是2、3、4或6等;将最小尺度的第一中间特征图作为最小尺度的第一目标特征图,利用上采样模块对最小尺度的第一目标特征图进行N次上采样(upsampling)处理,并在每次上采样处理中与下采样处理得到的相同尺度的第一中间特征图进行跳跃连接,以得到其它尺度的第一目标特征图。
如图3a所示,假设,最大尺度的第一中间特征图的尺度是512*512*3,N=3,经过3次下采样处理可以得到尺度为256*256*3、128*128*3和64*64*3的第一中间特征图;然后,将尺度为64*64*3的第一中间特征图作为最小尺度的第一目标特征图,对第一目标特征图依次进行3次上采样处理,具体地,经过1次上采样处理可以得到尺度为128*128*3 的第二中间特征图,将该第二中间特征图与尺度为128*128*3的第一中间特征图进行跳跃连接,得到尺度为128*128*3的第一目标特征图;继续对尺度为128*128*3的第一目标特征图进行1次上采样处理得到尺寸为256*256*3的第二中间特征图,将尺度为256*256*3的第二中间特征图和256*256*3的第一中间特征图进行跳跃连接,得到尺度为256*256*3的第一目标特征图;继续对度为256*256*3的第一目标特征图进行1次上采样处理得到尺度为512*512*3的第二中间特征图,将尺度为512*512*3的第二中间特征图与尺度为512*512*3的第一中间特征图进行跳跃连接,得到尺度为512*512*3的第一目标特征图。
进一步如图3a所示,霍夫变换网络包括:霍夫变换模块和特征融合模块。基于此,一种利用基于极坐标的霍夫变换对多个尺度的第一目标特征图进行霍夫变换,以将目标图像中存在的直线映射为霍夫空间中的点的实施方式,包括:利用霍夫变换模块对多个尺度的第一目标特征图X进行霍夫变换,以得到基于极坐标的霍夫空间中多个尺度的第二目标特征图Y,其中,可以针对每个尺度的第一目标特征图X分别进行霍夫变换,得到该尺度下基于极坐标的霍夫空间中的第二目标特征图Y;利用特征融合模块对多个尺度的第二目标特征图进行尺度变换,以得到尺度相同的多个特征图,并对尺度相同的多个特征图进行拼接,得到第三目标特征图Z;对第三目标特征图进行卷积降维,以得到霍夫空间中的二维图像,二维图像使用的是极坐标系,二维图像中包含多个点,每个点的坐标为(ρ,θ),每个点与目标图像中存在的直线对应。
其中,本实施例并不对第三目标特征图的尺度进行限定,例如,第三目标特征图的尺度可以是第二目标特征图对应的多个尺度中的任意一个尺度,优先地,第三目标特征图的尺度是第二目标特征图对应的多个尺度中的最大尺度;基于此,在对多个尺度的第二目标特征图进行尺度变换时,可以对非最大尺度的第二目标特征度进行上采样处理,以将其尺寸变为最大尺度。
如图3a所示,展示了第一目标特征图、第二目标特征图、第三目标特征图以及霍夫空间中的二维图像,另外,在本实施例中,霍夫变换是在深度特征图的基础上进行的,故称为深度霍夫变换(Depth Hough transform,DHT)。在图3a中,特征融合模块通过带圈的c进行表示,该特征融合模块进一步包括上采样单元(upsample)和拼接单元(Concat),上采样单元用于对多个尺度的第二目标特征图进行尺度变换,以得到尺度相同的多个特征图;拼接单元(Concat)用于对尺度相同的多个特征图进行拼接,得到第三目标特征图Z。
需要说明的是,在使用基于霍夫变换的交界线检测模型之前,可以预先进行模型训练得到基于霍夫变换的交界线检测模型。在本实施例中,创新性地将深度霍夫变换与神经网络模型的结合应用于图像中交界线的检测上,具体地,获取大量空间对象(如室内场景)的图片,对图片中存在的墙线、地线以及天花板线进行标注得到样本数据集,利用该样本数据集训练基于霍夫变换的基础神经网络,该基础神经网络包括第一特征提取网络和霍夫变换网络,第一特征提取网络包括上采样模块、下采样模块以及跳跃连接模块,用以对像素空间中的样本图像进行特征提取得到多尺度的样本特征图,霍夫变换网络包括霍夫变换 模块和特征融合模块,用以通过霍夫变换将多尺度的样本特征图变换到霍夫空间中并得到霍夫空间中的样本图像,接着,在霍夫空间中利用二分类方法对霍夫空间中的样本图像进行分类,具体是指将该样本图像中的点划分为与交界线对应的第一类点和与其它直线对应的第二类点;然后,根据标注结果和分类结果生成损失函数,可选地,可以采用交叉熵损失(Binary CrossEntropy Loss,BCELoss)函数,在损失函数不达标的情况下继续进行训练,直至损失函数达标或者模型训练时间或次数达到设定的时间或次数为止,得到基于霍夫变换的交界线检测模型。需要说明的是,在图3a中,展示了基于标注结果和分类结果进行损失函数计算的部分,这部分用于模型训练(Training only)阶段;另外,在图3a中还展示了通过反向霍夫变换将多个目标点重新映射到图像空间中,以得到目标图像中存在的多条交界线的过程,该过程用于模型推理阶段。
基于霍夫变换的消影点检测:
在一可选实施例中,如图3b所示,基于霍夫变换的消影点检测模型包括第二特征提取网络,第二特征提取网络可以是任何能够进行特征提取的网络,例如,图像分类中的骨干(backbone)网络,例如,UNet,堆积沙漏(Stacked Hourglass)模型等,其中,UNet是全卷积网络(Fully Convolutional Networks,FCN)的一种变体,其网络结构是对称的,形似英文字母U,故而被称为UNet;Stacked Hourglass模型是一种利用多尺度特征来识别姿态的网络结构。其中,可以将目标图像输入基于霍夫变换的消影点检测模型中,利用第二特征提取网络对目标图像进行特征提取,得到第四目标特征图;如图3b所示,以目标图像是[512x512]x3为例,其中,[512x512]为目标图像的尺度,3为通道数;以第四目标特征图是[128x128]x3为例进行图示,但并不限于此。其中,可以将第四目标特征图映射至高斯球面空间,例如,将目标图像中的直线投影到以目标相机中心为球心的高斯球面空间,在高斯球面空间中,目标图像中相交于同一个点的直线会在一个点上响应最强,则可以根据高斯球面空间中各个点的响应值和角度来获取位于至少两个正交主方向上的消影点。进一步,如图3b所示,基于霍夫变换的消影点检测模型还包括基于高斯球面的霍夫变换网络。基于此,第四目标特征图被送入基于高斯球面的霍夫变换网络中,在该网络中,利用基于高斯球面的霍夫变换对第四目标特征图进行霍夫变换,以将目标图像中存在的直线映射为高斯球面空间中的点;然后,在高斯球面空间中选择概率值和角度值满足要求的至少两个点作为至少两个正交主方向上的消影点,并重新映射到图像空间中,以得到目标图像中至少两个正交主方向上的消影点。其中,高斯球面空间中包括多个点,每个点是由目标图像中的直线得到的,每个点具有亮度值和角度值两个属性,亮度值表示对应点是由直接相交得到的点(即消影点)的概率值,对应点在高斯球面上的亮度越大,表示该点是消影点的概率就越大;角度值表示该点对应的直线在图像空间中与x轴之间的夹角,根据该角度值可以确定高斯球面上各点对应的直线之间是否相互垂直。基于此,可以从在高斯球面空间中选择概率值和角度值满足要求的至少两个点作为至少两个正交主方向上的消影点。
进一步可选地,基于高斯球面的霍夫变换网络包括霍夫变换模块和高斯球面变换模块。 基于此,在对目标图像进行特征提取得到第四目标特征图后,可以将第四目标特征图依次变换到霍夫空间和高斯球面空间。在霍夫空间中,目标图像中的直线会被映射为一个点(简称霍夫点),但无法直接区分霍夫点是否是由多条直线相交的点;进一步,将霍夫点映射到高斯球面空间中,在高斯球面空间,每个点(简称高斯点)的响应值(即亮度值)会因为高斯点对应直线数量的不同而有所不同,对于由一条直线形成的高斯点的响应值小于由多条直线相交而形成的高斯点的响应值,因此可以选择响应值最强的若干个高斯点作为消影点。具体地,将第四目标特征图输入霍夫变换模块,在该模块内部,对第四目标特征图进行霍夫变换,以得到基于极坐标的霍夫空间中的第五目标特征图,如图3b所示,以第五目标特征图的维度是[184x180]x128为例进行图示,但并不限于此,其中,184是ρ所代表的距离维度上的最大值,180是θ角度维度的180°,128是通道数,图3b中,HT表示霍夫空间,(ρ,θ)是霍夫空间中的坐标。进一步,如图3b所示,在霍夫空间中对第五目标特征图进行霍夫卷积,以得到第六目标特征图,其中,第六目标特征图相对第五目标特征图的维度不变,在图3b中,霍夫卷积表示为HT Conv。接着,如图3b所示,将第六目标特征图输入高斯球面变换模块,在该模块内部,对第六目标特征图进行球面高斯球面变换,以得到高斯球面空间中的第七目标特征图,在图3b中,(α,β)为高斯球面空间中的变量,以第七目标特征图的维度是[32768]x128为例进行图示,128为通道数,[32768]表示高斯球面被离散为32768个点,通过该球面高斯球面变换,可以确定霍夫空间中的点对应于高斯球面空间中的离散点;在高斯球面空间中对第七目标特征图进行球面卷积,得到第八目标特征图,第八目标特征图中包括多个高斯点,每个高斯点与目标图像中存在的直线对应,在图3b中以球面卷积为Shperical Conv,第八目标特征图的维度是[32768]x128为例进行图示,但并不限于此。
可选地,根据第八目标特征图中各个高斯点的概率值,从多个高斯点中选择概率值满足设定要求的点作为消影点,例如,将概率值超过设定概率阈值的点作为消影点,概率阈值可以是80%、90%或95%等;从消影点中选择角度大于设定角度的至少两个消影点,作为至少两个正交主方向上的消影点。这里说明,可以从消影点中选择角度最大的两个或三个消影点,作为至少两个正交主方向上的消影点,在图3b所示的高斯球面中,以选择3个消影点为例进行图示。
在本实施例中,上述根据至少两个正交主方向上的消影点,确定目标相机的相机内参数和相机坐标系下的重力方向,包括:在至少两个正交主方向上的消影点中包括至少两个有限消影点的情况下,从至少两个有限消影点中选择两个目标消影点;根据两个目标消影点与相机内参数的约束关系,确定目标相机的相机内参数;根据相机内参数,将至少两个正交主方向中位于重力方向上的消影点转换到相机坐标系中,以得到相机坐标系下的重力方向。
其中,目标相机采集目标图像时,三维空间中两条平行直线经过透视变换会在目标图像中相交形成消影点,这些消影点可能相交于无穷远处,也可能未相较于无穷远处。在本 实施例中,对无穷远处进行了一定的定义,例如,在消影点的像素坐标到像素坐标系的原点的距离大于目标图像的对角线长度的设定倍数(如10倍)时,认为消影点相交于无穷远处。为了便于区分和描述,将相交于无穷远处的消影点称为无限消影点,将未相交于无穷远处的消影点称为有限消影点。基于此,本申请实施例中的至少两个正交主方向上的消影点中,也可能包含无限消影点,也可能包含有限消影点。具体地,在至少两个正交主方向上的消影点中包括至少两个有限消影点的情况下,从至少两个有限消影点中选择两个消影点作为目标消影点。其中,从至少两个有限消影点中选择两个目标消影点的实施方式并不限定,例如,可以从至少两个有限消影点中随机选择两个消影点作为目标消影点;又例如,可以从至少两个有限消影点中选择距离相机光心最近的两个消影点作为目标消影点。其中,无论选择的目标消影点是哪两个,这两个目标消影点中可能包含位于重力方向上的消影点,也可能不包含位于重力方向上的消影点,对此不做限定。
其中,在选择出两个目标消影点之后,可以根据两个目标消影点与相机内参数的约束关系,确定目标相机的相机内参数。其中,两个目标消影点与相机内参数的约束关系为:(K-1 *VP1)·(K-1 *VP2)=0,该约束关系表示两个目标消影点在三维空间中对应的直线方向的点积运算的结果为0,简单来说就是两个目标消影点在三维空间中对应的直线方向相互垂直,VP1:表示第一个目标消影点的三维向量,该三维向量是第一个目标消影点的图像坐标后面增加一个维度形成的三维向量,如可以在图像坐标后面增加1形成第一个目标消影点的三维向量;K:表示目标相机的内参数,是3X3的矩阵,包括焦距和主点位置,主点位置就是相机光心位置,例如,可以取图像坐标为(0.5,0.5)的位置作为相机光心位置,但不限于此,其中,焦距是待求解的,主点位置是已知的;K-1:表示相机内参数的逆;VP2:是第二个目标消影点的三维向量,该三维向量是第二个目标消影点的图像坐标后面增加一个维度形成的三维向量,例如可以在图像坐标后面增加1形成第二个目标消影点的三维的向量。
其中,VP1和VP2为已知,则可以通过两个目标消影点与相机内参数的约束关系,求解得到相机内参数。接着,可以利用相机内参数以及位于重力方向上的消影点,计算相机坐标系下的重力方向diry=K-1 *VPy,其中,y方向表示重力方向,VPy表示图像空间中位于重力方向上的消影点的三维的向量,diry表示相机坐标系下的重力方向。
进一步,消影点约束了对应主方向的透视关系,假定某一物理主体结构(如墙面)的法线方向为Ni,Ni为待求量,该物理主体结构对应的消影点为VPi,则需要满足该物理主体结构的法线方向和相交得到该消影点的平行直线所属的主方向互相平行的约束,即Ni·(K-1 *VPi)=1。
在一可选实施例中,在重建第一空间对象对应的初始三维模型时,可以确定第一空间对象中包含的基准物理主体结构,基准物理主体结构可以是地面,也可以是天花板,具体视目标图像中包含的物理主体结构而定,例如,若目标图像中包含地面,则将地面作为基准物理主体结构,若目标图像中不包含地面,而包含天花板,则将天花板作为基准物理主 体结构。在确定基准物理主体结构的基础上,根据相机坐标系下的重力方向,在相机坐标系下构建三维空间的基准平面,该基准平面与基准物理主体结构对应;从多条交界线中识别与基准物理主体结构相交的多条基准交界线,若基准物理主体结构为地面,则多条基准交界线可以是多个墙面与地面相交得到的多条地线,若基准物理主体结构为天花板,则多条基准交界线可以是多个墙面与天花板相交得到的多条天花板线;根据多条基准交界线之间的交点位置,在基准平面上构建与基准物理主体结构对应的基准模型主体结构;根据预设的相机光心到基准平面的高度和多条基准交界线之间的交点位置,在基准模型主体结构上构建与其它物理主体结构对应的其它模型主体结构,以得到第一空间对象对应的初始三维模型,其中,若基准物理主体结构为地面,则其它物理主体结构可以是墙面、天花板等。如图3e所示为基准模型主体结构以及其它模型主体结构的示例性展示。示例性地,在图3e中,基准模型主体结构是指初始三维模型中的地面,其它模型主体结构是指初始三维模型中的墙面。
可选地,上述根据多条基准交界线之间的交点位置,在基准平面上构建与基准物理主体结构对应的基准模型主体结构的一种实施方式包括:从多条基准交界线中选择有效基准交界线,有效基准交界线的数量为多条,根据有效基准交界线与图像坐标系中x轴的夹角对有效基准交界线进行排序,以得到有效基准交界线之间的相邻关系;根据有效基准交界线之间的相邻关系,确定有效基准交界线之间的交点位置,并将交点位置划分为与目标图像的边界相交的第一交点位置以及未与目标图像的边界相交的第二交点位置,在图3c中示例性展示了第一交点位置和第二交点位置,但并不限于此;根据第一交点位置和第二交点位置,在基准平面上绘制基准物理主体结构对应的模型边界,以得到基准模型主体结构,如图3d所示。在图3d中,网格区域为基准平面,网格区域上的白色区域为基准物理主体结构对应的模型边界所形成的基准模型主体结构,接续于图3c所示的交点位置,图3d中的基准模型主体结构为白色区域形成的地面。
进一步可选地,本申请实施例对从多条基准交界线中选择有效基准交界线的方式并不限定,下面进行举例说明。
示例B1:根据多条基准交界线与图像坐标系中x轴的夹角,从多条基准交界线中选择第一基准交界线;例如,从多条基准交界线中选择与图像坐标系中x轴夹角小于设定角度阈值的基准交界线作为第一基准交界线,角度阈值可以是10度、15度或20度等,进一步,在夹角小于该角度阈值的交界线为多条的情况下,可以从夹角小于该角度阈值的交界线中选择夹角最小的交界线作为第一基准交界线,或者,也可以从夹角小于该角度阈值的交界线中随机选择一条交界线作为第一基准交界线;再根据第一基准交界线剔除一些误检测的基准交界线,具体地,根据其它基准交界线与第一基准交界线之间的夹角,将夹角小于设定夹角阈值的基准交界线剔除,角度阈值可以是3度、5度或10度等;将未被剔除的基准交界线和第一基准交界线作为有效基准交界线。
示例B2:根据多条基准交界线的长度,从多条基准交界线中选择第一基准交界线;例 如,从条基准交界线中选择长度大于设定长度阈值的基准交界线作为第一基准交界线,长度阈值并不限定,具体视目标图像的大小而定;再根据第一基准交界线剔除一些误检测的基准交界线,具体地,根据其它基准交界线与第一基准交界线之间的夹角,将夹角小于设定夹角阈值的基准交界线剔除,将未被剔除的基准交界线和第一基准交界线作为有效基准交界线。
示例B3:根据多条基准交界线与图像坐标系中x轴的夹角和多条基准交界线的长度,从多条基准交界线中选择第一基准交界线;例如,从多条基准交界线中选择与图像坐标系中x轴夹角小于设定角度阈值的候选基准交界线,再从候选基准交界线中选择长度大于设定长度阈值的基准交界线作为第一基准交界线;根据其它基准交界线与第一基准交界线之间的夹角,将夹角小于设定夹角阈值的基准交界线剔除,将未被剔除的基准交界线和第一基准交界线作为有效基准交界线。
可选地,有效基准交界线之间的交点位置划分为第一交点位置和第二交点位置的情况下,根据预设的相机光心到基准平面的高度和多条基准交界线之间的交点位置,在基准模型主体结构上构建与其它物理主体结构对应的其它模型主体结构,以得到第一空间对象对应的初始三维模型的一种实施方式,包括:根据第二交点位置,确定基准模型主体结构上相交于第二交点位置的相邻模型边界;根据预设的相机光心到基准平面的高度和预设的缩放比例,确定其它模型主体结构的初始高度;根据其它模型主体结构的初始高度,在相交于第二交点位置的相邻模型边界上构建其它模型主体结构,以得到第一空间对象对应的初始三维模型。例如,以基准物理主体结构为地面,则在相交于第二交点位置的相邻模型边界上构建其它模型主体结构具体为,在相交于第二交点位置的相邻模型边界上构建墙面,进一步,在相邻墙面上方补充天花板,以得到第一空间对象对应的初始三维模型。
在一可选实施例中,上述根据多条交界线之间的约束关系、至少两个正交主方向上的消影点之间的约束关系以及相机内参数,对初始三维模型进行优化,以得到目标三维模型的一种实施方式包括:根据多条交界线之间的约束关系、至少两个正交主方向上的消影点之间的约束关系以及相机内参数,构建以各模型主体结构的位置参数和/或高度参数为优化变量的优化函数;采用最小二乘算法对优化函数进行求解,以得到各模型主体结构的优化后的位置参数和/或高度参数,根据优化后的位置参数和/或高度参数调整各模型主体结构的位置,以得到目标三维模型。
在本实施例中,基准模型主体结构和其它模型主体结构可以具有位置参数和/或高度参数,其中,位置参数包括模型主体结构的法向量和模型主体结构到相机光心的距离,这两个信息可以唯一确定模型主体结构在初始三维模型中的位置以及与其它模型主体结构之间的相对位置;高度参数为模型主体结构到基准平面的高度信息,例如,天花板的高度信息、墙面的高度信息以及地面的高度信息,其中,若以地面为基准平面,地面的高度信息为0。
可选地,在构建以各模型主体结构的位置参数和/或高度参数为优化变量的优化函数时, 可以根据多条交界线之间的约束关系、至少两个正交主方向上的消影点之间的约束关系以及相机内参数,构建第一类优化项、第二类优化项和第三类优化项中的至少一种优化项,根据至少一种优化项生成优化函数。优选地,可以同时根据第一类优化项、第二类优化项和第三类优化项,生成优化函数。
其中,针对每一个模型主体结构,根据相机内参数和与该模型主体结构垂直的消影点,生成该模型主体结构在相机坐标系下的参考法向量,其中,理论上参考法向量与该模型主体结构的法向量应该是平行的,因此,以该模型主体结构的法向量与参考法向量的点积为1作为优化目标可以构建第一类优化项。
其中,针对任意相邻模型主体结构,任意相邻模型主体结构的法向量应当是垂直的,例如,地面与墙面之间垂直,墙面与天花板之间垂直,相邻墙面与墙面之间垂直,因此,以任意相邻模型主体结构之间的法向量的点积为0作为优化目标可以构建第二类优化项。
其中,针对任意相邻模型主体结构,根据相机内参数、任意相邻模型主体结构的法向量以及任意相邻模型主体结构到相机光心的距离,生成任意相邻模型主体结构在图像坐标系下的交界线;另外,基于霍夫变换检测目标图像中对应相邻物理主体结构之间的交界线,理论上讲,任意相邻模型主体结构在图像坐标系下的交界线(这里的交界线可以视为是初始三维模型中任意相邻模型主体结构之间的交界线在图像空间中的投影)与从目标图像中检测到对应相邻物理主体结构之间的交界线(这里的交界线可以视为是第一空间对象中相邻物理主体结构之间的交界线在图像空间中的投影)应当相同,因此,以任意相邻模型主体结构在图像坐标系下的交界线和目标图像中对应相邻物理主体结构之间的交界线相同为优化目标可以构建第三类优化项。
下面以结合模型主体结构对应的平面方程对上述优化项的构建过程进行示例性说明。其中,待优化的模型主体结构为墙面、天花板以及地面,假设待优化的墙面方程为{ni,di},由消影点估计出的第i面墙的法向量为从图像中检测出的第i面墙与地面的交界线记做从图像中检测出的第i面墙与第j面墙的交界线记做从图像中检测出的第i面墙和天花板的交界线记做其中,||·||表示范数,ni是第i个墙面的法向量,属于待优化量,该法向量具有初始值(即初始三维模型中确定的值),di是第i个墙面到相机光心的距离,天花板的高度dceiling是构建初始三维模型时假设的值,需要进一步优化,因此,优化变量为墙面的位置参数(如墙面方程),以及天花板的高度参数dceiling。一种示例性的优化函数可表示如下:
其中,T表示图像中出现的墙面集合。采用最小二乘算法对优化函数进行求解,是以优化函数中的所有项之和最小为优化目标,得到墙面的位置参数(如墙面方程)以及天花 板的高度参数。优化函数中包含六项,其中,第一项属于第一类优化项,第二项和第三项属于第二类优化项,第四项、第五项以及第六项属于第三类优化项。下面进行详细说明:
第一项,第i墙面的法向量ni,以及根据消影点得到的该墙面的法向量二者的方向应该是一致的,二者两者点积应该为1,再减1之后应该为0,其中,ni是待优化量。
第二项,ni和nj属于相邻的墙面的法向量,相邻墙面的法向量垂直,二者法向量的点积为0,ni和nj均为待优化量。
第三项,ni和nfloor分别表示墙面和地面,墙面和地面是相互垂直的,ni和nfloor的点积为0,nfloor是已知的,即相机坐标系下的重力方向,ni是待优化量。
第四项,表示第i个墙面和地面之间的交界线,表示将该交界线投影至图像坐标系中,投影至图像坐标系的交界线与从图像中检测出的交界线的差为0,是已知的,是基于霍夫变换检测目标图像中存在的墙面与地面的交界线。
第五项,天花板交界线与从图像中检测出的天花板线之间的距离为0。
第六项,相邻的第i个墙面和第j个墙面之间的交界线,与从图像中检测出的第i个墙面和第j个墙面之间的交界线之间的差为0。
在本实施例中,经过上述优化之后,可以得到第一空间对象对应的目标三维模型。在本申请实施例中,首先并不直接从目标图像中检测3D结构信息,而是从目标图像中直接检测出交界线和消影点等2D视觉信息,然后,利用这些2D视觉信息和室内场景的曼哈顿约束、消影点约束等生成3D结构信息,相对于直接基于目标图像回归3D结构信息的方法,本申请实施例的稳定性更高,可解释性更强,对于用户拍摄角度倾斜大、场景杂乱的情况效果更具鲁棒性。另外,本申请实施例中,直接从目标图像中检测墙线、地线以及天花板线等交界线,并以这些交界线作为约束来优化墙面的三维方程,最后生成的目标三维模型和实际的墙面、地面以及天花板的分割线贴合程度更高,模型质量更高。可选地,在目标三维模型中还可以体现2D图像中的纹理等其它真实场景信息。
在得到第一空间对象对应的目标三维模型之后,基于目标三维模型可以开展各种应用。例如,可以将该目标三维模型展示给用户,以供用户查看或了解第一空间对象的三维结构。又例如,将目标三维模型应用于在线家装场景中,供用于查看家装对象(如挂画、沙发等)与第一空间对象真实环境的搭配效果,根据搭配效果选择家装方案。又例如,基于目标三维模型进行商品选购,以查看待选购商品与第一空间对象真实环境的搭配效果,根据搭配效果进行商品选购。下面实施例中将以在线家装场景和在线购物场景为例进行说明。
图4a为本申请示例性实施例提供的一种在线家装方法的流程示意图。如图4a所示,该方法包括:
401a、响应图像上传操作,获取第一空间对象对应的目标图像,第一空间对象是目标物理空间中的至少部分空间对象;
402a、响应目标家装对象在目标图像上的放置操作,将目标家装对象融合到第一空间 对象对应的目标三维模型中,以得到融合目标家装对象的目标三维模型;
403a、将融合目标家装对象的目标三维模型投影至目标图像上,以得到包含目标家装对象的家装效果图。在本实施例中,在线家装方法可以在终端设备上实现,也可以由终端设备和服务端设备配合实现,对此不做限定。下面以终端设备和服务端设备配合实现为例进行描述。
其中,用户可以打开终端设备上的家装App(Application),进入在线家装页面;在在线家装页面上,包括多种家装方式,例如基于样板间的家装方式,基于真实背景图的家装方式等;用户可以选择基于真实背景图的家装方式,则响应于用户选择基于真实背景图的家装方式操作,可以展示真实背景图的添加页面。用户可以通过终端设备采集第一空间对象的目标图像,也可以直接从图库中选择第一空间对象的目标图像;家装App响应于图像采集操作或图像选择操作,可获取第一空间对象对应的目标图像,第一空间对象可以是目标物理空间中的至少部分空间对象。例如,目标物理空间是用户待装修的房屋,该房屋中包括卧室、客厅、厨房或卫生间等,第一空间对象可以是卧室和客厅,也可以是客厅的局部空间,对此不做限定,详细内容可参见前述实施例,在此暂不赘述。
在本实施例中,家装App获取第一空间对象对应的目标图像之后,可以将目标图像上传至服务端设备,服务端设备采用前述方法实施例中描述的三维场景重建方法,生成第一空间对象对应的目标三维模型,关于目标三维模型的生成方式可参见前述实施例,在此不再赘述。在此说明,生成目标三维模型的过程对用户来说无感知的。如图4b为本申请实施例示出的目标图像以及从目标图像到生成目标三维模型的过程示意图。在图4b中,从模型生成原理的角度对从目标图像到生成目标三维模型的带颜色的中间态模型和同时带颜色和纹理的中间态模型进行了示意,但实际过程中通常不会产生中间态模型。需要说明的是,中间态模型和目标三维模型的视角不同,但并不限于此。另外,图4b中的颜色未做明确体现,仅以灰度颜色进行示意。
在用户选择目标图像之后,家装App可以在目标图像的关联区域中展示多种家装对象,例如,家装对象可以是家电、家具或装饰物品等。在图4c和图4d中,以在目标图像的下发区域中展示沙发、茶几、地毯、单椅、挂画为例进行图示。在图4c和图4d中,用户可以从多种家装对象中选择目标家装对象,将目标家装对象放置在目标图像上,例如,可以拖动选中的目标家装对象的图像到目标图像上的相应位置,然后松开。家装App可以响应目标家装对象在目标图像上的放置操作,将目标家装对象的标识信息(例如名称、商品ID或图像等)以及目标家装对象在目标图像上的位置范围信息提供给服务端设备;服务端设备根据目标家装对象的标识信息,获取目标家装对应的三维模型,根据目标家装对象在目标图像上的位置范围信息,将目标家装对象的三维模型融合到该目标三维模型上中,并将融合目标家装对象的目标三维模型投影至目标图像上,以得到包含目标家装对象的家装效果图,并将该家装效果图提供给终端设备;终端设备展示该家装效果图。在图4c中,以用户选择挂画,并将该挂画放置到主体墙面的中心位置为例,图4c所示的家装效果图中透视 变形正确,同时可以限制挂画的移动区域为墙面部分,具有较高的精确度。在图4d中,以用户先后选择沙发和地毯,并将沙发放置在主体墙面和地面的中线上,将地毯铺贴在地面上为例,图4d所示的家装效果图中透视关系正确,放置位置合理。
在上述实施例中,以终端设备与服务端设备相互配合实施在线家装过程为例,但并不限于此。当然,也可以由终端设备独立实施。具体地,家装App在获取第一空间对象对应的目标图像之后,可以采用前述方法实施例中描述的三维场景重建方法,生成第一空间对象对应的目标三维模型;与此同时,家装App可以在目标图像的关联区域中展示多种家装对象,家装App可以响应目标家装对象在目标图像上的放置操作,获取目标家装对象的标识信息以及目标家装对象在目标图像上的位置范围信息;根据目标家装对象的标识信息,获取目标家装对应的三维模型,根据目标家装对象在目标图像上的位置范围信息,将目标家装对象的三维模型融合到该目标三维模型上中,并将融合目标家装对象的目标三维模型投影至目标图像上,以得到包含目标家装对象的家装效果图并展示该家装效果图。
在此说明,在上述实施例中,以实时构建目标三维模型为例进行说明,但并不限于此,例如,可以在首次上传目标图像时构建目标三维模型,在后续在线家装时,可以直接使用预先构建的目标三维模型。
图4e为本申请示例性实施例提供的一种商品选择方法的流程示意图。如图4e所示,该方法包括:
401b、响应商品页面上的选择操作,确定被选择的目标商品,目标商品具有商品三维模型;
402b、响应搭配效果查看操作,选择待与目标商品进行搭配的第一空间对象对应的目标图像;
403b、将商品三维模型添加至第一空间对象对应的目标三维模型中,以得到融合目标商品的目标三维模型;
404b、将融合目标商品的目标三维模型投影至目标图像上,以得到目标商品与第一空间对象的搭配效果图。
在本实施例中,商品选择方法可以在终端设备上实现,也可以由终端设备和服务端设备配合实现,对此不做限定。下面以商品选择方法在终端设备实现为例进行描述,但并不限于此。
在本实施例中,在购物场景中,用户在购买一些家具、家电或装饰类商品时,为了便于用户选购与实际环境更加适配的商品,用户可以将选择的商品添加到第一空间对象对应的目标图像中,查看添加后的搭配效果。其中,第一空间对象可以是家居环境、商场环境、办公场景等任何具有空间概念的区域,商品可以是任何需要放置在第一空间对象中的物品,例如,在家居环境中,商品可以是衣柜、挂画或桌椅等,商场环境中,商品可以是储物柜、置物架或货架等。
在本实施例中,终端设备上可以展示商品页面,商品页面上展示有多种商品,用户可 以在商品页面上选择目标商品,终端设备可以响应商品页面上的选择操作,确定被选择的目标商品,目标商品对应有商品三维模型。
在本实施例中,在购物车页面、下单页面或商品详情页面上可以增设搭配效果查看控件,用户通过该控件可以发起查看商品搭配效果的操作,电商App响应于该操作,可以让用户选择第一空间对象对应的目标图像;用户可以通过终端设备采集第一空间对象的目标图像,也可以直接从图库中选择第一空间对象的目标图像;电商App获取目标图像之后,可以采用上述实施例提供的三维场景重建方法构建出第一空间对象对应的目标三维模型,并将用户选择的目标商品的商品三维模型添加至第一空间对象对应的目标三维模型中,以得到融合目标商品的目标三维模型;将融合目标商品的目标三维模型投影至目标图像上,以得到目标商品与第一空间对象的搭配效果图,并展示搭配效果图。用户通过该搭配效果图可以确定是否选购该目标商品。进一步可选地,还可以根据预设的搭配规则或策略,自动对搭配效果图进行评分并给出搭配分值,以辅助用户确定是否选购该目标商品,有利于提高用户的购物体验,提高用户成功选购所需商品的概率,进而降低退换货概率。
需要说明的是,上述实施例所提供方法的各步骤的执行主体均可以是同一设备,或者,该方法也由不同设备作为执行主体。比如,步骤101至步骤103的执行主体可以为设备A;又比如,步骤101和102的执行主体可以为设备A,步骤103的执行主体可以为设备B;等等。
另外,在上述实施例及附图中的描述的一些流程中,包含了按照特定顺序出现的多个操作,但是应该清楚了解,这些操作可以不按照其在本文中出现的顺序来执行或并行执行,操作的序号如101、102等,仅仅是用于区分开各个不同的操作,序号本身不代表任何的执行顺序。另外,这些流程可以包括更多或更少的操作,并且这些操作可以按顺序执行或并行执行。需要说明的是,本文中的“第一”、“第二”等描述,是用于区分不同的消息、设备、模块等,不代表先后顺序,也不限定“第一”和“第二”是不同的类型。
图5为本申请示例性实施例提供的一种三维场景重建装置的结构示意图,如图5所示,该装置包括:获取模块51、检测模块52、确定模块53、重建模块54和优化模块55。
获取模块51,用于获取第一空间对象对应的目标图像,第一空间对象是目标物理空间中的至少部分空间对象;
检测模块52,用于检测目标图像中存在的多条交界线以及至少两个正交主方向上的消影点,交界线是第一空间对象中相邻物理主体结构之间的交线;
确定模块53,用于根据至少两个正交主方向上的消影点,确定目标相机的相机内参数和相机坐标系下的重力方向,目标相机是指拍摄目标图像使用的相机;
重建模块54,用于根据多条交界线和相机坐标系下的重力方向,重建第一空间对象对应的初始三维模型,初始三维模型中包括与相邻物理主体结构对应的相邻模型主体结构;
优化模块55,用于根据多条交界线之间的约束关系、至少两个正交主方向上的消影点之间的约束关系以及相机内参数,对初始三维模型进行优化,以得到目标三维模型。
在一可选实施例中,检测模块52具体用于:将目标图像输入基于霍夫变换的交界线检测模型进行交界线的检测,以获取目标图像中存在的多条交界线;将目标图像输入基于霍夫变换的消影点检测模型进行消影点检测,以获取目标图像中至少两个正交主方向上的消影点。
在一可选实施例中,检测模块52具体用于:将目标图像输入基于霍夫变换的交界线检测模型中,在该模型中,利用融合跳跃连接和多尺度特征的第一特征提取网络对目标图像进行特征提取,得到多个尺度的第一目标特征图;利用基于极坐标的霍夫变换对多个尺度的第一目标特征图进行霍夫变换,以将目标图像中存在的直线映射为霍夫空间中的点;在霍夫空间中选择与交界线特征匹配的多个目标点,并将多个目标点重新映射到图像空间中,以得到目标图像中存在的多条交界线。
在一可选实施例中,检测模块52具体用于:对目标图像进行特征提取,得到最大尺度的第一中间特征图,对最大尺度的第一中间特征图进行N次下采样处理,得到多个其它尺度的第一中间特征图;将最小尺度的第一中间特征图作为最小尺度的第一目标特征图,对最小尺度的第一目标特征图进行N次上采样处理,并在每次上采样处理中与相同尺度的第一中间特征图进行跳跃连接,以得到其它尺度的第一目标特征图。
在一可选实施例中,检测模块52具体用于:对多个尺度的第一目标特征图进行霍夫变换,以得到基于极坐标的霍夫空间中多个尺度的第二目标特征图;对多个尺度的第二目标特征图进行尺度变换,以得到尺度相同的多个特征图,对尺度相同的多个特征图进行拼接,得到第三目标特征图;对第三目标特征图进行卷积降维,以得到霍夫空间中的二维图像,二维图像中包含多个点,每个点与目标图像中存在的直线对应。
在一可选实施例中,检测模块52具体用于:将目标图像输入基于霍夫变换的消影点检测模型中,在该模型中,利用第二特征提取网络对目标图像进行特征提取,得到第四目标特征图;利用基于高斯球面的霍夫变换对第四目标特征图进行霍夫变换,以将目标图像中存在的直线映射为高斯球面空间中的点;在高斯球面空间中选择概率值和角度值满足要求的至少两个正交主方向上的消影点,并重新映射到图像空间中,以得到目标图像中至少两个正交主方向上的消影点。
在一可选实施例中,检测模块52具体用于:对第四目标特征图进行霍夫变换,以得到基于极坐标的霍夫空间中的第五目标特征图;在霍夫空间中对第五目标特征图进行霍夫卷积,以得到第六目标特征图;对第六目标特征图进行球面高斯球面变换,以得到高斯球面空间中的第七目标特征图;在高斯球面空间中对第七目标特征图进行球面卷积,得到第八目标特征图,第八目标特征图中包括多个点,每个点与目标图像中存在的直线对应。
在一可选实施例中,检测模块52具体用于:根据第八目标特征图中各个点的概率值,从多个点中选择概率值满足设定要求的点作为消影点;从消影点中选择角度大于设定角度的至少两个消影点,作为至少两个正交主方向上的消影点。
在一可选实施例中,确定模块53具体用于:在至少两个正交主方向上的消影点中包 括至少两个有限消影点的情况下,从至少两个有限消影点中选择距离相机光心最近的两个目标消影点;根据两个目标消影点与相机内参数的约束关系,确定目标相机的相机内参数;根据相机内参数,将至少两个正交主方向中位于重力方向上的消影点转换到相机坐标系中,以得到相机坐标系下的重力方向。
在一可选实施例中,重建模块54具体用于:根据相机坐标系下的重力方向,在相机坐标系下构建三维空间的基准平面,基准平面与第一空间对象中包含的基准物理主体结构对应;从多条交界线中识别与基准物理主体结构相交的多条基准交界线,根据多条基准交界线之间的交点位置,在基准平面上构建与基准物理主体结构对应的基准模型主体结构;根据预设的相机光心到基准平面的高度和多条基准交界线之间的交点位置,在基准模型主体结构上构建与其它物理主体结构对应的其它模型主体结构,以得到第一空间对象对应的初始三维模型。
在一可选实施例中,重建模块54具体用于:从多条基准交界线中选择有效基准交界线,根据有效基准交界线与图像坐标系中x轴的夹角对有效基准交界线进行排序,以得到有效基准交界线之间的相邻关系;根据有效基准交界线之间的相邻关系,确定有效基准交界线之间的交点位置,并将交点位置划分为与目标图像的边界相交的第一交点位置以及未与目标图像的边界相交的第二交点位置;根据第一交点位置和第二交点位置,在基准平面上绘制基准物理主体结构对应的模型边界,以得到基准模型主体结构。
在一可选实施例中,重建模块54具体用于:根据多条基准交界线与图像坐标系中x轴的夹角和/或多条基准交界线的长度,从多条基准交界线中选择第一基准交界线;根据其它基准交界线与第一基准交界线之间的夹角,将夹角小于设定夹角阈值的基准交界线剔除,将未被剔除的基准交界线和第一基准交界线作为有效基准交界线。
在一可选实施例中,重建模块54具体用于:根据第二交点位置,确定基准模型主体结构上相交于第二交点位置的相邻模型边界;根据预设的相机光心到基准平面的高度和预设的缩放比例,确定其它模型主体结构的初始高度;根据初始高度,在相交于第二交点位置的相邻模型边界上构建其它模型主体结构,以得到第一空间对象对应的初始三维模型。
在一可选实施例中,优化模块55具体用于:根据多条交界线之间的约束关系、至少两个正交主方向上的消影点之间的约束关系以及相机内参数,构建以各模型主体结构的位置参数和/或高度参数为优化变量的优化函数,位置参数包括模型主体结构的法向量和到相机光心的距离;采用最小二乘算法对优化函数进行求解,以得到各模型主体结构的优化后的位置参数和/或高度参数,根据优化后的位置参数和/或高度参数调整各模型主体结构的位置,以得到目标三维模型。
在一可选实施例中,优化模块55具体用于:针对每一个模型主体结构,根据相机内参数和与该模型主体结构垂直的消影点,生成该模型主体结构在相机坐标系下的参考法向量,以该模型主体结构的法向量与参考法向量的点积为1作为优化目标构建第一类优化项;针对任意相邻模型主体结构,以任意相邻模型主体结构之间的法向量的点积为0作为优化 目标构建第二类优化项;针对任意相邻模型主体结构,根据相机内参数和任意相邻模型主体结构的法向量和到相机光心的距离,生成任意相邻模型主体结构在图像坐标系下的交界线;以任意相邻模型主体结构在图像坐标系下的交界线和目标图像中对应相邻物理主体结构之间的交界线相同为优化目标构建第三类优化项;根据第一类优化项、第二类优化项和第三类优化项,生成优化函数。其中,上述各操作的详细实施方式,可参见前述实施例,在此不再赘述。
本申请实施例还提供一种在线家装装置,该装置包括:获取模块、融合模块和投影模块。其中,获取模块,用于响应图像上传操作,获取第一空间对象对应的目标图像,第一空间对象是目标物理空间中的至少部分空间对象。融合模块,用于响应目标家装对象在目标图像上的放置操作,将目标家装对象融合到第一空间对象对应的目标三维模型中,以得到融合目标家装对象的目标三维模型。投影模块,用于将融合目标家装对象的目标三维模型投影至目标图像上,以得到包含目标家装对象的家装效果图;其中,目标三维模型是根据本申请实施例提供的三维场景重建方法中的步骤构建的。其中,上述各操作的详细实施方式,可参见前述实施例,在此不再赘述。
本申请实施例还提供一种商品选择装置,该装置包括:确定模块,选择模块、添加模块和投影模块。确定模块,用于响应商品页面上的选择操作,确定被选择的目标商品,目标商品具有商品三维模型。选择模块,用于响应搭配效果查看操作,选择待与目标商品进行搭配的第一空间对象对应的目标图像。添加模块,用于将商品三维模型添加至第一空间对象对应的目标三维模型中,以得到融合目标商品的目标三维模型。融合模块,用于将融合目标商品的目标三维模型投影至目标图像上,以得到目标商品与第一空间对象的搭配效果图;其中,目标三维模型是根据本申请实施例提供的三维场景重建方法中的步骤构建的。其中,上述各操作的详细实施方式,可参见前述实施例,在此不再赘述。
图6为本申请示例性实施例提供的一种电子设备的结构示意图。如图6所示,该设备包括:存储器64和处理器65。
存储器64,用于存储计算机程序,并可被配置为存储其它各种数据以支持在计算平台上的操作。这些数据的示例包括用于在计算平台上操作的任何应用程序或方法的指令。
存储器64可以由任何类型的易失性或非易失性存储设备或者它们的组合实现,如静态随机存取存储器(SRAM),电可擦除可编程只读存储器(EEPROM),可擦除可编程只读存储器(EPROM),可编程只读存储器(PROM),只读存储器(ROM),磁存储器,快闪存储器,磁盘或光盘。
处理器65,与存储器64耦合,用于执行存储器64中的计算机程序,以用于:获取第一空间对象对应的目标图像,第一空间对象是目标物理空间中的至少部分空间对象;检测目标图像中存在的多条交界线以及至少两个正交主方向上的消影点,交界线是第一空间对象中相邻物理主体结构之间的交线;根据至少两个正交主方向上的消影 点,确定目标相机的相机内参数和相机坐标系下的重力方向,目标相机是指拍摄目标图像使用的相机;根据多条交界线和相机坐标系下的重力方向,重建第一空间对象对应的初始三维模型,初始三维模型中包括与相邻物理主体结构对应的相邻模型主体结构;根据多条交界线之间的约束关系、至少两个正交主方向上的消影点之间的约束关系以及相机内参数,对初始三维模型进行优化,以得到目标三维模型。
在一可选实施例中,处理器65在基于霍夫变换检测目标图像中存在的多条交界线以及至少两个正交主方向上的消影点时,具体用于:将目标图像输入基于霍夫变换的交界线检测模型进行交界线的检测,以获取目标图像中存在的多条交界线;将目标图像输入基于霍夫变换的消影点检测模型进行消影点检测,以获取目标图像中至少两个正交主方向上的消影点。
在一可选实施例中,处理器65在将目标图像输入基于霍夫变换的交界线检测模型进行交界线的检测,以获取目标图像中存在的多条交界线时,具体用于:将目标图像输入基于霍夫变换的交界线检测模型中,在该模型中,利用融合跳跃连接和多尺度特征的第一特征提取网络对目标图像进行特征提取,得到多个尺度的第一目标特征图;利用基于极坐标的霍夫变换对多个尺度的第一目标特征图进行霍夫变换,以将目标图像中存在的直线映射为霍夫空间中的点;在霍夫空间中选择与交界线特征匹配的多个目标点,并将多个目标点重新映射到图像空间中,以得到目标图像中存在的多条交界线。
在一可选实施例中,处理器65在利用融合跳跃连接和多尺度特征的第一特征提取网络对目标图像进行特征提取,得到多个尺度的第一目标特征图时,具体用于:对所述目标图像进行特征提取,得到最大尺度的第一中间特征图,对所述最大尺度的第一中间特征图进行N次下采样处理,得到其它尺度的第一中间特征图;将最小尺度的第一中间特征图作为最小尺度的第一目标特征图,对所述最小尺度的第一目标特征图进行N次上采样处理,并在每次上采样处理中与相同尺度的第一中间特征图进行跳跃连接,以得到其它尺度的第一目标特征图。
在一可选实施例中,处理器65在利用基于极坐标的霍夫变换对多个尺度的第一目标特征图进行霍夫变换,以将目标图像中存在的直线映射为霍夫空间中的点时,具体用于:对多个尺度的第一目标特征图进行霍夫变换,以得到基于极坐标的霍夫空间中多个尺度的第二目标特征图;对多个尺度的第二目标特征图进行尺度变换,以得到尺度相同的多个特征图,对尺度相同的多个特征图进行拼接,得到第三目标特征图;对第三目标特征图进行卷积降维,以得到霍夫空间中的二维图像,二维图像中包含多个点,每个点与目标图像中存在的直线对应。
在一可选实施例中,处理器65在将目标图像输入基于霍夫变换的消影点检测模型进行消影点检测,以获取目标图像中至少两个正交主方向上的消影点时,具体用于:将目标图像输入基于霍夫变换的消影点检测模型中,在该模型中,利用第二特征提取 网络对目标图像进行特征提取,得到第四目标特征图;利用基于高斯球面的霍夫变换对第四目标特征图进行霍夫变换,以将目标图像中存在的直线映射为高斯球面空间中的点;在高斯球面空间中选择概率值和角度值满足要求的至少两个正交主方向上的消影点,并重新映射到图像空间中,以得到目标图像中至少两个正交主方向上的消影点。
在一可选实施例中,处理器65在利用基于高斯球面的霍夫变换对第四目标特征图进行霍夫变换,以将目标图像中存在的直线映射为高斯球面空间中的点时,具体用于:对第四目标特征图进行霍夫变换,以得到基于极坐标的霍夫空间中的第五目标特征图;在霍夫空间中对第五目标特征图进行霍夫卷积,以得到第六目标特征图;对第六目标特征图进行球面高斯球面变换,以得到高斯球面空间中的第七目标特征图;在高斯球面空间中对第七目标特征图进行球面卷积,得到第八目标特征图,第八目标特征图中包括多个点,每个点与目标图像中存在的直线对应。
在一可选实施例中,处理器65在高斯球面空间中选择概率值和角度值满足要求的至少两个正交主方向上的消影点时,具体用于:根据第八目标特征图中各个点的概率值,从多个点中选择概率值满足设定要求的点作为消影点;从消影点中选择角度大于设定角度的至少两个消影点,作为至少两个正交主方向上的消影点。
在一可选实施例中,处理器65在根据至少两个正交主方向上的消影点,确定目标相机的相机内参数和相机坐标系下的重力方向时,具体用于:在至少两个正交主方向上的消影点中包括至少两个有限消影点的情况下,从至少两个有限消影点中选择距离相机光心最近的两个目标消影点;根据两个目标消影点与相机内参数的约束关系,确定目标相机的相机内参数;根据相机内参数,将至少两个正交主方向中位于重力方向上的消影点转换到相机坐标系中,以得到相机坐标系下的重力方向。
在一可选实施例中,处理器65在根据多条交界线和相机坐标系下的重力方向,重建第一空间对象对应的初始三维模型时,具体用于:根据相机坐标系下的重力方向,在相机坐标系下构建三维空间的基准平面,基准平面与第一空间对象中包含的基准物理主体结构对应;从多条交界线中识别与基准物理主体结构相交的多条基准交界线,根据多条基准交界线之间的交点位置,在基准平面上构建与基准物理主体结构对应的基准模型主体结构;根据预设的相机光心到基准平面的高度和多条基准交界线之间的交点位置,在基准模型主体结构上构建与其它物理主体结构对应的其它模型主体结构,以得到第一空间对象对应的初始三维模型。
在一可选实施例中,处理器65在根据多条基准交界线之间的交点位置,在基准平面上构建与基准物理主体结构对应的基准模型主体结构时,具体用于:从多条基准交界线中选择有效基准交界线,根据有效基准交界线与图像坐标系中x轴的夹角对有效基准交界线进行排序,以得到有效基准交界线之间的相邻关系;根据有效基准交界线之间的相邻关系,确定有效基准交界线之间的交点位置,并将交点位置划分为与目标图像的边界相交的第一交点位置以及未与目标图像的边界相交的第二交点位置;根据 第一交点位置和第二交点位置,在基准平面上绘制基准物理主体结构对应的模型边界,以得到基准模型主体结构。
在一可选实施例中,处理器65在从多条基准交界线中选择有效基准交界线时,具体用于:根据多条基准交界线与图像坐标系中x轴的夹角和/或多条基准交界线的长度,从多条基准交界线中选择第一基准交界线;根据其它基准交界线与第一基准交界线之间的夹角,将夹角小于设定夹角阈值的基准交界线剔除,将未被剔除的基准交界线和第一基准交界线作为有效基准交界线。
在一可选实施例中,处理器65在根据预设的相机光心到基准平面的高度和多条基准交界线之间的交点位置,在基准模型主体结构上构建与其它物理主体结构对应的其它模型主体结构,以得到第一空间对象对应的初始三维模型时,具体用于:根据第二交点位置,确定基准模型主体结构上相交于第二交点位置的相邻模型边界;根据预设的相机光心到基准平面的高度和预设的缩放比例,确定其它模型主体结构的初始高度;根据初始高度,在相交于第二交点位置的相邻模型边界上构建其它模型主体结构,以得到第一空间对象对应的初始三维模型。
在一可选实施例中,处理器65在根据多条交界线之间的约束关系、至少两个正交主方向上的消影点之间的约束关系以及相机内参数,对初始三维模型进行优化,以得到目标三维模型时,具体用于:根据多条交界线之间的约束关系、至少两个正交主方向上的消影点之间的约束关系以及相机内参数,构建以各模型主体结构的位置参数和/或高度参数为优化变量的优化函数,位置参数包括模型主体结构的法向量和到相机光心的距离;采用最小二乘算法对优化函数进行求解,以得到各模型主体结构的优化后的位置参数和/或高度参数,根据优化后的位置参数和/或高度参数调整各模型主体结构的位置,以得到目标三维模型。
在一可选实施例中,处理器65在根据多条交界线之间的约束关系、至少两个正交主方向上的消影点之间的约束关系以及相机内参数,构建以各模型主体结构的位置参数和/或高度参数为优化变量的优化函数时,具体用于:针对每一个模型主体结构,根据相机内参数和与该模型主体结构垂直的消影点,生成该模型主体结构在相机坐标系下的参考法向量,以该模型主体结构的法向量与参考法向量的点积为1作为优化目标构建第一类优化项;针对任意相邻模型主体结构,以任意相邻模型主体结构之间的法向量的点积为0作为优化目标构建第二类优化项;针对任意相邻模型主体结构,根据相机内参数和任意相邻模型主体结构的法向量和到相机光心的距离,生成任意相邻模型主体结构在图像坐标系下的交界线;以任意相邻模型主体结构在图像坐标系下的交界线和目标图像中对应相邻物理主体结构之间的交界线相同为优化目标构建第三类优化项;根据第一类优化项、第二类优化项和第三类优化项,生成优化函数。
进一步,如图6所示,该电子设备还包括:通信组件66、显示器67、电源组件68、音频组件69等其它组件。图6中仅示意性给出部分组件,并不意味着电子设备只包括图6 所示组件。
本申请实施例还提供一种电子设备,该电子设备的结构与图6所示电子设备的结构相同或相似,具体可参见图6所示电子设备的结构实现,本实施例提供的电子设备与图6所示实施例中电子设备的区别主要在于:处理器执行存储器中存储的计算机程序所实现的功能不同。对本实施例提供的电子设备来说,其处理器执行存储器中存储的计算机程序,可用于:响应图像上传操作,获取第一空间对象对应的目标图像,第一空间对象是目标物理空间中的至少部分空间对象;响应目标家装对象在目标图像上的放置操作,将目标家装对象融合到第一空间对象对应的目标三维模型中,以得到融合目标家装对象的目标三维模型;将融合目标家装对象的目标三维模型投影至目标图像上,以得到包含目标家装对象的家装效果图;其中,目标三维模型是根据本申请实施例提供的三维场景重建方法中的步骤构建的。
本申请实施例还提供一种电子设备,该电子设备的结构与图6所示电子设备的结构相同或相似,具体可参见图6所示电子设备的结构实现,本实施例提供的电子设备与图6所示实施例中电子设备的区别主要在于:处理器执行存储器中存储的计算机程序所实现的功能不同。对本实施例提供的电子设备来说,其处理器执行存储器中存储的计算机程序,可用于:响应商品页面上的选择操作,确定被选择的目标商品,目标商品具有商品三维模型;响应搭配效果查看操作,选择待与目标商品进行搭配的第一空间对象对应的目标图像;将商品三维模型添加至第一空间对象对应的目标三维模型中,以得到融合目标商品的目标三维模型;将融合目标商品的目标三维模型投影至目标图像上,以得到目标商品与第一空间对象的搭配效果图;其中,目标三维模型是根据本申请实施例提供的三维场景重建方法中的步骤构建的。
相应地,本申请实施例还提供一种存储有计算机程序的计算机可读存储介质,计算机程序被执行时能够实现上述图1b、图4a和图4e所示方法中的各步骤。
上述图6中的通信组件被配置为便于通信组件所在设备和其他设备之间有线或无线方式的通信。通信组件所在设备可以接入基于通信标准的无线网络,如WiFi,2G、3G、4G/LTE、5G等移动通信网络,或它们的组合。在一个示例性实施例中,通信组件经由广播信道接收来自外部广播管理系统的广播信号或广播相关信息。在一个示例性实施例中,通信组件还包括近场通信(NFC)模块,以促进短程通信。例如,在NFC模块可基于射频识别(RFID)技术,红外数据协会(IrDA)技术,超宽带(UWB)技术,蓝牙(BT)技术和其他技术来实现。
上述图6中的显示器包括屏幕,其屏幕可以包括液晶显示器(LCD)和触摸面板(TP)。如果屏幕包括触摸面板,屏幕可以被实现为触摸屏,以接收来自用户的输入信号。触摸面板包括一个或多个触摸传感器以感测触摸、滑动和触摸面板上的手势。触摸传感器可以不仅感测触摸或滑动动作的边界,而且还检测与触摸或滑动操作相关的持续时间和压力。
上述图6中的电源组件,为电源组件所在设备的各种组件提供电力。电源组件可以包 括电源管理系统,一个或多个电源,及其他与为电源组件所在设备生成、管理和分配电力相关联的组件。
上述图6中的音频组件,可被配置为输出和/或输入音频信号。例如,音频组件包括一个麦克风(MIC),当音频组件所在设备处于操作模式,如呼叫模式、记录模式和语音识别模式时,麦克风被配置为接收外部音频信号。所接收的音频信号可以被进一步存储在存储器或经由通信组件发送。在一些实施例中,音频组件还包括一个扬声器,用于输出音频信号。
本领域内的技术人员应明白,本申请的实施例可提供为方法、系统、或计算机程序产品。因此,本申请可采用完全硬件实施例、完全软件实施例、或结合软件和硬件方面的实施例的形式。而且,本申请可采用在一个或多个其中包含有计算机可用程序代码的计算机可用存储介质(包括但不限于磁盘存储器、CD-ROM、光学存储器等)上实施的计算机程序产品的形式。
本申请是参照根据本申请实施例的方法、设备(系统)、和计算机程序产品的流程图和/或方框图来描述的。应理解可由计算机程序指令实现流程图和/或方框图中的每一流程和/或方框、以及流程图和/或方框图中的流程和/或方框的结合。可提供这些计算机程序指令到通用计算机、专用计算机、嵌入式处理机或其他可编程数据处理设备的处理器以产生一个机器,使得通过计算机或其他可编程数据处理设备的处理器执行的指令产生用于实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能的装置。
这些计算机程序指令也可存储在能引导计算机或其他可编程数据处理设备以特定方式工作的计算机可读存储器中,使得存储在该计算机可读存储器中的指令产生包括指令装置的制造品,该指令装置实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能。
这些计算机程序指令也可装载到计算机或其他可编程数据处理设备上,使得在计算机或其他可编程设备上执行一系列操作步骤以产生计算机实现的处理,从而在计算机或其他可编程设备上执行的指令提供用于实现在流程图一个流程或多个流程和/或方框图一个方框或多个方框中指定的功能的步骤。
在一个典型的配置中,计算设备包括一个或多个处理器(CPU)、输入/输出接口、网络接口和内存。
内存可能包括计算机可读介质中的非永久性存储器,随机存取存储器(RAM)和/或非易失性内存等形式,如只读存储器(ROM)或闪存(flash RAM)。内存是计算机可读介质的示例。
计算机可读介质包括永久性和非永久性、可移动和非可移动媒体可以由任何方法或技术来实现信息存储。信息可以是计算机可读指令、数据结构、程序的模块或其他数据。计算机的存储介质的例子包括,但不限于相变内存(PRAM)、静态随机存取存储器(SRAM)、 动态随机存取存储器(DRAM)、其他类型的随机存取存储器(RAM)、只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、快闪记忆体或其他内存技术、只读光盘只读存储器(CD-ROM)、数字多功能光盘(DVD)或其他光学存储、磁盒式磁带,磁带磁盘存储或其他磁性存储设备或任何其他非传输介质,可用于存储可以被计算设备访问的信息。按照本文中的界定,计算机可读介质不包括暂存电脑可读媒体(transitory media),如调制的数据信号和载波。
还需要说明的是,术语“包括”、“包含”或者其任何其他变体意在涵盖非排他性的包含,从而使得包括一系列要素的过程、方法、商品或者设备不仅包括那些要素,而且还包括没有明确列出的其他要素,或者是还包括为这种过程、方法、商品或者设备所固有的要素。在没有更多限制的情况下,由语句“包括一个……”限定的要素,并不排除在包括所述要素的过程、方法、商品或者设备中还存在另外的相同要素。
以上所述仅为本申请的实施例而已,并不用于限制本申请。对于本领域技术人员来说,本申请可以有各种更改和变化。凡在本申请的精神和原理之内所作的任何修改、等同替换、改进等,均应包含在本申请的权利要求范围之内。

Claims (16)

  1. 一种三维场景重建方法,其特征在于,包括:
    获取第一空间对象对应的目标图像,所述第一空间对象是目标物理空间中的至少部分空间对象;
    检测所述目标图像中存在的多条交界线以及至少两个正交主方向上的消影点,所述交界线是所述第一空间对象中相邻物理主体结构之间的交线;
    根据所述至少两个正交主方向上的消影点,确定目标相机的相机内参数和相机坐标系下的重力方向,所述目标相机是指拍摄所述目标图像使用的相机;
    根据所述多条交界线和所述相机坐标系下的重力方向,重建所述第一空间对象对应的初始三维模型,所述初始三维模型中包括与所述相邻物理主体结构对应的相邻模型主体结构;
    根据所述多条交界线之间的约束关系、所述至少两个正交主方向上的消影点之间的约束关系以及所述相机内参数,对所述初始三维模型进行优化,以得到目标三维模型。
  2. 根据权利要求1所述的方法,其特征在于,检测所述目标图像中存在的多条交界线以及至少两个正交主方向上的消影点,包括:
    将所述目标图像输入基于霍夫变换的交界线检测模型进行交界线的检测,以获取所述目标图像中存在的多条交界线;
    将所述目标图像输入基于霍夫变换的消影点检测模型进行消影点检测,以获取所述目标图像中至少两个正交主方向上的消影点。
  3. 根据权利要求2所述的方法,其特征在于,将所述目标图像输入基于霍夫变换的交界线检测模型进行交界线的检测,以获取所述目标图像中存在的多条交界线,包括:
    在基于霍夫变换的交界线检测模型中,利用融合跳跃连接和多尺度特征的第一特征提取网络对所述目标图像进行特征提取,得到多个尺度的第一目标特征图;
    利用基于极坐标的霍夫变换对所述多个尺度的第一目标特征图进行霍夫变换,以将所述目标图像中存在的直线映射为霍夫空间中的点;
    在所述霍夫空间中选择与交界线特征匹配的多个目标点,并将所述多个目标点重新映射到图像空间中,以得到所述目标图像中存在的多条交界线。
  4. 根据权利要求3所述的方法,其特征在于,利用基于极坐标的霍夫变换对所述多个尺度的第一目标特征图进行霍夫变换,以将所述目标图像中存在的直线映射为霍夫空间中的点,包括:
    对所述多个尺度的第一目标特征图进行霍夫变换,以得到基于极坐标的霍夫空间中多个尺度的第二目标特征图;
    对所述多个尺度的第二目标特征图进行尺度变换,以得到尺度相同的多个特征图, 对所述尺度相同的多个特征图进行拼接,得到第三目标特征图;
    对所述第三目标特征图进行卷积降维,以得到霍夫空间中的二维图像,所述二维图像中包含多个点,每个点与所述目标图像中存在的直线对应。
  5. 根据权利要求2所述的方法,其特征在于,将所述目标图像输入基于霍夫变换的消影点检测模型进行消影点检测,以获取所述目标图像中至少两个正交主方向上的消影点,包括:
    在基于霍夫变换的消影点检测模型中,利用第二特征提取网络对所述目标图像进行特征提取,得到第四目标特征图;
    利用基于高斯球面的霍夫变换对所述第四目标特征图进行霍夫变换,以将所述目标图像中存在的直线映射为高斯球面空间中的点;
    在所述高斯球面空间中选择概率值和角度值满足要求的至少两个正交主方向上的消影点,并重新映射到图像空间中,以得到所述目标图像中至少两个正交主方向上的消影点。
  6. 根据权利要求5所述的方法,其特征在于,利用基于高斯球面的霍夫变换对所述第四目标特征图进行霍夫变换,以将所述目标图像中存在的直线映射为高斯球面空间中的点,包括:
    对所述第四目标特征图进行霍夫变换,以得到基于极坐标的霍夫空间中的第五目标特征图;
    在所述霍夫空间中对所述第五目标特征图进行霍夫卷积,以得到第六目标特征图;
    对所述第六目标特征图进行高斯球面变换,以得到高斯球面空间中的第七目标特征图;
    在所述高斯球面空间中对所述第七目标特征图进行球面卷积,得到第八目标特征图,所述第八目标特征图中包括多个点,每个点与所述目标图像中存在的直线对应。
  7. 根据权利要求1所述的方法,其特征在于,根据所述至少两个正交主方向上的消影点,确定目标相机的相机内参数和相机坐标系下的重力方向,包括:
    在所述至少两个正交主方向上的消影点中包括至少两个有限消影点的情况下,从所述至少两个有限消影点中选择距离相机光心最近的两个目标消影点;
    根据所述两个目标消影点与相机内参数的约束关系,确定所述目标相机的相机内参数;
    根据所述相机内参数,将所述至少两个正交主方向中位于重力方向上的消影点转换到相机坐标系中,以得到相机坐标系下的重力方向。
  8. 根据权利要求1所述的方法,其特征在于,根据所述多条交界线和所述相机坐标系下的重力方向,重建所述第一空间对象对应的初始三维模型,包括:
    根据所述相机坐标系下的重力方向,在相机坐标系下构建三维空间的基准平面,所述基准平面与所述第一空间对象中包含的基准物理主体结构对应;
    从所述多条交界线中识别与所述基准物理主体结构相交的多条基准交界线,根据所述多条基准交界线之间的交点位置,在所述基准平面上构建与所述基准物理主体结构对应的基准模型主体结构;
    根据预设的相机光心到所述基准平面的高度和所述多条基准交界线之间的交点位置,在所述基准模型主体结构上构建与其它物理主体结构对应的其它模型主体结构,以得到所述第一空间对象对应的初始三维模型。
  9. 根据权利要求8所述的方法,其特征在于,根据所述多条基准交界线之间的交点位置,在所述基准平面上构建与所述基准物理主体结构对应的基准模型主体结构,包括:
    从所述多条基准交界线中选择有效基准交界线,根据所述有效基准交界线与图像坐标系中x轴的夹角对所述有效基准交界线进行排序,以得到所述有效基准交界线之间的相邻关系;
    根据所述有效基准交界线之间的相邻关系,确定所述有效基准交界线之间的交点位置,并将所述交点位置划分为与所述目标图像的边界相交的第一交点位置以及未与所述目标图像的边界相交的第二交点位置;
    根据所述第一交点位置和所述第二交点位置,在所述基准平面上绘制所述基准物理主体结构对应的模型边界,以得到基准模型主体结构。
  10. 根据权利要求9所述的方法,其特征在于,根据预设的相机光心到所述基准平面的高度和所述多条基准交界线之间的交点位置,在所述基准模型主体结构上构建与其它物理主体结构对应的其它模型主体结构,以得到所述第一空间对象对应的初始三维模型,包括:
    根据所述第二交点位置,确定所述基准模型主体结构上相交于所述第二交点位置的相邻模型边界;
    根据预设的相机光心到所述基准平面的高度和预设的缩放比例,确定所述其它模型主体结构的初始高度;
    根据所述初始高度,在相交于所述第二交点位置的相邻模型边界上构建其它模型主体结构,以得到所述第一空间对象对应的初始三维模型。
  11. 根据权利要求1-10任一项所述的方法,其特征在于,根据所述多条交界线之间的约束关系、所述至少两个正交主方向上的消影点之间的约束关系以及所述相机内参数,对所述初始三维模型进行优化,以得到目标三维模型,包括:
    根据所述多条交界线之间的约束关系、所述至少两个正交主方向上的消影点之间的约束关系以及所述相机内参数,构建以各模型主体结构的位置参数和/或高度参数为优化变量的优化函数,所述位置参数包括模型主体结构的法向量和到相机光心的距离;
    采用最小二乘算法对所述优化函数进行求解,以得到各模型主体结构的优化后的 位置参数和/或高度参数;
    根据所述优化后的位置参数和/或高度参数调整各模型主体结构的位置,以得到目标三维模型。
  12. 根据权利要求11所述的方法,其特征在于,根据所述多条交界线之间的约束关系、所述至少两个正交主方向上的消影点之间的约束关系以及所述相机内参数,构建以各模型主体结构的位置参数和/或高度参数为优化变量的优化函数,包括:
    针对每一个模型主体结构,根据所述相机内参数和与该模型主体结构垂直的消影点,生成该模型主体结构在相机坐标系下的参考法向量,并以该模型主体结构的法向量与参考法向量的点积为1作为优化目标构建第一类优化项;
    针对任意相邻模型主体结构,以任意相邻模型主体结构之间的法向量的点积为0作为优化目标构建第二类优化项;
    针对任意相邻模型主体结构,根据所述相机内参数和所述任意相邻模型主体结构的法向量和到相机光心的距离,生成任意相邻模型主体结构在图像坐标系下的交界线;并以任意相邻模型主体结构在图像坐标系下的交界线和所述目标图像中对应相邻物理主体结构之间的交界线相同为优化目标构建第三类优化项;
    根据所述第一类优化项、所述第二类优化项和所述第三类优化项,生成所述优化函数。
  13. 一种在线家装方法,其特征在于,包括:
    响应图像上传操作,获取第一空间对象对应的目标图像,所述第一空间对象是目标物理空间中的至少部分空间对象;
    响应目标家装对象在所述目标图像上的放置操作,将所述目标家装对象融合到所述第一空间对象对应的目标三维模型中,以得到融合所述目标家装对象的目标三维模型;
    将融合所述目标家装对象的目标三维模型投影至所述目标图像上,以得到包含所述目标家装对象的家装效果图;
    其中,所述目标三维模型是根据权利要求1-12任一项所述方法中的步骤构建的。
  14. 一种商品选择方法,其特征在于,包括:
    响应商品页面上的选择操作,确定被选择的目标商品,所述目标商品具有商品三维模型;
    响应搭配效果查看操作,选择待与所述目标商品进行搭配的第一空间对象对应的目标图像;
    将所述商品三维模型添加至所述第一空间对象对应的目标三维模型中,以得到融合所述目标商品的目标三维模型;
    将融合所述目标商品的目标三维模型投影至所述目标图像上,以得到所述目标商品与所述第一空间对象的搭配效果图;
    其中,所述目标三维模型是根据权利要求1-12中任一项所述方法中的步骤构建的。
  15. 一种电子设备,其特征在于,包括:存储器和处理器;所述存储器,用于存储计算机程序;所述处理器与所述存储器耦合,用于执行所述计算机程序,以用于执行权利要求1-12、权利要求13以及权利要求14中任一项所述方法中的步骤。
  16. 一种存储有计算机程序的计算机可读存储介质,其特征在于,当所述计算机程序被处理器执行时,致使所述处理器能够实现权利要求1-12、权利要求13以及权利要求14中任一项所述方法中的步骤。
PCT/CN2023/072122 2022-12-14 2023-01-13 三维场景重建、在线家装与商品获取方法、设备及介质 Ceased WO2024124653A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202211599959.2A CN115937422A (zh) 2022-12-14 2022-12-14 三维场景重建、在线家装与商品获取方法、设备及介质
CN202211599959.2 2022-12-14

Publications (1)

Publication Number Publication Date
WO2024124653A1 true WO2024124653A1 (zh) 2024-06-20

Family

ID=86557097

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2023/072122 Ceased WO2024124653A1 (zh) 2022-12-14 2023-01-13 三维场景重建、在线家装与商品获取方法、设备及介质

Country Status (2)

Country Link
CN (1) CN115937422A (zh)
WO (1) WO2024124653A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20240290061A1 (en) * 2023-02-28 2024-08-29 Lemon Inc. Pixel perspective estimation and refinement in an image

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120747336B (zh) * 2024-07-01 2026-04-21 荣耀终端股份有限公司 三维重建方法和电子设备

Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200005538A1 (en) * 2018-06-29 2020-01-02 Factualvr, Inc. Remote Collaboration Methods and Systems
CN110782524A (zh) * 2019-10-25 2020-02-11 重庆邮电大学 基于全景图的室内三维重建方法
CN114529566A (zh) * 2021-12-30 2022-05-24 北京城市网邻信息技术有限公司 图像处理方法、装置、设备及存储介质
CN114549765A (zh) * 2022-02-28 2022-05-27 北京京东尚科信息技术有限公司 三维重建方法及装置、计算机可存储介质
CN114756919A (zh) * 2021-12-29 2022-07-15 每平每屋(上海)科技有限公司 数据处理方法、家装设计方法、设备及存储介质
CN115359192A (zh) * 2022-10-14 2022-11-18 阿里巴巴(中国)有限公司 三维重建与商品信息处理方法、装置、设备及存储介质
CN115439607A (zh) * 2022-09-01 2022-12-06 中国民用航空总局第二研究所 一种三维重建方法、装置、电子设备及存储介质

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN103198524B (zh) * 2013-04-27 2015-08-12 清华大学 一种大规模室外场景三维重建方法
EP3086284A1 (en) * 2015-04-23 2016-10-26 Application Solutions (Electronics and Vision) Limited Camera extrinsic parameters estimation from image lines
CN108731645A (zh) * 2018-04-25 2018-11-02 浙江工业大学 基于全景图的室外全景相机姿态估计方法
CN109242958A (zh) * 2018-08-29 2019-01-18 广景视睿科技(深圳)有限公司 一种三维建模的方法及其装置
WO2022060873A1 (en) * 2020-09-16 2022-03-24 Wayfair Llc Techniques for virtual visualization of a product in a physical scene

Patent Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200005538A1 (en) * 2018-06-29 2020-01-02 Factualvr, Inc. Remote Collaboration Methods and Systems
CN110782524A (zh) * 2019-10-25 2020-02-11 重庆邮电大学 基于全景图的室内三维重建方法
CN114756919A (zh) * 2021-12-29 2022-07-15 每平每屋(上海)科技有限公司 数据处理方法、家装设计方法、设备及存储介质
CN114529566A (zh) * 2021-12-30 2022-05-24 北京城市网邻信息技术有限公司 图像处理方法、装置、设备及存储介质
CN114549765A (zh) * 2022-02-28 2022-05-27 北京京东尚科信息技术有限公司 三维重建方法及装置、计算机可存储介质
CN115439607A (zh) * 2022-09-01 2022-12-06 中国民用航空总局第二研究所 一种三维重建方法、装置、电子设备及存储介质
CN115359192A (zh) * 2022-10-14 2022-11-18 阿里巴巴(中国)有限公司 三维重建与商品信息处理方法、装置、设备及存储介质

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20240290061A1 (en) * 2023-02-28 2024-08-29 Lemon Inc. Pixel perspective estimation and refinement in an image
US12586340B2 (en) * 2023-02-28 2026-03-24 Lemon Inc. Pixel perspective estimation and refinement in an image

Also Published As

Publication number Publication date
CN115937422A (zh) 2023-04-07

Similar Documents

Publication Publication Date Title
US11657419B2 (en) Systems and methods for building a virtual representation of a location
CN110310175B (zh) 用于移动增强现实的系统和方法
US10977818B2 (en) Machine learning based model localization system
JP7569439B2 (ja) クロスリアリティシステムにおけるスケーラブル3次元オブジェクト認識
Li et al. Database‐assisted object retrieval for real‐time 3d reconstruction
US20210133850A1 (en) Machine learning predictions of recommended products in augmented reality environments
CN103975365B (zh) 用于俘获和移动真实世界对象的3d模型和真实比例元数据的方法和系统
CN113689578B (zh) 一种人体数据集生成方法及装置
CN112950759B (zh) 基于房屋全景图的三维房屋模型构建方法及装置
CN112927353A (zh) 基于二维目标检测和模型对齐的三维场景重建方法、存储介质及终端
US20150029219A1 (en) Information processing apparatus, displaying method and storage medium
WO2020024569A1 (zh) 动态生成人脸三维模型的方法、装置、电子设备
WO2024124653A1 (zh) 三维场景重建、在线家装与商品获取方法、设备及介质
AU2023232170B2 (en) System and method of object detection and interactive 3d models
US12560450B2 (en) Method and server for generating spatial map
Xiao et al. Coupling point cloud completion and surface connectivity relation inference for 3D modeling of indoor building environments
Sankar et al. In situ CAD capture
KR20240049102A (ko) 3차원(3d) 모델을 생성하는 방법 및 전자 장치
Maghoumi et al. Gemsketch: Interactive image-guided geometry extraction from point clouds
CN112070175A (zh) 视觉里程计方法、装置、电子设备及存储介质
CN119229385B (zh) 一种多角度融合的厅店场景热力图生成方法和装置
KR20240049096A (ko) 공간 맵을 생성하는 방법 및 서버
CN115147520A (zh) 基于视觉语义驱动虚拟人物的方法及设备
De Sorbier et al. Stereoscopic augmented reality with pseudo-realistic global illumination effects
CN120655815A (zh) 生成三维空间结构的方法及电子设备

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23901873

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 23901873

Country of ref document: EP

Kind code of ref document: A1