CN116645497B - Method and system for acquiring plane material database based on neural network - Google Patents
Method and system for acquiring plane material database based on neural networkInfo
- Publication number
- CN116645497B CN116645497B CN202310783568.4A CN202310783568A CN116645497B CN 116645497 B CN116645497 B CN 116645497B CN 202310783568 A CN202310783568 A CN 202310783568A CN 116645497 B CN116645497 B CN 116645497B
- Authority
- CN
- China
- Prior art keywords
- network
- neural network
- acquisition
- camera
- vector
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Active
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/10—Image acquisition
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/50—Information retrieval; Database structures therefor; File system structures therefor of still image data
- G06F16/51—Indexing; Data structures therefor; Storage structures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/77—Processing image or video features in feature spaces; using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]; Blind source separation
- G06V10/774—Generating sets of training patterns; Bootstrap methods, e.g. bagging or boosting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
- Y02P—CLIMATE CHANGE MITIGATION TECHNOLOGIES IN THE PRODUCTION OR PROCESSING OF GOODS
- Y02P90/00—Enabling technologies with a potential contribution to greenhouse gas [GHG] emissions mitigation
- Y02P90/30—Computing systems specially adapted for manufacturing
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Evolutionary Computation (AREA)
- Software Systems (AREA)
- Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Computing Systems (AREA)
- General Health & Medical Sciences (AREA)
- General Engineering & Computer Science (AREA)
- Databases & Information Systems (AREA)
- Multimedia (AREA)
- Data Mining & Analysis (AREA)
- Medical Informatics (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Life Sciences & Earth Sciences (AREA)
- Biomedical Technology (AREA)
- Biophysics (AREA)
- Computational Linguistics (AREA)
- Molecular Biology (AREA)
- Mathematical Physics (AREA)
- Image Analysis (AREA)
Abstract
The invention discloses a method and a system for acquiring a plane material database based on a neural network, wherein the method comprises a training stage, an acquisition stage and a reconstruction stage, the method designs the neural network, and the neural network comprises an optimized illumination pattern part, a gate network part, an expert network part and a nonlinear mapping network part, the gate network can adaptively select the optimal hidden vector expression of the reflection attribute of the expert network prediction material space independence through the information in a target material photo shot under illumination of an illumination pattern, the hidden vector is further optimized under the constraint of the photo shot under a group of other illumination patterns, finally the hidden vector is restored into a high-dimensional Lumitexel vector through the nonlinear mapping network, and the vector is fitted into a BRDF model, and the parameters of the vector are stored as a texture map. The method can robustly, high-quality and high-efficiency acquire the near-plane anisotropy SVBRDFs.
Description
Technical Field
The invention relates to a method and a system for acquiring a plane material database based on a neural network, belonging to the fields of computer graphics and computer vision.
Background
The high quality of the material appearance shows the complex physical interaction between the object and the light, and is usually expressed by a six-dimensional bi-directional reflection distribution function (SVBRDF) which varies with space, which is one of the core problems of computer graphics and vision, and has important application in the fields of cultural heritage, electronic commerce, computer games, film production and the like. In computer graphics, high quality digitized material appearance can realistically render a complex physical appearance of an object that varies with position, illumination, and viewing angle. On the other hand, in the field of computer vision, the appearance of texture may help machines better understand the real world from images.
Over the past decades, as the demand for accurate, diverse digital appearances has grown in academia and industry, tremendous efforts have been made to build high value texture reflection databases. Some pioneering work was therefore born, including actual acquisition BRDFs/SVBRDFs and synthesis SVBRDFs. However, the number of collected material reflectance data available for the current disclosure remains limited, which has hampered the development of related research in this data-driven era. For example, ,Wojciech Matusik,Hanspeter Pfister,Matt Brand,and Leonard McMillan.2003.A Data-driven Reflectance Model.ACM Trans.Graph.22,3(July 2003),759–769. discloses 100 measured isotropy BRDFs data, which are still used in many studies today after 20 years of their publication.
The main reason for this is the technical difficulty of acquiring large-scale datasets with the prior art, the intensive sampling of a single SVBRDF six-dimensional physical space despite the high acquisition quality, is time consuming and therefore cannot be extended to the creation of a large database. The method based on strong priori is used for reconstructing quality to obtain acquisition efficiency, and when the priori is not established, the quality of the result cannot be ensured. For illumination multiplexing techniques with high quality and efficiency, quality is not satisfactory even when some challenging material appearances, such as wire drawn metal and polished wood, are reconstructed, even for the most advanced work.
Disclosure of Invention
The invention aims to provide a method for collecting near-plane high-dimensional materials on a large scale aiming at the defects of the prior art. The method can robustly, high-quality and high-efficiency acquire the near-plane anisotropy SVBRDFs.
The invention designs a neural network which comprises an optimized illumination pattern part, a gate network part, an expert network part and a nonlinear mapping network part. Through the information in the target material photo taken under illumination of the illumination pattern, the gate network can adaptively select the best expert network to predict the hidden vector expression of the spatially independent reflection attribute of the material, the hidden vector is further optimized under the constraint of the photo taken under a group of other illumination patterns, finally the hidden vector is restored to a high-dimensional Lumitexel vector through a nonlinear mapping network, the vector is fitted into the BRDF model, and the parameters of the vector are stored as texture maps.
The method can be used for shooting target materials at a single view angle or multiple view angles, the photo alignment method for multiple view angles is not limited to a specific method, other methods for carrying out pixel matching on photos shot by two cameras at two view angles are also suitable, the method for outputting probability of a gate network in the method is not limited to a specific method, other methods for outputting a group of sum probabilities of 1 are also suitable, such as a Softmax function, the illumination condition of the photos for optimizing hidden vectors in the method is not limited to a specific illumination pattern, such as a linear light source pattern adopted by the method, other illumination patterns which can be illuminated by acquisition equipment are also suitable, such as a point light source pattern, a surface light source pattern and the like, a neural network is not limited to a fully connected network, the expression of the object material attribute is not limited to a Lumitexel vector, the BRDF model fitted to the BRDF model is not limited to a GGX BRDF model, the fitting method is not limited to the differentiable nerve fitting used in the method, and other fitting methods are also suitable, such as a traditional numerical method L-BFGS-B.
According to a first aspect of the present specification, there is provided a method for acquiring a planar texture database based on a neural network, the method comprising a training phase, an acquisition phase and a reconstruction phase;
The training phase comprises the following steps:
(1) Acquiring parameters of acquisition equipment, and generating acquisition results of an analog camera as training data;
(2) Training a neural network using the generated training data, the neural network being characterized by:
The input of the neural network is Lumitexel vectors under all observation directions;
The first part of the neural network is a linear full-connection layer and is used for simulating an illumination pattern used in actual acquisition and converting Lumitexel vectors into acquisition results of corresponding cameras;
The second part of the neural network comprises a gate network and a plurality of expert networks, the gate network takes the acquisition results of all cameras as input, outputs a group of probability that the expert network is selected, and each expert network takes the acquisition results of all cameras as input to predict hidden vector representation of materials in an implicit space;
the third part of the neural network is a nonlinear mapping network and is used for recovering high-dimensional material information according to the hidden vector;
the acquisition phase comprises the following steps:
(1) The neural network illumination pattern acquisition, wherein the acquisition equipment sequentially irradiates a target near-plane sample according to a group of illumination patterns, and all cameras respectively acquire a group of photos;
(2) The material optimization illumination pattern acquisition, wherein the acquisition equipment sequentially irradiates a target near-plane sample according to a group of preset linear illumination patterns to obtain a group of photos shot by a main camera;
the reconstruction phase comprises the steps of:
(1) According to the photos acquired in the acquisition stage (1), taking the results of samples acquired by all cameras under different illumination patterns as the input of a gate network of a neural network, outputting a group of probabilities of selecting an expert network by the gate network, and predicting hidden vectors of high-dimensional material information by taking the acquisition results of all cameras as the input of the expert network with the highest probability;
(2) Material optimization, namely according to the photo acquired in the acquisition stage (2), taking a predicted hidden vector as an initial value, recovering the predicted hidden vector into a Lumitexel vector through a nonlinear mapping network, according to the linear relation between Lumitexel and the luminous intensity of a light source, simulating an acquisition process by using vector multiplication, taking the result of a sample acquired by a main camera under irradiation of different linear illumination patterns as a target, and optimizing the hidden vector;
(3) And training a differential rendering neural network for each near-plane sample to obtain GGX BRDF model parameters and a local coordinate system of the sample as a material acquisition result of the near-plane sample.
Further, the acquisition device is provided with at least one camera facing the target near-plane material, and when the cameras are a plurality of cameras, the camera facing the target near-plane material is set as a main camera, and the other cameras are secondary cameras.
Further, each of the values Lumitexel describes the reflected light intensity of the sample point for the incident light from each light source along a certain viewing direction, lumitexel is linear with the light source luminous intensity, modeled with a linear fully connected layer.
Further, the gate network of the neural network is formed by a plurality of sub-gate networks, the number of the sub-gate networks is related to the method of outputting probabilities by the gate network, for the mode of encoding the expert network index in a binary coding form, the gate network is formed by log 2 n one-bit sub-gate networks, n is the number of the expert networks, one sub-gate network takes the acquisition results of all cameras as input, g (b) is outputted, the probability of b bit of binary representing the expert network index is 1, and the probability of selecting the expert network with index a is formalized as follows:
wherein a b represents the b-th position of a.
Further, the neural network has a loss function as follows:
The Lumitexel vector of the sampling point under the observation angle of the main camera is m p, the hidden vector output by the expert network with the index a is m a after passing through the nonlinear mapping network, and the Loss function Loss of the material characteristic part is expressed as follows:
wherein n is the number of expert networks, pr(s) is the probability of the expert network with index a being selected, As a nonlinear mapping function, acting on each dimension of the vector, l represents the light source l;
After training, according to the parameters of the linear full-connection layer, an illumination matrix is obtained through transformation to serve as an illumination pattern.
Further, the acquisition stage is used for the situation that the total number of cameras is larger than 1, and the method further comprises a mapping acquisition step, wherein the acquisition equipment irradiates the target near-plane sample with full white light, each camera respectively obtains a low dynamic range photo, and the photo of the main camera and the photo of a certain secondary camera to be aligned are used as input to obtain the corresponding relation of pixels of the near-plane sample in the two photos.
Further, the mapping acquisition step specifically includes:
According to the augmented reality marks on the photos shot by the cameras, calculating homography matrixes from each secondary camera photo to the primary camera photo, and converting the secondary camera photos to the primary camera observation angles;
calculating dense SIFT feature vectors of the photos shot by the primary camera and the photos shot by the secondary camera after transformation;
and taking the photo shot by the main camera as a reference, performing block matching calculation according to the dense SIFT feature vectors of the two photos, and finding out the corresponding pixel of each pixel on the main camera photo on the secondary camera photo.
Further, in the material fitting step, the differentiable rendering neural network is characterized as follows:
for each effective texture coordinate of the near-plane sample, the input of the network is a high-dimensional neural parameter vector which is an optimizable variable;
the network consists of a plurality of layers of nonlinear fitting networks represented by full connection layers;
the output of the network is GGX BRDF model parameters and a local coordinate system, the vector Lumitexel rendered by the GGX BRDF model parameters and the local coordinate system is used for optimizing the high-dimensional neural parameter vector according to the error of the Lumitexel vector output by the nonlinear mapping network, and the GGX BRDF model parameters and the local coordinate system after the optimization are used as the material acquisition result of the near-plane sample.
The acquisition stage further comprises the step that the acquisition equipment irradiates the target near-plane sample from the bottom according to a preset illumination pattern, a photo shot by the main camera is obtained, and the transparency of the near-plane sample is calculated in the reconstruction stage according to the obtained photo.
According to a second aspect of the present disclosure, there is provided a system for collecting a planar material database implemented by the above method, including:
The preparation module is used for acquiring parameters of acquisition equipment, generating acquisition results of the simulation camera as training data, and training the neural network by using the generated training data;
The acquisition module is used for acquiring the neural network illumination pattern and optimizing the material quality;
And the recovery module takes the results of the samples acquired by all cameras under the irradiation of different illumination patterns as input, loads a trained neural network, predicts the hidden vectors of the material characteristics, optimizes the hidden vectors according to the photos for tuning, and fits the coordinate system and the material parameters by using the differential rendering neural network.
The method has the beneficial effect that the method can robustly, high-quality and high-efficiency acquire the near-plane anisotropy SVBRDFs.
Drawings
FIG. 1 is a three-dimensional schematic diagram of an acquisition device in an embodiment of the invention;
FIG. 2 is an exterior elevation view of a collection device in accordance with an embodiment of the present invention;
FIG. 3 is an external side view of a collection device in an embodiment of the invention;
FIG. 4 is an interior side view of a collection device in an embodiment of the invention;
FIG. 5 is an expanded view of an acquisition device in an embodiment of the invention;
FIG. 6 is a flowchart of an acquisition method according to an embodiment of the present invention;
FIG. 7 is a schematic diagram of a neural network according to an embodiment of the present invention;
FIG. 8 is a graph of illumination patterns of a computed map photograph, with gray values representing luminous intensity, according to an embodiment of the present invention;
FIG. 9 is a partial illustration of an illumination pattern obtained according to an embodiment of the present invention, with gray values representing luminous intensity;
FIG. 10 is a partial illustration of a linear illumination pattern according to an embodiment of the present invention, with gray scale values representing luminous intensity;
FIG. 11 is a graph showing the illumination pattern required for calculating the transparency according to the embodiment of the present invention, wherein the gray value represents the luminous intensity;
FIG. 12 is a Lumitexel vector result of system recovery using an embodiment of the present invention;
FIG. 13 is a graph showing the results of material properties of a sample object recovered using the system of an embodiment of the present invention.
Detailed Description
The present invention will be described in detail below with reference to the accompanying drawings, in order to make the objects, technical solutions and advantages of the present invention more apparent.
The method for acquiring the near-plane anisotropy SVBRDFs in a large-scale, robust, high-quality and high-efficiency mode provided by the invention can be concretely implemented as the following steps:
1.a training phase comprising the steps of:
1. generating training data
The acquisition equipment is provided with at least one camera facing the target near-plane material, when the camera is a plurality of cameras, the camera facing the target near-plane material is set as a main camera, the other cameras are secondary cameras, and parameters of the acquisition equipment are acquired, wherein the parameters comprise the distance and angle from a light source to a sampling space origin, the characteristic curve of the light source, the distance and angle from the camera to the sampling space origin, and internal parameters and external parameters of the camera. And generating an acquisition result of the simulation actual camera by using the parameters as training data. In this embodiment, the rendering model adopted when generating training data is a GGX model, and the generation formula is as follows:
′′′′′′
Wherein, f r(ωi,ωo; P) is a four-dimensional reflection function related to omega i,ωo, omega i represents the incident light direction under the world coordinate system, omega o represents the emergent light direction under the world coordinate system, omega i is the incident direction under the local coordinate system, omega o is the emergent direction under the local coordinate system, and omega h is a half-way vector under the local coordinate system. And P comprises parameter information of the sampling points, wherein the parameter information comprises material parameters n, t and alpha x,αy,ρd,ρs of the sampling points, n represents a normal vector under a world coordinate system, t represents an x-axis direction of a local coordinate system of the sampling points under the world coordinate system, and n and t are used for converting an incident direction and an emergent direction from the world coordinate system to the local coordinate system. Alpha x,αy denotes a roughness coefficient, ρ d denotes a diffuse reflectance, ρ s denotes a specular reflectance, ρ d and ρ s are both scalar quantities in a single channel, three scalar quantities in a color case, respectively AndD GGX is a micro-surface distribution term, F is a fresnel term, and G GGX represents a shading coefficient function.
2. Using the generated training data, the neural network shown in fig. 7 is trained. The neural network is characterized as follows:
(1) The relationship between the reflection function f r and the light intensity of each light source can be described as:
Wherein I represents the luminous information of each light source l, including the spatial position x l of the light source l, the normal vector n l of the light source l, the luminous intensity I (l) of the light source l, P comprises the parameter information of the sampling point P, including the spatial position x p of the sampling point, the material parameter n p,t,αx,αy,ρd,ρs.Ψ(xl,) describes the light intensity distribution of the light source l in different incidence directions, V represents the binary function of x l for the visibility of x p, and (-) + is the dot product operation of two vectors, and the negative value is truncated to 0.f r(ω′i;ω′o, P) is a two-dimensional reflection function for ω 'i when ω' o is fixed.
The input of the neural network is Lumitexel vector sampled in the observation direction of the main camera, which is denoted as m p (l; P) and Lumitexel vector sampled in the observation direction of the secondary camera, which is denoted as m s (l; P), wherein each value describes the reflection light intensity of the sampling point on the incident light from each light source along a certain observation direction, lumitexel is in linear relation with the light source luminous intensity, and the linear full-connection layer is used for simulation;
Wherein, the AndThe directions of view of the primary and secondary cameras, respectively.
(2) The first part of the neural network is a linear full-connection layer which is used for simulating an illumination pattern used in actual acquisition and converting Lumitexel vectors into acquisition results of corresponding cameras, wherein the linear full-connection layer only consists of a parameter matrix which does not comprise a nonlinear activation function, and the parameter matrix of the linear full-connection layer is obtained by training the following formula:
Wl=fW(Wraw)
The method comprises the steps of obtaining a light source, wherein W raw is a parameter to be trained, W l is an illumination matrix, the size of the illumination matrix is 1 multiplied by n l for a single-channel light source, the size of the illumination matrix is k multiplied by n l;nl for a color light source, the length of a vector m is the sampling precision of Lumitexel, k is the number of illumination patterns, and f W is a mapping used for transforming W raw, so that the generated illumination matrix can correspond to the possible luminous intensity of the light source.
The linear fully-connected layer is represented as follows:
y1=m·Wl
Where y 1 is the output of the first tier network.
(3) The second part of the neural network comprises a gate network and a plurality of expert networks, the gate network takes the acquisition results of all cameras obtained in the step (2) as input, outputs a group of probability that the expert network is selected, and each expert network takes the acquisition results of all cameras as input to predict hidden vector representation of materials in an implicit space;
The gate network may directly output a set of probabilities of being selected by the expert network through a softmax function, or construct a multi-sub-gate network structure, where in this embodiment, the gate network is formed by γ sub-gate networks, the value of γ is related to the method of outputting probabilities by the gate network, and for the way of encoding the expert network index in binary coding form, the gate network is formed by log 2 n one-bit sub-gate networks, n is the number of expert networks, in this embodiment n is 128, but not limited to 128, each sub-gate network takes the output of the first part of the neural network as input, outputs g (b), which represents the probability of b bit of binary of the expert network index being 1, and the probability of selecting the expert prediction network with index binary a is formalized as follows:
Wherein a b represents the b-th position.
Each expert network takes the acquisition results of all cameras as input, and predicts the hidden vector representation of the material in the implicit space;
wherein i is the mapping function of the ith layer network, W i is the parameter matrix of the ith layer network, b i is the offset vector of the ith layer network, y i is the output of the ith layer network, Z A and Z S respectively represent two branches of the albedo part and the shape part, and input AndFor the acquisition results of all the cameras,AndThe maximum layer numbers of the two branches are respectively, the output Z A and Z S are respectively an albedo hidden vector of 8 dimensions and a shape hidden vector of 48 dimensions, and the albedo hidden vector and the shape hidden vector of 48 dimensions are synthesized into a 56-dimension hidden vector, and expressed as follows:
Z=concat[ZA,ZS]
in the above formula, Z is expressed under a single channel, and when three-channel materials are expressed, Z is expanded into the following form:
Z3c=concat[ZA_R,ZA_G,ZA_B,ZS]
The dimensions of the albedo hidden vector and the shape hidden vector are not limited to 8 and 48.
(4) The third part of the neural network is a nonlinear mapping network after the expert network, and is used for recovering high-dimensional material information according to the hidden vector, and the expression is as follows:
yi=fi(yi-1Wi+bi),nr≥i≥nl
Wherein f i is the mapping function of the ith layer network, W i is the parameter matrix of the ith layer network, b i is the offset vector of the ith layer network, y i is the output and input of the ith layer network Is the hidden vector output by the expert network,N r is the maximum number of layers of the network.
(5) The loss function of the neural network is designed as follows:
The Lumitexel vector of the sampling point under the observation angle of the main camera is m p, the hidden vector output by the expert network with index a is output by the vector m a after passing through the nonlinear mapping network, wherein the length of m a is the same as that of m p, and the loss function of the material characteristic part is expressed as follows:
Wherein, the The vector is a nonlinear mapping function, acts on each dimension of the vector, can use a logarithmic function in actual use, can also use functions of other compression ranges, and n is the number of expert networks.
3. After training, the parameters W raw of the linear full-connection layer of the network are taken out and converted by the formula W l=fW(Wraw) to be used as the illumination pattern.
2. Acquisition phase
The acquisition stage can be further divided into mapping acquisition and material acquisition, and the material acquisition can be further divided into a neural network illumination pattern acquisition stage, a material optimization illumination pattern acquisition stage and a transparency illumination pattern acquisition stage.
1. Mapping acquisition phase
Aiming at the condition that the total number of cameras is greater than 1, the acquisition equipment irradiates the target near-plane sample with full white light, each camera respectively obtains a photo with a low dynamic range, the photo is used as input, and the corresponding relation of pixels of the near-plane sample in the two photos is obtained, wherein the steps are as follows:
(1) According to the augmented reality mark ARTags on the photo shot by the camera, a homography matrix from each secondary camera photo to the primary camera photo is calculated, and the secondary camera photo is converted into the primary camera photo under the observation angle, and the conversion process can be described as follows:
H=findHomography(Imgp,Imgs)
wherein findHomography is a function of calculating a homography matrix, img p and Img s are pictures obtained by the primary camera and the secondary camera respectively, and H is the homography matrix obtained by calculation;
Img′s=warpPerspective(Imgs,H)
Wherein WARPPERSPECTIVE is the perspective transformation function, img' s represents the resulting photograph of the secondary camera transformed into the primary camera view angle.
(2) Calculating dense SIFT feature vectors of the photos shot by the primary camera and the photos shot by the secondary camera after transformation;
(3) And taking the photo shot by the main camera as a reference, performing block matching PATCHMATCH calculation according to the dense SIFT feature vectors of the two photos, and finding out the corresponding pixel of each pixel on the main camera photo on the secondary camera photo.
2. Material collection stage
(1) The method comprises the steps of acquiring a neural network illumination pattern, wherein an acquisition device irradiates a target near-plane sample according to a group of illumination patterns, and all cameras respectively acquire a group of photos;
(2) The material optimization illumination pattern acquisition, wherein the acquisition equipment sequentially irradiates a target near-plane sample according to a group of preset optimization illumination patterns to obtain a group of photos shot by a main camera;
(3) The acquisition equipment irradiates the target near-plane sample from the bottom according to a preset illumination pattern, and a photo shot by the main camera is obtained.
3. Reconstruction stage
The reconstruction stage can be subdivided into a material prediction stage, a material tuning stage, a material fitting stage and a transparency calculation stage.
1. Texture prediction stage
All cameras acquire several groups of picturesWherein p and s 1...sn represent the primary camera and all secondary cameras respectively, k represents k illumination patterns, first calculate the mapping relation of the photo of each secondary camera to the primary camera, and transform the photo intoFor effective texture coordinates on the sampling plane sample, finding the pixel values of the effective texture coordinates in all camera photos to form a vectorAnd v is used as input of a gate network in the neural network, the gate network outputs a group of probability of selecting the expert network, and the expert network with the highest probability is selected to predict hidden vectors of high-dimensional material information.
2. Quality adjusting stage
According to the photo collected in the material collection stage (2), the predicted hidden vector is used as an initial value, the predicted hidden vector is restored to Lumitexel vectors through a nonlinear mapping network, the collection process is simulated by vector multiplication according to the linear relation between Lumitexel and the luminous intensity of the light source, the result of a sample collected by a main camera under the irradiation of different linear illumination patterns is used as a target, and the hidden vector is optimized.
Specifically, according to the photo { Li p1,Lip2,…,Lipj }, j represents the j-th linear light source illumination pattern acquired in the material acquisition stage (2), taking the predicted hidden vector as an initial value, firstly converting the photo of the sample acquired by the main camera under the illumination of different linear illumination patterns into a gray photo, optimizing each effective texture coordinate u on the sampling near-plane sample, and adjusting the hidden vector of 56 dimensions of the effective texture coordinate u:
Wherein, the The pixel value corresponding to the valid texture coordinate u in the gray photograph is represented,The hidden vector predicted value of the effective texture coordinate u output by the neural network is represented, LT represents a nonlinear mapping part in the neural network, WL j represents the jth linear illumination pattern, and the optimization result is recorded as
Next, the process will be describedThree replicates were replicated and R, G, B three channels in control { Li p1,Lip2,…,Lipj } were optimized separately8-Dimensional hidden vector part representing reflectivity in the middleHidden vector portion for fixing representation shapeThe expression is as follows:
Wherein the method comprises the steps of R, G, B three-channel values respectively representing pixels corresponding to effective texture coordinates u in the photo, and recording the optimization result as
Next, the process will be describedAs part of sharing, optimizationAndThe expression is as follows:
recording the optimization result as Synthesizing a hidden vector representing three channels:
3. Material fitting stage
The optimized hidden vector can recover high-dimensional material information Lumitexel through a nonlinear mapping network, and for each near-plane sample, a differential rendering neural network is trained to obtain GGX BRDF model parameters of the sample, and the differential rendering neural network is characterized in that:
(1) For each effective texture coordinate of a near-plane sample, the input of the network is a high-dimensional neural parameter vector which is an optimizable variable;
(2) The network consists of a plurality of layers of nonlinear fitting networks represented by full connection layers;
(3) The network output is GGX BRDF model parameters and a local coordinate system, the high-dimensional neural parameter vector is optimized according to Lumitexel vectors rendered by the GGX BRDF model parameters and the local coordinate system and errors of Lumitexel vectors output by the nonlinear mapping network, the GGX BRDF model parameters and the local coordinate system after the optimization are used as material acquisition results of near-plane samples, and specifically, the loss function of the neural network is designed as follows:
Wherein, the The three single channels Lumitexel are synthesized into three channels Lumitexel, θ is the high-dimensional nerve parameter input of the network, G is a nonlinear fitting network, G (θ) outputs GGX BRDF model parameters and local coordinate system parameters, R is a traditional rendering equation, G (θ) is rendered into three channels Lumitexel under the main camera view angle, errors of the network propagate through gradients, parameters in θ and G can be optimized, and when the network converges, G (θ) is saved as a fitting result to be a texture map.
4. Transparency calculation stage
And (3) calculating the transparency of the near-plane sample according to the photo acquired in the material acquisition stage (3).
In particular, methods already disclosed in the field of use (Andrew Gardner,Chris Tchou,Tim Hawkins,and Paul Debevec.2003.Linear light source reflectometry.ACM Trans.Graph.22,3(2003),749–
758. ) And calculating the transparency of the near-plane sample according to two photos, wherein the illumination of the two photos is the bright white light of the light source at the bottom of the acquisition device, one photo is a photo which is taken in advance and is not placed on the sample placing table, the other photo is a photo which is taken after the sample is placed, and the transparency of each effective texture coordinate on the sample is determined by the quotient of the pixel on the photo which is placed on the sample and the pixel on the photo which is not placed on the sample.
Specifically, in the material prediction, material tuning and transparency calculation process in the reconstruction stage, firstly, flat field correction, distortion removal and color correction are performed on a photo shot by a camera.
An example of a specific system of the collecting device is given below, as shown in fig. 1 is a three-dimensional representation of the system example, fig. 2 is an external front view of the system example, fig. 3 is an external side view of the system example, fig. 4 is an internal side view of the system example, and fig. 5 is an internal expanded view of the system example. A push-pull drawer is arranged at a position of 10cm above a lamp panel at the bottom inside the collection equipment, a platform capable of placing samples is arranged on the drawer, and when the sample is actually collected, the near-plane sample can be replaced by pulling out and pushing in the drawer. The top camera looks at the center of the sample placement stage at 90 degrees vertically, which is referred to as the primary camera, and the side camera looks at the center of the sample placement stage at 45 degrees, which is referred to as the secondary camera. LED lamp beads are densely arranged on the lamp panels of six faces of the device, wherein 4096 are arranged on the top face and the bottom face respectively, 2048 are arranged on the four side faces respectively, and 16384 are arranged on the total. The lamp beads are controlled by the FPGA, so that the luminous brightness and the luminous time can be adjusted.
An example of an acquisition system applying the method of the invention is given below, the system being generally divided into the following modules:
And the preparation module is used for providing a data set for network training, and inputting a set of BRDF parameters, the spatial positions of points and the positions of two cameras by using the GGX model to obtain two reflection conditions. The network training section uses Pytorch open source framework and trains using Adam optimizer. The network structure is shown in fig. 7, each rectangle represents a layer of neurons, and the numbers in the rectangle represent the number of neurons in that layer. The leftmost layer is the input layer and the rightmost layer is the output layer. Solid arrows between layers represent full connection.
The acquisition module is shown in fig. 1,2, 3, 4 and 5, and the specific constitution is described above.
And the recovery module is used for estimating a geometric model of the sample by using the pixel position of the sample in the photo and the calibrated drawer position, calculating the geometric model of the sample with texture coordinates by using the model, loading a trained neural network, predicting a material characteristic hidden vector for each vertex on the near-plane sample geometric model with the texture coordinates, optimizing the hidden vector according to the photo for optimization, and fitting a coordinate system and material parameters by using a differential fitting network.
Fig. 6 is a workflow of the present embodiment. Firstly, training data are generated, 2 hundred million groups of material parameters are obtained through random sampling, lumitexel corresponding to two cameras is rendered, 80% of the training set is taken as a training set, and the rest of the training set is taken as a verification set. When the network is trained, the Xavier method is used for initializing parameters, and the learning rate is 1e-4. The illumination pattern is single-channel light, and the size of the illumination matrix is (64,16384). After training, the illumination matrix is taken out and converted into illumination patterns, the parameters of each column specify the position, the luminous intensity of the light source, and fig. 7 shows a schematic diagram of the neural network structure. The next process comprises the steps of 1, illuminating the acquisition equipment lamp panel according to the illumination pattern of fig. 8, shooting an object by two cameras at the same time to obtain two shooting results, illuminating the acquisition equipment lamp panel according to the illumination pattern of fig. 9, shooting the object by the two cameras at the same time to obtain a group of shooting results, illuminating the acquisition equipment lamp panel according to the illumination pattern of fig. 10, shooting the object by the main camera to obtain a group of shooting results, and illuminating the acquisition equipment lamp panel according to the illumination pattern of fig. 11, and shooting the object by the main camera to obtain a group of shooting results. 2. The geometric model with texture coordinates was obtained using Isochart for the geometric model of the sampled planar object. 3. The method comprises the steps of calculating the mapping relation between secondary camera pixels and primary camera pixels according to two photos in an illumination pattern figure 8, converting the photos of a secondary camera in the illumination pattern figure 9 into primary camera view angles according to the mapping relation, loading a neural network, taking out pixel values of the photos of the two cameras in the illumination pattern figure 9 as input of a second part of the neural network for each vertex on a geometric model with texture coordinates, and recovering hidden vectors. 4. Optimizing the hidden vector of the network output according to the photographed picture of the main camera under the illumination pattern figure 10, and calculating the sample transparency according to the photographed picture of the main camera under the illumination pattern figure 11. 5. A differentiable fitting network is trained to fit a coordinate system for rendering and roughness, specular reflectivity, and diffuse reflectivity to each vertex on the sample.
Fig. 12 shows two Lumitexel vectors recovered from the validation set using the system described above, one on the left as m p and one on the right as the corresponding m a.
FIG. 13 shows texture property results recovered from texture appearance scanning of a sample object using the system described above, with the first row representing the sample object, respectivelyThree components, the second row representing the sample object respectivelyThe third row represents the sampled object roughness coefficient alpha x,αy, the gray value represents the numerical value, the fourth row represents the sampled sample transparency coefficient, and the gray value represents the numerical value.
The present invention is not limited to the above embodiments, and the technical effects of the present invention can be achieved by the same means, and the present invention should be considered as the scope of the present invention. Various modifications and variations are possible in the technical solution and/or in the embodiments within the scope of the invention.
Claims (9)
1. The method for acquiring the plane material database based on the neural network is characterized by comprising a training stage, an acquisition stage and a reconstruction stage;
The training phase comprises the following steps:
(1) Acquiring parameters of acquisition equipment, and generating acquisition results of an analog camera as training data;
The acquisition equipment is provided with at least one camera facing the target near-plane material, and when the cameras are a plurality of cameras, the camera facing the target near-plane material is set as a main camera, and the other cameras are secondary cameras;
(2) Training a neural network using the generated training data, the neural network being characterized by:
The input of the neural network is Lumitexel vectors under all observation directions;
The first part of the neural network is a linear full-connection layer and is used for simulating an illumination pattern used in actual acquisition and converting Lumitexel vectors into acquisition results of corresponding cameras;
The second part of the neural network comprises a gate network and a plurality of expert networks, the gate network takes the acquisition results of all cameras as input, outputs a group of probability that the expert network is selected, and each expert network takes the acquisition results of all cameras as input to predict hidden vector representation of materials in an implicit space;
the third part of the neural network is a nonlinear mapping network and is used for recovering high-dimensional material information according to the hidden vector;
the acquisition phase comprises the following steps:
(1) The neural network illumination pattern acquisition, wherein the acquisition equipment sequentially irradiates a target near-plane sample according to a group of illumination patterns, and all cameras respectively acquire a group of photos;
(2) The material optimization illumination pattern acquisition, wherein the acquisition equipment sequentially irradiates a target near-plane sample according to a group of preset linear illumination patterns to obtain a group of photos shot by a main camera;
the reconstruction phase comprises the steps of:
(1) Material prediction, namely acquiring a plurality of groups of photos by all cameras
Wherein p and s 1...sn represent the primary camera and all secondary cameras respectively, k represents the kth illumination pattern, first calculate the mapping relationship of each secondary camera photo to the primary camera, transform the photo intoFor effective texture coordinates on the sampling plane sample, finding the pixel values of the effective texture coordinates in all camera photos to form a vectorTaking v as input of a gate network in the neural network, outputting a group of probability of selecting the expert network by the gate network, taking all camera acquisition results as input by the expert network with the maximum probability, and predicting to obtain hidden vectors of high-dimensional material information;
(2) Material optimization, namely according to the photo acquired in the acquisition stage (2), taking a predicted hidden vector as an initial value, recovering the predicted hidden vector into a Lumitexel vector through a nonlinear mapping network, according to the linear relation between Lumitexel and the luminous intensity of a light source, simulating an acquisition process by using vector multiplication, taking the result of a sample acquired by a main camera under irradiation of different linear illumination patterns as a target, and optimizing the hidden vector;
(3) And training a differential rendering neural network for each near-plane sample to obtain GGX BRDF model parameters and a local coordinate system of the sample as a material acquisition result of the near-plane sample.
2. The method of claim 1, wherein each value Lumitexel describes the reflected light intensity of the sample point for the incident light from each light source along a certain observation direction, lumitexel is linear with the light source luminous intensity, and the linear full-connection layer is used for simulation.
3. The method for collecting a planar texture database based on a neural network according to claim 1, wherein the gate network of the neural network is composed of a plurality of sub-gate networks, the number of the sub-gate networks is related to the method for outputting probabilities by the gate network, for encoding the expert network index in a binary coding form, the gate network is composed of log 2 n one-bit sub-gate networks, n is the number of the expert networks, one sub-gate network takes the collection results of all cameras as input, outputs g (b), represents the probability that the b-th bit of the binary of the expert network index is 1, and the probability that the expert network with index a is selected is formalized as follows:
wherein a b represents the b-th position of a.
4. The method for collecting a planar texture database based on a neural network according to claim 1, wherein a loss function of the neural network is as follows:
The Lumitexel vector of the sampling point under the observation angle of the main camera is m p, the hidden vector output by the expert network with the index a is m a after passing through the nonlinear mapping network, and the Loss function Loss of the material characteristic part is expressed as follows:
Wherein n is the number of expert networks, pr (a) is the probability of the expert network with index a being selected, As a nonlinear mapping function, acting on each dimension of the vector, l represents the light source l;
After training, according to the parameters of the linear full-connection layer, an illumination matrix is obtained through transformation to serve as an illumination pattern.
5. The method for acquiring the planar texture database based on the neural network according to claim 1, wherein the acquisition stage is used for acquiring the target near-plane sample by the acquisition device in a mapping way according to the situation that the total number of cameras is larger than 1, wherein each camera respectively acquires a low dynamic range photo, and the photo of the main camera and the photo of a certain secondary camera to be aligned are taken as inputs to acquire the corresponding relation of pixels of the near-plane sample in the two photos.
6. The method for collecting a planar texture database based on a neural network according to claim 5, wherein the mapping collecting step specifically comprises:
According to the augmented reality marks on the photos shot by the cameras, calculating homography matrixes from each secondary camera photo to the primary camera photo, and converting the secondary camera photos to the primary camera observation angles;
calculating dense SIFT feature vectors of the photos shot by the primary camera and the photos shot by the secondary camera after transformation;
and taking the photo shot by the main camera as a reference, performing block matching calculation according to the dense SIFT feature vectors of the two photos, and finding out the corresponding pixel of each pixel on the main camera photo on the secondary camera photo.
7. The method for collecting planar texture database based on neural network according to claim 1, wherein in the texture fitting step, the differentiable rendering neural network is characterized as follows:
for each effective texture coordinate of the near-plane sample, the input of the network is a high-dimensional neural parameter vector which is an optimizable variable;
the network consists of a plurality of layers of nonlinear fitting networks represented by full connection layers;
the output of the network is GGX BRDF model parameters and a local coordinate system, the vector Lumitexel rendered by the GGX BRDF model parameters and the local coordinate system is used for optimizing the high-dimensional neural parameter vector according to the error of the Lumitexel vector output by the nonlinear mapping network, and the GGX BRDF model parameters and the local coordinate system after the optimization are used as the material acquisition result of the near-plane sample.
8. The method according to claim 1, wherein the collecting step further comprises the step of irradiating the target near-plane sample from the bottom according to a preset illumination pattern by the collecting device to obtain a photograph taken by the main camera, and calculating the transparency of the near-plane sample in the reconstructing step according to the obtained photograph.
9. A system for collecting a planar texture database implemented according to the method of any one of claims 1-8, the system comprising:
The preparation module is used for acquiring parameters of acquisition equipment, generating acquisition results of the simulation camera as training data, and training the neural network by using the generated training data;
The acquisition module is used for acquiring the neural network illumination pattern and optimizing the material quality;
And the recovery module takes the results of the samples acquired by all cameras under the irradiation of different illumination patterns as input, loads a trained neural network, predicts the hidden vectors of the material characteristics, optimizes the hidden vectors according to the photos for tuning, and fits the coordinate system and the material parameters by using the differential rendering neural network.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310783568.4A CN116645497B (en) | 2023-06-29 | 2023-06-29 | Method and system for acquiring plane material database based on neural network |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310783568.4A CN116645497B (en) | 2023-06-29 | 2023-06-29 | Method and system for acquiring plane material database based on neural network |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| CN116645497A CN116645497A (en) | 2023-08-25 |
| CN116645497B true CN116645497B (en) | 2025-08-19 |
Family
ID=87619812
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| CN202310783568.4A Active CN116645497B (en) | 2023-06-29 | 2023-06-29 | Method and system for acquiring plane material database based on neural network |
Country Status (1)
| Country | Link |
|---|---|
| CN (1) | CN116645497B (en) |
Families Citing this family (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN119152235A (en) * | 2024-09-13 | 2024-12-17 | 嘉杰科技有限公司 | Unmanned ship and unmanned plane based collaborative mapping control method and system |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110570503A (en) * | 2019-09-03 | 2019-12-13 | 浙江大学 | Method for acquiring normal vector, geometry and material of three-dimensional object based on neural network |
| CN112634156A (en) * | 2020-12-22 | 2021-04-09 | 浙江大学 | Method for estimating material reflection parameter based on portable equipment collected image |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US9299188B2 (en) * | 2013-08-08 | 2016-03-29 | Adobe Systems Incorporated | Automatic geometry and lighting inference for realistic image editing |
| CN115221396B (en) * | 2021-04-21 | 2026-03-20 | 腾讯科技(深圳)有限公司 | AI-based information recommendation methods, devices, and electronic equipment |
| CN113362466B (en) * | 2021-06-07 | 2022-06-21 | 浙江大学 | Free type collection method for high-dimensional material |
| CN115600635A (en) * | 2021-07-08 | 2023-01-13 | 华为技术有限公司(Cn) | Training method of neural network model, and data processing method and device |
| CN114512114B (en) * | 2021-12-30 | 2025-05-09 | 浙江大学 | Acoustic model post-processing method, server and readable memory based on probability diffusion model |
| CN114663377A (en) * | 2022-03-16 | 2022-06-24 | 广东时谛智能科技有限公司 | Texture SVBRDF (singular value decomposition broadcast distribution function) acquisition method and system based on deep learning |
-
2023
- 2023-06-29 CN CN202310783568.4A patent/CN116645497B/en active Active
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110570503A (en) * | 2019-09-03 | 2019-12-13 | 浙江大学 | Method for acquiring normal vector, geometry and material of three-dimensional object based on neural network |
| CN112634156A (en) * | 2020-12-22 | 2021-04-09 | 浙江大学 | Method for estimating material reflection parameter based on portable equipment collected image |
Also Published As
| Publication number | Publication date |
|---|---|
| CN116645497A (en) | 2023-08-25 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN117274760B (en) | Infrared and visible light image fusion method based on multi-scale mixed converter | |
| CN111462120B (en) | Defect detection method, device, medium and equipment based on semantic segmentation model | |
| CN110570503B (en) | A method for obtaining normal vector, geometry and material of 3D object based on neural network | |
| CN118429284B (en) | Industrial appearance defect detection method and equipment based on photometric stereo method and diffusion model | |
| CN116958420A (en) | A high-precision modeling method for the three-dimensional face of a digital human teacher | |
| CN111507357B (en) | Defect detection semantic segmentation model modeling method, device, medium and equipment | |
| CN118097662B (en) | A Pap smear cervical cell image classification method based on CNN-SPPF and ViT | |
| CN116374114B (en) | A ship resistance prediction method based on image learning | |
| CN115937807A (en) | Method and system for identifying and detecting road defects | |
| CN108985333A (en) | A kind of material acquisition methods neural network based and system | |
| CN110188621B (en) | Three-dimensional facial expression recognition method based on SSF-IL-CNN | |
| CN116645497A (en) | A method and system for collecting plane material database based on neural network | |
| CN120198764B (en) | Multi-task inverse imaging method, system, terminal and readable storage medium based on expert mixture cooperative diffusion operator learning | |
| CN119000565B (en) | A spectral reflectance image acquisition method and system based on intrinsic decomposition | |
| CN119338881B (en) | Three-dimensional object size measuring method based on depth vision camera | |
| CN121236525A (en) | Flotation foam image feature extraction method based on self-adaptive multi-scale weighted fusion | |
| CN120976475A (en) | A 3D model acquisition and repair system and method | |
| CN120471853A (en) | Product appearance defect detection method and system | |
| CN113822825A (en) | Optical building target three-dimensional reconstruction method based on 3D-R2N2 | |
| CN120107316A (en) | Generalizable radiation field representation method based on index optimization in weak light environment of mine | |
| CN116597273B (en) | Multi-scale encoding and decoding essential image decomposition network, method and application based on self-attention | |
| CN118196577A (en) | A method, device, terminal equipment and medium for completing depth image of transparent object | |
| CN113326924B (en) | Photometric localization method of key targets in sparse images based on deep neural network | |
| CN114663377A (en) | Texture SVBRDF (singular value decomposition broadcast distribution function) acquisition method and system based on deep learning | |
| CN119295846B (en) | A PBAT degradation film material type identification method and system based on image processing technology |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PB01 | Publication | ||
| PB01 | Publication | ||
| SE01 | Entry into force of request for substantive examination | ||
| SE01 | Entry into force of request for substantive examination | ||
| GR01 | Patent grant | ||
| GR01 | Patent grant |