EP4634936A1 - Reinforcement learning for final setups and intermediate staging in clear tray aligners - Google Patents

Reinforcement learning for final setups and intermediate staging in clear tray aligners

Info

Publication number
EP4634936A1
EP4634936A1 EP23828816.1A EP23828816A EP4634936A1 EP 4634936 A1 EP4634936 A1 EP 4634936A1 EP 23828816 A EP23828816 A EP 23828816A EP 4634936 A1 EP4634936 A1 EP 4634936A1
Authority
EP
European Patent Office
Prior art keywords
setups
tooth
oral care
representation
implementations
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP23828816.1A
Other languages
German (de)
French (fr)
Inventor
Mariah Sonja Pereira Penha
Jonathan D. Gandrud
Seyed Amir Hossein Hosseini
Francis J. T. YATES
Haleh HAGH-SHENAS
Nitsan BEN-GAL NGUYEN
Haozhu Wang
Alireza SADEGHI
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Solventum Intellectual Properties Co
Original Assignee
Solventum Intellectual Properties Co
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Solventum Intellectual Properties Co filed Critical Solventum Intellectual Properties Co
Publication of EP4634936A1 publication Critical patent/EP4634936A1/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/30ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for calculating health indices; for individual health risk assessment
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16HHEALTHCARE INFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR THE HANDLING OR PROCESSING OF MEDICAL OR HEALTHCARE DATA
    • G16H50/00ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics
    • G16H50/20ICT specially adapted for medical diagnosis, medical simulation or medical data mining; ICT specially adapted for detecting, monitoring or modelling epidemics or pandemics for computer-aided diagnosis, e.g. based on medical expert systems

Definitions

  • Patent Applications is incorporated herein by reference: 63/432,627; 63/366,492; 63/366,495; 63/352,850; 63/366,490; 63/366,494; 63/370,160; 63/366,507; 63/352,877; 63/366,514; 63/366,498; 63/366,514; and 63/264,914.
  • This disclosure relates to configurations and training of neural networks to improve the accuracy of automatically generated clear tray aligner (CT A) devices used in orthodontic treatments.
  • CT A clear tray aligner
  • the present disclosure describes systems and techniques for training and using one or more machine learning models, such as neural networks, to produce intermediate stages and final setups for CTAs, in a manner which is customized to the treatment needs of the patient.
  • a neural network is termed herein as a “setups prediction neural network” or simply a “setups prediction model.”
  • One or more neural networks may be trained using reinforcement learning (RL) to implement a decision-making agent which may generate setups predictions. These predictions may be generated through a system of rewarding desired behavior and/or punishing undesirable behavior.
  • the agent may become trained to generate actions and learn through a process of interacting with its environment, which has the advantage of enabling the setups prediction model to learn about that environment.
  • RL models for the placement of 3D oral care representations may encode learning from past rounds of training in tuples, which may be stored in a replay buffer.
  • RL is particularly well suited to the prediction of transforms for 3D oral care representations, because RL may learn from experience from a minimal dataset, and RL may correct errors that occur though training. Once an error is corrected, the error may be less likely to recur, because the knowledge of the error is captured in the replay buffer for later use.
  • a final setup is a target configuration of 3D tooth representations (such as 3D tooth meshes) such as the teeth appear at the end of treatment.
  • An intermediate setup also referred to as an “intermediate stage” or as “intermediate staging” describes a configuration of teeth during one of the several stages of treatment, after the teeth leave their maloccluded poses (e.g., positions and/or orientations) and before the teeth reach their final setup poses.
  • a final setup may be used to generate, at least in part, one or more intermediate stages. Each stage may be used in the generation of a clear tray aligner. Such aligners may incrementally move the patient's teeth from the initial or maloccluded poses to the final poses represented by the final setup.
  • a first computer-implemented method for generating setups for orthodontic alignment treatment including the steps of receiving, by one or more computer processors, a first digital representation of a patient’s teeth, using, by the one or more computer processors and to determine a prediction for one or more tooth movements for a final setup, a generator that is a machine learning model, such as comprising one or more neural networks (e.g., a 3D encoder, 3D decoder, an MLP, an encoder-decoder structure, a neural network with an attention layer or other neural networks disclosed herein) that has been initially trained to predict one or more tooth movements for a final setup, further training, by the one or more computer processors, the setups prediction model based on the using, and where the training of the setups prediction model is modified by performing operations including predicting, by the generator, one or more tooth movements for a final setup based on the first digital representation of the patient’s teeth, computing a loss function which quantifies the difference between predicted tooth movements
  • a generator that is a machine learning
  • An encoder-decoder structure may comprise at least one encoder or at least one decoder.
  • Non-limiting examples of an encoder-decoder structure include a 3D U-Net, a transformer, a pyramid encoder-decoder or an autoencoder, among others.
  • the first aspect can optionally include additional features.
  • the method can produce, by the one or more processors, an output state for the final setup.
  • the method can determine, by the one or more computer processors, a difference between the one or more predicted tooth movements and the one or more reference tooth movements.
  • the determined difference between the one or more predicted tooth movements and the one or more reference tooth movements can be used to modify the training of the generator.
  • Modifying the training of the generator can include adjusting one or more weights of the generator’s neural network.
  • the method can generate, by the one or more computer processors, one or more lists specifying mesh elements of the first digital representation of the patient’s teeth. At least one of the one or more lists can specify one or more edges in the first digital representation of the patient’s teeth.
  • At least one of the one or more lists can specify one or more polygonal faces in the digital representation of the patient’ s teeth. At least one of the one or more lists can specify one or more vertices in the first digital representation of the patient’s teeth (e.g., such as derived from a 3D mesh). At least one of the one or more lists can specify one or more points in the first digital representation of the patient’s teeth (e.g., such as derived from a 3D point cloud).
  • a 3D point cloud may, in some instances, comprise the plurality of vertices extracted from a 3D mesh.
  • At least one of the one or more lists can specify one or more voxels in the first digital representation of the patient’s teeth (e.g., such as derived from a sparse representation).
  • the method can compute, by the one or more computer processors, one or more mesh element features.
  • the one or more mesh element features can include edge endpoints, edge curvatures, edge normal vectors, edges movement vectors, edge normalized lengths, vertices, faces of associated three-dimensional representations, voxels, and combinations thereof.
  • Other mesh element features for edges are disclosed herein.
  • Mesh element features for each of vertices, points, faces and voxels are also disclosed herein.
  • the method can generate, by the one or more computer processors, a digital representation predicting the position and orientation of the patient’s teeth based on the one or more predicted tooth movements.
  • a prediction for the movement of a tooth may comprise a transform (e.g., such as one or more of an affine transformation matrix, a translation vector, a quaternion, or one or more Euler angles).
  • the setups prediction model may predict each of tooth position and tooth orientation information. In some non-limiting examples, the network may predict the orientation and position information substantially concurrently.
  • the setups prediction model may predict a setup transform for each tooth in the arch, to place each tooth in the final setup pose.
  • the method can generate, by the one or more computer processors, a digital representation of the patient’s teeth based on the one or more reference tooth movements.
  • the generator of a setups prediction model may be trained, at least in part, with the assistance of a discriminator.
  • the discriminator may determine whether a representation of the one or more tooth movements predicted by the generator is distinguishable from a representation of one or more reference tooth movements can include the steps of receiving the representation of the one or more tooth movements predicted by the generator, the representation of the one or more reference tooth movements, and the first digital representation of the patient’s teeth, comparing the representation of the one or more tooth movements predicted by the generator, the representation of the one or more reference tooth movements, wherein the comparison is based at least in part on the first digital representation of the patient’s teeth, and determining, by the one or more computer processors, a probability that the representation of the one or more tooth movements predicted by the generator is the same as the representation of one or more reference tooth movements.
  • a second computer-implemented method for generating setups for orthodontic alignment treatment pertains to intermediate staging prediction.
  • Intermediate staging of teeth from a malocclusion stage to a final stage requires determining accurate individual teeth movements in a way that teeth are not colliding with each other, the teeth move toward their final state, and the teeth follow optimal and preferably short trajectories. Because each tooth has six degrees-of-freedom and an average arch has about fourteen teeth, finding the optimal teeth trajectory from initial to final stage is a large and complex problem.
  • the second computer-implemented method is customized to the treatment needs of the patient (e.g., as specified by a clinician, which may include technician or healthcare professional) and is described including the steps of receiving, by one or more computer processors, a first digital representation of a patient’s teeth, and a representation of a final setup, using, by the one or more computer processors and to determine a prediction for one or more tooth movements for one or more intermediate stages, a generator that is a machine learning model, such as a neural network, included in a setups prediction machine learning model, such as comprising one or more neural networks (e.g., a 3D encoder, 3D decoder, a 3D U-Net, a multilayer perceptron (MLP), a transformer, an autoencoder, a pyramid encoder-decoder, and other neural networks disclosed herein), and that has been initially trained to predict one or more tooth movements for one or more intermediate stages, further training, by the one or more computer processors, the setups prediction model based
  • Techniques of this disclosure relate to the automatic generation of transformation for a three- dimensional (3D) representations of oral care data (e.g., teeth, etc.).
  • the methods involve receiving oral care data describing the current state of a treatment and providing that data as input to a reinforcement learning (RL) model.
  • RL reinforcement learning
  • Trained neural networks are then executed using the oral care data to generate a representation that specifies a predicted action associated with the treatment of the patient.
  • the methods may generate a predicted next state of the treatment based on the current state and the action.
  • the oral care data describing the current state of the treatment can include 3D representations of the oral care data (e.g., tooth meshes, appliance components, fixture model components, etc.) or transforms that modify the poses of the 3D representations of oral care data (e.g,. tooth transforms).
  • Reward values may be computed to quantify the difference between a predicted next state and a corresponding ground truth next state.
  • Loss values may be computed to quantify the difference between a predicted action and a corresponding ground truth action, in some implementations, a loss may be computed, at least in part, by a reward (i.e., a reward may be considered by a loss calculation).
  • a 3D representation of oral care data can include a 3D mesh, a 3D point cloud, a 3D voxelized representation, or a 3D surface.
  • a 3D representation of oral care data may represent teeth, appliance components, or fixture model components, among other 3D oral care representations described herein.
  • the methods can be used in digital oral care, such as orthodontic alignment treatment.
  • the trained neural network can be executed according to a policy -based or policy -free methodology, or the neural network can be trained using various paradigms such as Q-leaming, soft actor critic (SAC) learning, or generative adversarial imitation learning (GAIL).
  • SAC soft actor critic
  • GAIL generative adversarial imitation learning
  • the computing device can be deployed in a clinical context and perform the method in near real-time during a patient encounter.
  • the RL model may, in some non-limiting implementations, be trained through reinforcement learning by human feedback (RLHF). Additionally, the RL model can be used to generate designs for orthodontic appliances, or dental restoration appliances.
  • the method can also involve providing additional input data to the RL model, such as 3D geometries describing teeth, vectors containing values for computing dimensions and distances between teeth, latent vector information, position and orientation vectors, or tooth-related information.
  • the computing device includes interface hardware for receiving the 3D representation of oral care data and processing circuitry for executing the method steps.
  • methods of this disclosure may generate transformations for a three- dimensional (3D) representations of oral care data (e.g., 3D meshes of teeth, or appliance components).
  • the methods may be trained, at least in part, using reinforcement learning (RL).
  • the methods may operate, in deployment, through the use of the reinforcement learning architecture described herein.
  • RL models e.g., RL Setups for orthodontic setup generation
  • a 3D representation of oral care data is at least one of a 3D mesh, a 3D point cloud, a 3D voxelized representation, or a 3D surface.
  • a 3D mesh may represent a tooth in a dental arch, contained within a patient case.
  • an actor of the RL model may be configmed to issue one or more actions associated with moving at least one tooth represented in the 3D oral care representation to generate an orthodontic setup transform.
  • the actor of the RL model may be configmed to predict a transform associated with moving a tooth (e.g., for orthodontic setups generation), an appliance component (e.g., dental restoration appliance generation, etc.), or a fixtme model component (e.g., for fixture model generation).
  • the actor may contain one or more nemal networks.
  • the actor may be trained using on at least one of: (i) positive rewards that are issued to reward desired behaviors, or (ii) negative rewards that me issued to penalize undesired behaviors.
  • the RL models of this disclosure may, in some implementations, be executed according to a policy -based methodology, or according to a policy -free methodology.
  • the RL models may, in some implementations, be trained according to a Q-leaming paradigm.
  • the RL models may, in some implementations, be trained according to a soft actor critic (SAC) learning paradigm.
  • the RL models may, in some implementations, be trained according to a generative adversarial imitation learning (GAIL) paradigm.
  • the transformations predicted by the methods may be used for orthodontic setups generation (e.g., for alignment treatment).
  • a setup transformation may be associated with transforming one or more teeth of the patient into a final setup (or intermediate stage).
  • a final setup may describe the poses of the patient’s teeth after completion of an orthodontic treatment.
  • An intermediate stage may describe the poses of the patient’s teeth during the course of an orthodontic treatment of a patient.
  • the methods may, in some instances, be deployed at a clinical context.
  • the RL model may be trained, at least in part, through reinforcement learning by human feedback (RLHF).
  • the RL model may, in some implementations, be used to generate a design for a dental restoration appliance, or an orthodontic appliance (e.g., a clear tray aligner), among others.
  • Inputs which may be provided to the RL models of this disclosure, in some implementations, include one or more of: (i) one or more 3D geometries describing one or more teeth, (ii) one or more vectors P containing at least one value pertaining to at least one method of computing a dimension of at least one tooth, (iii) one or more vectors Q containing at least one value pertaining to at least one method of computing a distance between adjacent teeth, (iv) one or more vectors B containing latent vector information about one or more teeth, (v) one or more vectors N containing at least one value pertaining to the position of at least one tooth, (vi) one or more vectors O containing at least one value pertaining to the orientation of at least one tooth, (vii) one or more vectors R at least one of tooth name, designation, tooth type and tooth classification.
  • FIG. 1 shows a method of augmenting training data for use in training machine learning (ML) models of this disclosure.
  • FIG. 2 shows a summary of some of the setups prediction methods described herein.
  • FIG. 3 shows a method of using reinforcement learning to train an ML model to generate transforms for 3D representations (e.g., tooth transforms, or appliance component transforms).
  • 3D representations e.g., tooth transforms, or appliance component transforms.
  • FIG. 4 shows an example environment.
  • FIG. 5 summarizes a method of training an ML model using reinforcement learning, according to techniques of this disclosure.
  • FIG. 6 various neural networks trained for use in the reinforcement learning methods of this disclosure.
  • FIG. 7 shows transformer which may be configured to generate orthodontic setups transforms.
  • FIG. 8 shows an example method of using reinforcement learning with human feedback to train ML models of this disclosure.
  • Described herein are techniques for the automatic prediction of setups, which may provide the advantage of improving accuracy in comparison to existing techniques, enable new clinicians to be trained in the generation of effective setups, enable customized setups to be produced (e.g., which align with the specifications of clinicians), and provide the technical improvement of enhanced data precision in the formulation of these setups.
  • a setups prediction model of this disclosure may receive a variety of input data, which, as described herein, may include tooth meshes representing one or both arches of the patient.
  • the tooth data may be presented in the form of 3D representations, such as meshes or point clouds.
  • These data may be preprocessed, for example, by arranging the constituent mesh elements into lists and computing an optional mesh element feature vector for each mesh element.
  • Such vectors may impart valuable information of the shape and/or structure of the tooth to the setups prediction neural network. Additional inputs may enable the setups prediction neural network to better understand the distribution of the inputted data (e.g., tooth meshes), which provides the technical improvement of enabling customization to the specific medical/dental needs of the patient when the setups prediction model is deployed.
  • one or more oral care metrics may be computed.
  • Oral care metrics may be used for measuring one or more physical aspects of a setup (e.g., physical relationships within a tooth or between teeth).
  • an orthodontic metric may be computed for a ground truth setup which is then used in the training of a machine learning model (e.g., a setups prediction model).
  • the metric value may be received at the input of the setups prediction model, as a way of training the model to encode a distribution of such a metric over the several examples of the training dataset.
  • an “overbiteleff ’ metric may be computed for a setup which is received by the setups prediction model (e.g., at least one of mal and approved setup).
  • the network may then receive this metric value as an input, to assist in training the network to link that inputted metric value to the physical aspects of the received setup (e.g., to learn a distribution over the possible values of that metric across the examples of the training dataset).
  • the metric may be computed for the mal setup, and that metric value be providedas an input the network during training, alongside the malocclusion transforms and/or tooth meshes.
  • the metric may also (or alternatively) be computed for the approved setup, and that metric be provided as an input to the network during training, alongside the approved setup transforms and/or tooth meshes (e.g., for application during loss calculation time).
  • Such a loss calculation may quantify the difference between a prediction and a ground truth example (e.g., between a predicted setup and a ground truth setup).
  • the network may, through the course of loss calculation and subsequent backpropagation, learn to encode a distribution of that metric.
  • a technical improvement provided by the setups prediction techniques described herein is the customization of orthodontic treatment to the patient.
  • Oral care parameters may enable a clinician to customize specific desired aspects of the dimensions, proportions and other physical aspects of a predicted setup.
  • one or more oral care parameters may be defined and provided to the trained setups prediction model as part of the execution-phase input to specify one or more aspects of an intended setup upon an execution run.
  • a procedure parameter may be defined which corresponds to an oral care metric (e.g., such as the overbiteleft metric described above), which may be received at the input to a deployed setups prediction neural network and be taken as an instruction to the setups prediction neural network to generate a setup with the specified quantity of the metric (e.g., overbiteleft).
  • the setups prediction model may be especially suited to generating a setup with a prescribed value of a procedure parameter in the circumstance where that prescribed value falls within the distribution of the corresponding metric value that appeared in the training dataset.
  • Other procedure parameters may also be defined corresponding to other orthodontic metrics and be taken as instructions to the setups prediction model for the quantity of the relevant metric that is to be imparted to the predicted setup. This interplay between oral care metrics and oral care parameters may also apply to the training and deployment of other predictive models in oral care as well.
  • aspects of this disclosure are directed to forming training data that have a distribution which describes the kind of setup that the setups prediction neural network is configmed to produce. For example, to produce a final setup with an overbite of approximately 2.0 mm, one approach is to use ground truth training data with an overbite of approximately 2.0 mm. This approach may lead to a clean training signal and may produce useful results, and an alternative method may enable the network to learn to account for differences in overbite among the various ground truth training samples in the training dataset. An overbite metric may be computed for the malocclusion arches of a training sample (a patient case).
  • This overbite value may be received as an input to the setups prediction neural network at training time, along with the maloccluded tooth data, and serve as a signal to the neural network regarding the magnitude of overbite present in that mal arch.
  • the network thereby learns that different cases have different overbite magnitudes and can encode a distribution of possible overbite magnitudes, which can then be imparted to the predicted setup.
  • the trained neural network may receive the maloccluded tooth data as input and may also receive an input to indicate a magnitude of the overbite (e.g., or some other oral care metric) that is desired in the predicted setup (e.g., in the form of a procedure parameter which has been defined for the purpose).
  • This approach may enable the setups prediction neural network to account for differences in the distribution of the training dataset without excluding patient cases from the training dataset (e.g., as may be done in the case of filtering the training dataset), with the added benefit of enabling the deployed setups prediction neural network to customize the predicted setup, according to the specification of the clinician who uses the setups prediction model.
  • Other orthodontic metrics e.g., those disclosed herein
  • Corresponding procedure parameters e.g., those disclosed herein or those defined to correspond to specific metrics
  • Other techniques disclosed herein, besides setups prediction may also be trained with this use of oral care metrics and procedure parameters being received as inputs to a predictive model.
  • a setups prediction neural network of this disclosure may be trained, at least in part, by the calculation of one or more loss values (e.g., reconstruction loss or other loss values described herein). Such loss values may quantify the difference between a predicted setup and a corresponding ground truth setup. In some instances, these setups may be registered with each other (e.g., using iterative closest point (ICP) or singular value decomposition (SVD)) before the loss is computed, to reduce noise and improve the accuracy of the resulting trained setups prediction neural network. Such a registration may alternatively or additionally be performed between the maloccluded setup and the corresponding ground truth setup, with the advantage of reducing noise in the loss measurement and improving the accuracy of the trained network.
  • ICP iterative closest point
  • SSD singular value decomposition
  • the setups prediction neural network may compute a transform for each tooth, to move that tooth into a pose which is suitable for the end of orthodontic treatment (e.g., the final setup).
  • the pose of the tooth may include a change in position in 3D space and may also include a change in orientation (e.g., with respect to one or more coordinate axes - e.g., local coordinate axes with origin at the crown centroid).
  • the transform may effect the change in orientation by pivoting the tooth mesh relative to a pivot point or tooth origin. This pivot point may be chosen to lie within the crown centroid.
  • Alternatives include at the apex of the root tip, origin of malocclusion transform or at a point along an archform.
  • the setups prediction neural network may be trained conditionally on interproximal reduction (IPR) information.
  • IPR may be applied to the teeth, to enable greater packing of teeth a in final setup.
  • the setups model may be trained to account to IPR quantities (e.g., millimeters of offset in from either or both of the mesial and distal sides of a tooth) and/or IPR cut planes (which may be used in conjunction with mesh Boolean operations to remove material on either or both of the mesial and distal sides of a tooth).
  • IPR cut planes may be used to modify one or more tooth meshes for one or more patient cases which are used to train the setups prediction model.
  • IPR may be applied to a trial patient case, to modify the shapes of the teeth before the case is received as input to the setups prediction model. In some instances, IPR may be applied to one or more tooth meshes of a patient case before the computation of orthodontic metrics.
  • an anterior posterior (AP) shift may involve a sagittal shift of the mandible (lower arch), moving the mandible either forward or backwards.
  • the application of the AP Shift may improve the class relationship of the teeth.
  • Class may describe the patient’s malocclusion. Possible classes include: class 1, class 2 or class 3.
  • Elastics may aid in the shift of the mandible. Such elastics may attach to hardware on the teeth, such as buttons.
  • the setups prediction model of this disclosure may directly receive an AP shift transform as an input, which may improve the data precision of the resulting model.
  • an AP shift transform may first be applied to the patient case data before the patient case data are received as input to the setups prediction model of this disclosure.
  • the predictive models of the present disclosure may, in some implementations, may produce more accurate results by the incorporation of one or more of the following inputs: archform information V, interproximal reduction (IPR) information U, tooth dimension information P, tooth gap information Q, latent capsule representations of oral care meshes T, latent vector representations of oral care meshes A, procedure parameters K (which may describe a clinician’s intended treatment of the patient), doctor preferences L (which may describe the typical procedure parameters chosen by a doctor), flags regarding tooth status M (such as for fixed or pinned teeth), tooth position information N, tooth orientation information O, tooth name/dental notation R, oral care metrics S (comprising at least one of oral care metrics and restoration design metrics).
  • IPR interproximal reduction
  • Systems of this disclosure may, in some instances, be deployed at a clinical context (such as a dental or orthodontic office) for use by clinicians (e.g., doctors, dentists, orthodontists, nurses, hygienists, oral care technicians).
  • clinicians e.g., doctors, dentists, orthodontists, nurses, hygienists, oral care technicians.
  • Such systems which are deployed at a clinical context may enable clinicians to process oral care data (such as dental scans) in the clinic environment, or in some instances, in a "chairside" context (where the patient is present in the clinical environment).
  • a non-limiting list of examples of techniques may include: segmentation, mesh cleanup, coordinate system prediction, CTA trimline generation, restoration design generation, appliance component generation or placement or assembly, generation of other oral care meshes, the validation of oral care meshes, setups prediction, removal of hardware from tooth meshes, hardware placement on teeth, imputation of missing values, clustering on oral care data, oral care mesh classification, setups comparison, metrics calculation, or metrics visualization.
  • the execution of these techniques may, in some instances, enable patient data to be processed, analyzed and used in appliance generate by the clinician before the patient leaves the clinical environment (which may facilitate treatment planning because feedback may be received from the patient during the treatment planning process).
  • Systems of this disclosure may automate operations in digital orthodontics (e.g., setups prediction, hardware placement, setups comparison), in digital dentistry (e.g., restoration design generation) or in combinations thereof. Some techniques may apply to either or both of digital orthodontics and digital dentistry. A non-limiting list of examples is as follows: segmentation, mesh cleanup, coordinate system prediction, oral care mesh validation, imputation of oral care parameters, oral care mesh generation or modification (e.g., using autoencoders, transformers, continuous normalizing flows or denoising diffusion models), metrics visualization, appliance component placement or appliance component generation or the like. In some instances, systems of this disclosure may enable a clinician or technician to process oral care data (such as scanned dental arches).
  • the systems of this disclosure may enable orthodontic treatment planning, which may involve setups prediction as at least one operation.
  • Systems of this disclosure may also enable restoration design generation, where one or more restored tooth designs are generated and processed in the course of creating oral care appliances.
  • Systems of this disclosure may enable either or both of orthodontic or dental treatment planning, or may enable automation steps in the generation of either or both of orthodontic or dental appliances. Some appliances may enable both of dental and orthodontic treatment, while other appliances may enable one or the other.
  • a cohort patient case may include a set of tooth crown meshes, a set of tooth root meshes, or a data file containing attributes of the case (e.g., a JSON file).
  • a typical example of a cohort patient case may contain up to 32 crown meshes (e.g., which may each contain tens of thousands of vertices or tens of thousands of faces), up to 32 root meshes (e.g., which may each contain tens of thousands of vertices or tens of thousands of faces), multiple gingiva mesh (e.g., which may each contain tens of thousands of vertices or tens of thousands of faces) or one or more JSON files which may each contain tens of thousands of values (e.g., objects, arrays, strings, real values, Boolean values or Null values).
  • values e.g., objects, arrays, strings, real values, Boolean values or Null values
  • aspects of the present disclosure can provide a technical solution to the technical problem of predicting, using a machine learning model which has been trained using reinforcement learning, orthodontic setups for use in oral care appliance generation (e.g., intermediate stages or final setups for the generation of aligner trays).
  • aspects of the present disclosure may need to be executed in a time- constrained manner, such as when an oral care appliance must be generated for a patient immediately after intraoral scanning (e.g., while the patient waits in the clinician’s office).
  • aspects of the present disclosure are necessarily rooted in the underlying computer technology of setups transform prediction for oral care appliance generation and cannot be performed by a human, even with the aid of pen and paper.
  • implementations of the present disclosure must be capable of: 1) storing thousands or millions of mesh elements of the patient’s dentition in a manner that can be processed by a computer processor; 2) performing calculation on thousands or millions of mesh elements, e.g., to quantify aspects of the shape and or/structure of an individual tooth in the 3D representation of the patient’s dentition; and 3) predicting, based on a machine learning model which has been trained using reinforcement learning, orthodontic setups for use in oral care appliance generation), and do so during the course of a short office visit.
  • This disclosure pertains to digital oral care, which encompasses the fields of digital dentistry and digital orthodontics.
  • This disclosure generally describes methods of processing three-dimensional (3D) representations of oral care data.
  • 3D representation is a 3D geometry.
  • a 3D representation may include, be, or be part of one or more of a 3D polygon mesh, a 3D point cloud (e.g., such as derived from a 3D mesh), a 3D voxelized representation (e.g., a collection of voxels - for sparse processing), or 3D representations which are described by mathematical equations.
  • 3D representation may describe elements of the 3D geometry and/or 3D structure of an object.
  • a first arch S 1 includes a set of tooth meshes arranged (e.g., using transforms) in their positions in the mouth, where the teeth are in the mal positions and orientations.
  • a second arch S2 includes the same set of tooth meshes from SI arranged (e.g., using transforms) in their positions in the mouth, where the teeth are in the ground truth setup positions and orientations.
  • a third arch S3 includes the same meshes as SI and S2, which are arranged (e.g., using transforms) in their positions in the mouth, where the teeth are in the predicted final setup poses (e.g., as predicted by one or more of the techniques of this disclosure).
  • S4 is a counterpart to S3, where the teeth are in the poses corresponding to one of the several intermediate stages of orthodontic treatment with clear tray aligners.
  • GDL geometric deep learning
  • RL reinforcement learning
  • VAE variational autoencoder
  • MLP multilayer perceptron
  • PT pose transfer
  • FDG force directed graphs
  • MLP Setups, VAE Setups and Capsule Setups each fall within the scope of Autoencoder Setups. Some implementations of MLP Setups may fall within the Scope of Transformer Setups.
  • FIG. 2 shows a non-limiting selection of models which may be trained for setups prediction.
  • Representation Setups refers to any of MLP Setups, VAE Setups, Capsule Setups and any other setups prediction machine learning model which uses an autoencoder to create the representation for at least one tooth.
  • Each of the setups prediction techniques of this disclosure is applicable to the fabrication of clear tray aligners and/or indirect bonding trays.
  • the setups predictions techniques may also be applicable to other products that involve final teeth poses, also.
  • a pose may comprise a position (or location) and a rotation (or orientation).
  • a 3D mesh is a data structure which may describe the geometry or shape of an object related to oral care, including but not limited to a tooth, a hardware element, or a patient’s gum tissue.
  • a 3D mesh may include one or more mesh elements such as one or more of vertices, edges, faces and combinations thereof.
  • mesh element may include voxels, such as in the context of sparse mesh processing operations.
  • Various spatial and structural features may be computed for these mesh elements and be provided to the predictive models of this disclosure, with the predictive models of this disclosure providing the technical advantage of improving data precision in the form of the models of this disclosure outputting more accurate predictions.
  • a patient’s dentition may include one or more 3D representations of the patient’s teeth (e.g., and/or associated transforms), gums and/or other oral anatomy.
  • An orthodontic metric may, in some implementations, quantify the relative positions and/or orientations of at least one 3D representation of a tooth relative to at least one other 3D representation of a tooth.
  • a restoration design metric may, in some implementations, quantify the at least one aspect of the structure and/or shape of a 3D representation of a tooth.
  • An orthodontic landmark (OL) may, in some implementations, locate one or more points or other structural regions of interest on a 3D representation of a tooth.
  • An OL may, in some implementations, be used in the generation of an orthodontic or dental appliance, such as a clear tray aligner or a dental restoration appliance.
  • a mesh element may, in some implementations, comprise at least one constituent element of a 3D representation of oral care data.
  • mesh elements may include at least: vertices, edges, faces and voxels.
  • a mesh element feature may, in some implementations, quantify some aspect of a 3D representation in proximity to or in relation with one or more mesh elements, as described elsewhere in this disclosure.
  • Orthodontic procedure parameters may, in some implementations, specify at least one value which defines at least one aspect of planned orthodontic treatment for the patient (e.g., specifying desired target attributes of a final setup in final setups prediction).
  • Orthodontic Doctor preferences may, in some implementations, specify at least one typical value for an OPP, which may, in some instances, be derived from past cases which have been treated by one or more oral care practitioners.
  • Restoration Design Parameters may, in some implementations, specify at least one value which defines at least one aspect of planned dental restoration treatment for the patient (e.g., specifying desired target attributes of a tooth which is to undergo treatment with a dental restoration appliance).
  • Doctor Restoration Design Preferences may, in some implementations, specify at least one typical value for an RDP, which may, in some instances, be derived from past cases which have been treated by one or more oral care practitioners.
  • 3D oral care representations may include, but are not limited to: 1) a set of mesh element labels which may be applied to the 3D mesh elements of teeth/gums/hardware/appliance meshes (or point clouds) in the course of mesh segmentation or mesh cleanup; 2) 3D representation(s) for one or more teeth/gums/hardware/appliances for which shapes have been modified (e.g., trimmed, distorted, or filled- in) in the course of mesh segmentation or mesh cleanup; 3) one or more coordinate systems (e.g., describing one, two, three or more coordinate axes) for a single tooth or a group of teeth (such as a full arch - as with the LDE coordinate system); 4) 3D representation(s) for one or more teeth for which shapes have been modified or otherwise made suitable for use in
  • the Setups Comparison tool may be used to compare the output of the GDL Setups model against ground truth data, compare the output of the RL Setups model against ground truth data, compare the output of the VAE Setups model against ground truth data and compare the output of the MLP Setups model against ground truth data.
  • the Metrics Visualization tool can enable a global view of the final setups and intermediate stages produced by one or more of the setups prediction models, with the advantage of enabling the selection of the best setups prediction model.
  • the Metrics Visualization tool furthermore, enables the computation of metrics which have a global scope over a set of intermediate stages. These global metrics may, in some implementations, be consumed as inputs to the neural networks for predicting setups (e.g., GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, among others). The global metrics may also be consumed by FDG Setups.
  • GDL Setups e.g., GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, among others.
  • the global metrics may also be consumed by FDG Setups.
  • the local metrics from this disclosure may, in some implementations, be consumed by the neural networks herein for predicting setups, with the advantage of improving predictive results.
  • the metrics described in this disclosure may, in some implementations, be visualized using the Metric Visualization tool.
  • the VAE and MAE models for mesh element labelling and mesh in-filling can be advantageously combined with the setups prediction neural networks, for the purpose of mesh cleanup ahead of or during the prediction process.
  • the VAE for mesh element labelling may be used to flag mesh elements for further processing, such as metrics calculation, removal or modification.
  • flagged mesh elements may be used as inputs to a setups prediction neural network, to inform that neural network about important mesh features, attributes or geometries, with the advantage of improving the performance of the resulting setups prediction model.
  • mesh in-filling may cause the geometry of a tooth to become more nearly complete, enabling the better functioning of a setups prediction model (i.e., improved correctness of prediction on account of better-formed geometry).
  • a neural network to classify a setup i.e., the Setups Classifier
  • the setups classifier tells that setups prediction neural network when the predicted setup is acceptable for use and can be provided to a method for aligner tray generation.
  • a Setups Classifier (e.g., GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups and FDG Setups, among others) may aid in the generation of final setups and also in the generation of intermediate stages.
  • a Setups Classifier neural network may be combined with the Metrics Visualization tool.
  • a Setups Classification neural network may be combined with the Setups Comparison tool (e.g., the Setup Comparison tool may output an indication of how a setup produced in part by the Setups Classifier compares to a setup produced by another setups prediction method).
  • the VAE for mesh element labelling may identify one or more mesh elements for use in a metrics calculation. The resulting metrics outputs may be visualized by the Metrics Visualization tool.
  • the Setups Classifier neural network may aid in the setups prediction technique described in U.S. Patent Application No. US20210259808A1 (which is incorporated herein by reference in its entirety) or the setups prediction technique described in PCT Application with Publication No. WO2021245480A1 (which is incorporated herein by reference in its entirety) or in PCT Application No. PCT/IB2022/057373 (which is incorporated herein by reference in its entirety).
  • the Setups Classifier would help one or more of those techniques to know when the predicted final setup is most nearly correct.
  • the Setups Classifier neural network may output an indication of how far away from final setup a given setup is (i.e., a progress indicator).
  • the latent space embedding vector(s) from the reconstruction VAE can be concatenated with the inputs to the setups prediction neural network described in WO2021245480A1.
  • the latent space vectors can also be incorporated as inputs to the other setups prediction models: GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups and Diffusion Setups, among others.
  • the advantage is to impart the reconstruction characteristics (e.g., latent vector dimensions of a tooth mesh) to that neural network, hence improving the generated setups prediction.
  • the various setups prediction neural networks of this disclosure may work together to produce the setups required for orthodontic treatment.
  • the GDL Setups model may produce a final setup, and the RL Setups model may use that final setup as input to produce a series of intermediate stages setups.
  • the VAE Setups model (or the MLP Setups model) may create a final setup which may be used by an RL Setups model to produce a series of intermediate stages setups.
  • a setup prediction may be produced by one setups prediction neural network, and then taken as input to another setups prediction neural network for further improvements and adjustments to be made, in some implementations, such improvements may be performed in iterative fashion.
  • a setups validation model such as the model disclosed in US Provisional Application No. US63/366495, may be involved in this iterative setups prediction loop.
  • a setup may be generated (e.g., using a model trained for setups prediction, such as GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups and FDG Setups, among others), then the setup undergoes validation. If the setup passes validation, the setup may be outputted for use. If the setup fails validation, the setup may be sent back to one or more of the setups prediction models for corrections, improvements and/or adjustments.
  • a model trained for setups prediction such as GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups and FDG Setups, among others.
  • the setups validation model may output an indication of what is wrong with the setup, enabling the setups generation model to make an improved version upon the next iteration. The process iterates until done.
  • two or more of the following techniques of the present disclosure may be combined in the course of orthodontic and/or dental treatment: GDL Setups, Setups Classification, Reinforcement Learning (RL) Setups, Setups Comparison, Autoencoder Setups (VAE Setups or Capsule Setups), VAE Mesh Element Labeling, Masked Autoencoder (MAE) Mesh Infilling, Multi-Layer Perceptron (MLP) Setups, Metrics Visualization, Imputation of Missing Oral Care Parameters Values, Tooth Classification Using Latent Vector, FDG Setups, Pose Transfer Setups, Restoration Design Metrics Calculation, Neural Network Techniques for Dental Restoration and/or Orthodontics (e.g., 3D Oral Care Represent
  • Oral care parameters may include one or more values that specify orthodontic procedure parameters, or restoration design parameters (RDP), as described herein.
  • Oral care parameters may define one or more intended aspects of a 3D oral care representation, and may be provided to an ML model to promote that ML model to generate output which may be used in the generation of oral care appliances that are suitable for the treatment of a patient.
  • Other types of values include doctor preferences and restoration design preferences, as described herein.
  • Doctor preferences and restoration design preferences may define the typical treatment choices or practices of a particular clinician. Restoration design preferences are subjective to a particular clinician, and so differ from restoration design parameters.
  • doctor preferences or restoration design preferences may be computed by unsupervised means, such as clustering, which may determine the typical values that a clinician uses in patient treatment. Those typical values may be stored in a datastore, and recalled to be provided to an automated ML model as default values (e.g., default values which may be modified before execution of the model).
  • one clinician may prefer one value for a restoration design parameter (RDP), while another clinician may prefer a different value for that RDP, when faced with a similar diagnosis or treatment protocol.
  • RDP restoration design parameter
  • Procedure parameters and/or doctor preferences may, in some implementations, be provided to a setups prediction model for orthodontic treatment, for the purpose of improving the customization of the resulting orthodontic appliance.
  • Restoration design parameters and doctor restoration preferences may in some implementations be used to design tooth geometry for use in the creation of a dental restoration appliance, for the purpose of improving the customization of that appliance.
  • some implementations of ML prediction models of this disclosure, in orthodontic treatment may also take as input a setup (e.g., an arrangement of teeth).
  • an ML prediction model of this disclosure may take as input a final setup (i.e., final arrangement of teeth), such as in the case of a prediction model trained to generate intermediate stages.
  • these preferences are referred to as doctor restoration preferences, but it is intended to be used in a non-limiting sense. Specifically, it should be appreciated that these preferences may be specified by any treating or otherwise appropriate medical professional and are not intended to be limited to doctor preferences per se (i.e., preferences from someone in possession of an M.D. or equivalent degree).
  • An oral care professional or clinician such as a dentist or orthodontist, may specify information about patient treatment in the form of a patient-specific set of procedure parameters.
  • an oral care professional may specify a set of general preferences (aka doctor preferences) for use over a broad range of cases, to use as default values in the set of procedure parameters specification process.
  • Oral care parameters may in some implementations be incorporated into the techniques described in this disclosure, such as one or more of GDL Setups, VAE Setups, RL Setups, Setups Comparison, Setups Classification, VAE Mesh Element Labelling, MAE Mesh In-Filling, Validation Using Autoencoders, Imputation of Missing Procedure Parameters Values, Metrics Visualization, or FDG Setups.
  • One or more of these models may take as input one or more procedure parameters vector K and/or one or more doctor preference vectors L.
  • one or more of these models may introduce to one or more of a neural network’s hidden layers one or more procedure parameters vector K and/or one or more doctor preferences vectors L.
  • one or more of these models may introduce either or both of K and L to a mathematical calculation, such as a force calculation, for the purpose of improving that calculation and the ultimate customization of the resulting appliance to the patient.
  • Some implementations of a neural network for predicting a setup may incorporate information from an oral care professional (aka doctor). This information may influence the arrangement of teeth in the final setup, bringing the positions and orientations of the teeth into conformance with a specification set by the doctor, within tolerances.
  • oral care parameters may be provided directly into the generator network as a separate input alongside the mesh data.
  • oral care parameters may be incorporated into the feature vector which is computed for each mesh element before the mesh elements are input to the generator for processing.
  • Some implementations of a VAE Setup model may incorporate oral care parameters into the setups predictions.
  • the procedure parameters K and/or the doctor preference information L may be concatenated with the latent space vector C.
  • a doctor’s preferences e.g., in an orthodontic context
  • doctor’s restoration preferences may be indicated in a treatment form, or they could be based upon characteristics in treatment plans such as final setup characteristics (e.g., amount of bite correction or midline correction in planned final setups), intermediate staging characteristics (e.g., treatment duration, tooth movement protocols, or overcorrection strategies), or outcomes (e.g., number of revisions/refinements).
  • Orthodontic procedure parameters may specify one or more of the following (with possible values shown in ⁇ ⁇ ).
  • Non-limiting categorical values for some example OPP are described below.
  • a real value may be specified for one or more of these OPP.
  • the Overbite OPP may specify a quantity of overbite (e.g., in millimeters) which is desired in a setup, and may be received as input of a setups prediction model to provide that setups prediction model information about the amount of overbite which is desired in the setup.
  • Some implementations may specify a numerical value for the Oveijet OPP, or other OPP.
  • one or more OPP may be defined which correspond to one or more orthodontic metrics (OM).
  • OM orthodontic metrics
  • a numerical value may be specified for such an OPP, for the purpose of controlling the output of a setups prediction model.
  • Tooth Movement Restrictions for each tooth, indicate if tooth is ⁇ DoNotMove, Missing, ToBeExtracted, Primary /Erupting, Clear ⁇
  • Oveijet ⁇ ShowResultingOverjetAfterAlignment, MaintainfnitialOveijet, ImproveResultingOveijet ⁇ Anterior/Posterior (AP) Relationship
  • doctor can specify an archform - selected from a set of options or custom-designed]
  • Other orthodontic procedure parameters may be defined, such as those which may be used to place standardized brackets at prescribed occlusal heights on the teeth.
  • one or more orthodontic procedure parameters may be defined to specify at least one of the 2 nd and 3 rd order rotation angles to be applied to a tooth (i.e., angulation and torque, respectively), which may enable a target setup arrangement where crown landmarks lie within a threshold distance of a common occlusal plane, for example.
  • one or more orthodontic procedure parameters may be defined to specify the position in global coordinates where at least one landmark (e.g., a centroid) of a tooth crown (or root) is to be placed in a setup arrangement of teeth.
  • an oral care parameter may be defined which corresponds to an oral care metric.
  • an orthodontic procedure parameter may be defined which corresponds to an orthodontic metric (e.g., to specify at the input of a setups prediction model an amount of a certain metric which is desired to appear in a predicted setup).
  • Doctor preferences may differ from orthodontic procedure parameters in that doctor preferences pertain to an oral care provider and may comprise of the means, modes, medians, minimums, or maximums (or some other statistic) of past settings associated with an oral care provider’s treatment decisions on past orthodontic cases.
  • Procedure parameters may pertain to a specific patient, and describe the needs of a particular patient’s treatment.
  • Doctor preferences may pertain to a doctor and the doctor’s past treatment practices, whereas procedure parameters may pertain to the treatment of a particular patient.
  • Doctor preferences (or “treatment preferences”) may specify one or more of the following (with some non-limiting possible values shown in ⁇ ⁇ ). Other possible values are found elsewhere in this disclosure.
  • Doctor preferences may specify one or more of the following (with other possible values found elsewhere in this disclosure).
  • Protocol A ⁇ protocol A, protocol B, protocol C ⁇
  • archform information V may be provided as an input to any of the GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups and Diffusion Setups prediction neural networks. In some implementations, archform information V may be provided directly to one or more internal neural network layers in one or more of those setups applications.
  • the additional procedure parameters may include text descriptions of the patient’s medical condition and of the intended treatment.
  • Such text descriptions may be analyzed via natural language processing operations, including tokenization, stop word removal, stemming, n-gram formation, text data vectorization, bag of words analysis, term frequency inverse document frequency (TF-IDF) analysis, sentiment analysis, naive Bayes classification, and/or logistic regression classification.
  • TF-IDF term frequency inverse document frequency
  • the outputs of such analysis techniques may be used as input to one or more of the neural networks of this disclosure with the advantage of customizing and improving the predicted outputs (e.g., the predicted setups or predicted mesh geometries).
  • a dataset used for training one or more of the neural network models of this disclosure may be filtered conditionally on one or more of the orthodontic procedure parameters described in this section.
  • patient cases which exhibit outlier values for one or more of these procedure parameters may be omitted from a dataset (alternatively used to form a dataset) for training one or more of the neural networks of this disclosure.
  • One or more procedure parameters and/or doctor preferences may be provided to a neural network during training. In this manner the neural network may be conditioned on the one or more procedure parameters and/or doctor preferences.
  • Examples of such neural networks include a conditional generative adversarial network (cGAN) and/or a conditional variational autoencoder (cVAE), either of which may be used for the various neural network-based applications of this disclosure.
  • tooth shape-based inputs may be provided to a neural network for setups predictions.
  • non-shape-based inputs can be used, such as a tooth name or designation, as it pertains to dental notation.
  • a vector R of flags may be provided to the neural network, where a ‘ 1 ’ value indicates that the tooth is present and a ‘0’ value indicates that the tooth is absent from the patient case (though other values are possible).
  • the vector R may comprise a 1- hot vector, where each element in the vector corresponds to a tooth type, name or designation.
  • Identifying information about a tooth can be provided to the predictive neural networks of this disclosure, with the advantage of enabling the neural network to become trained to handle different teeth in tooth-specific ways.
  • the setups prediction model may learn to make setups transformations predictions for a specific tooth designation (e.g., upper right central incisor, or lower left cuspid, etc.).
  • the mesh cleanup autoencoders either for labelling mesh element or for in-filling missing mesh data
  • the autoencoder may be trained to provide specialized treatment to a tooth according to that tooth’s designation, in this manner.
  • Tooth designation/name may be defined, for example, according to the Universal Numbering System, Palmer System, or the FDI World Dental Federation notation (ISO 3950).
  • a vector R may be defined as an optional input to the setups prediction neural networks of this disclosure, where there is a 0 in the vector element corresponding to each of the wisdom teeth, and a 1 in the elements corresponding to the following teeth: UR7, UR6, UR5, UR4, UR3, UR2, UR1, ULI, UL2, UL3, UL4, UL5, UL6, UL7, LL7, LL6, LL5, LL4, LL3, LL2, LL1, LR1, LR2, LR3, LR4, LR5, LR6, LR7 [0059]
  • the position of the tooth tip may be provided to a neural network for setups predictions.
  • the neural networks may take as input one or more indications of interproximal reduction (IPR) U, which may indicate the amount of enamel that is to be removed from a tooth during the course orthodontic treatment (either mesially or distally).
  • IPR information e.g., quantity of IPR that is to be performed on one or more teeth, as measured in millimeters, or one or more binary flags to indicate whether or not IPR is to be performed on each tooth identified by flagging
  • the vector(s) and/or capsule(s) resulting from such a concatenation may be provided to one or more of the neural networks of the present disclosure, with the technical improvement or added advantage of enabling that predictive neural network to account for IPR.
  • IPR is especially relevant to setups prediction methods, which may determine the positions and poses of teeth at the end of treatment or during one or more stages during treatment. It is important to account for the amount of enamel that is to be removed ahead of predicted tooth movements.
  • tooth dimensions P such as length, width, height, or circumference may be measured inside a plane, such as the plane that intersects the centroid of the tooth, or the plane that intersects a center point that is located midway between the centroid and either the incisal-most extent or the gingival-most extent of the tooth.
  • the tooth dimension of height may be measured as the distance from gums to incisal edge.
  • the tooth dimension of width may be measured as the distance from the mesial extent to the distal extent of the tooth.
  • the circularity or roundness of the tooth cross-section may be measured and included in the vector P. Circularity or roundness may be defined as the ratio of the radii of inscribed and circumscribed circles.
  • the distance Q between adjacent teeth can be implemented in different ways (and computed using different distance definitions, such as Euclidean or geodesic).
  • a distance QI may be measured as an averaged distance between the mesh elements of two adjacent teeth.
  • a distance Q2 may be measured as the distance between the centers or centroids of two adjacent teeth.
  • a distance Q3 may be measured between the mesh elements of closest approach between two adjacent teeth.
  • a distance Q4 may be measured between the cusp tips of two adjacent teeth. Teeth may, in some implementations, be considered adjacent within an arch. Teeth may, in some implementations, also be considered adjacent between opposing arches.
  • any of QI, Q2, Q3 and Q4 may be divided by a term for the purpose of normalizing the resulting value of Q.
  • the normalizing term may involve one or more of: the volume of a tooth, the count of mesh elements in a tooth, the surface area of a tooth, the cross-sectional area of a tooth (e.g., as projected into the XY plane), or some other term related to tooth size.
  • Other information about the patient’s dentition or treatment needs may be concatenated with the other input vectors to one or more of MLP, GAN, generator, encoder structure, decoder structure, transformer, VAE, conditional VAE, regularized VAE, 3D U-Net, capsule autoencoder, diffusion model, and/or any of the neural networks models listed elsewhere in this disclosure.
  • the vector M may contain flags which apply to one or more teeth.
  • M contains at least one flag for each tooth to indicate whether the tooth is pinned.
  • M contains at least one flag for each tooth to indicate whether the tooth is fixed.
  • M contains at least one flag for each tooth to indicate whether the tooth is pontic.
  • Other and additional flags are possible for teeth, as are combinations of fixed, pinned and pontic flags.
  • a flag that is set to a value that indicates that a tooth should be fixed is a signal to the network that the tooth should not move over the course of treatment.
  • the neural network loss function may be designed to be penalized for any movement in the indicated teeth (and in some particular cases, may be heavily penalized).
  • a flag to indicate that a tooth is pontic informs the network that the tooth gap is to be maintained, although that gap is allowed to move.
  • M may contain a flag indicating that a tooth is missing.
  • the presence of one or more fixed teeth in an arch may aid in setups prediction, because the one or more fixed teeth may provide an anchor for the poses of the other teeth in the arch (i.e., may provide a fixed reference for the pose transformations of one or more of the other teeth in the arch).
  • one or more teeth may be intentionally fixed, so as to provide an anchor against which the other teeth may be positioned.
  • a 3D representation (such as a mesh) which corresponds to the gums may be introduced, to provide a reference point against which teeth can be moved.
  • one or more of the optional input vectors K, L, M, N, O, P, Q, R, S, U and V described elsewhere in this disclosure may also be provided to the input or into an intermediate layer of one or more of the predictive models of this disclosure.
  • these optional vectors may be provided to the MLP Setups, GDL Setups, RL Setups, VAE Setups, Capsule Setups and/or Diffusion Setups, with the advantage of enabling the respective model to generate setups which better meet the orthodontic treatment needs of the patient.
  • such inputs may be provided, for example, by being concatenated with one or more latent vectors A which are also provided to one or more of the predictive models of this disclosure.
  • such inputs may be introduced, for example, by being concatenated with one or more latent capsules T which are also provided to one or more of the predictive models of this disclosure.
  • K, L, M, N, O, P, Q, R, S, U and V may be introduced to the neural network (e.g., MLP or Transformer) directly in a hidden layer of the network.
  • the neural network e.g., MLP or Transformer
  • K, L, M, N, O, P, Q, R, S, U and V may be introduced directly into the internal processing of an encoder structure.
  • a setups prediction model (such as GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, PT Setups, Similarity Setups and Diffusion Setups) may take as input one or more latent vectors A which correspond to one or more input oral care meshes (e.g., such as tooth meshes).
  • a setups prediction model (such as GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups and Diffusion Setups) may take as input one or more latent capsules T which correspond to one or more input oral care meshes (e.g., such as tooth meshes).
  • a setups prediction method may take as input both of A and T.
  • setups prediction neural networks e.g., GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, or FDG Setup, or other setups prediction network architectures
  • GDL Setups e.g., GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, or FDG Setup, or other setups prediction network architectures
  • Some implementations of the setups prediction neural networks may take additional inputs to aid in setups prediction. Some of these inputs may reflect the geometrical attributes of one or more teeth or of a whole arch.
  • an archform or arch curve may be provided to a setups prediction neural network, with the technical improvement of aiding that setups prediction neural network in finding a suitable set of final setups poses for the teeth in a patient case (with the technical improvements being directed to both resource footprint reduction by way of more efficient location capabilities and/or data precision in the form of locating a more pertinent final setup).
  • the archform or arch curve may be encoded as a spline, a B-spline, NonUniform Rational B-Splines (NURBS), polynomial spline, non-polynomial spline, parabolic curve, hyperbolic curve or other parameterized curve.
  • Such a curve may be computed as an average of multiple exemplars, such as exemplary final setups.
  • Another non-limiting example of an archform is a Beta curve.
  • the arch information may be provided to the encoder E2, as an additional input alongside E and D.
  • the arch information may be provided to the generator as an additional input to the mesh element lists and associated mesh element feature vectors.
  • an archform may be described by one or more 3D representations, such as a 3D mesh, a set of 3D control points and/or as a 3D polyline.
  • a Frenet frame may be overlaid onto an archform.
  • the Frenet frame may locally describe the coordinate system corresponding to each point along the archform.
  • a coordinate system may, in some implementations, be right-handed (or alternatively, in other implementations, left-handed).
  • Such a coordinate system may, in some implementations, be determined, at least in part, by at least one of the tangent to the archform at the point and the archform’s curvature.
  • a point may be described using an LDE coordinate frame relative to an archform, where L, D and E correspond to: 1) Length along the curve of the archform, 2) Distance away from the archform, and 3) Distance in the direction perpendicular to the L and D axes (which may be termed Eminence), respectively.
  • Other geometrical inputs may also aid in the training of a setups prediction neural network.
  • Various loss calculation techniques are generally applicable to the techniques of this disclosure (e.g., GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, Setups Classification, Tooth Classification, VAE Mesh Element Labelling, MAE Mesh In-Filling and the imputation of procedure parameters).
  • Losses include LI loss, L2 loss, mean squared error (MSE) loss, cross entropy loss, among others.
  • Losses may be computed and used in the training of neural networks, such as multi-layer perceptron’s (MLP), U-Net structures, generators and discriminators (e.g., for GANs), autoencoders, variational autoencoders, regularized autoencoders, masked autoencoders, transformer structures, or the like. Some implementations may use either triplet loss or contrastive loss, for example, in the learning of sequences.
  • MLP multi-layer perceptron’s
  • U-Net structures such as generators and discriminators (e.g., for GANs), autoencoders, variational autoencoders, regularized autoencoders, masked autoencoders, transformer structures, or the like.
  • Some implementations may use either triplet loss or contrastive loss, for example, in the learning of sequences.
  • Losses may also be used to train encoder structures and decoder structures.
  • a KL- Divergence loss may be used, at least in part, to train one or more of the neural networks of the present disclosure, such as a mesh reconstruction autoencoder or the generator of GDL Setups, which the advantage of imparting Gaussian behavior to the optimization space.
  • This Gaussian behavior may enable a reconstruction autoencoder to produce a better reconstruction (e.g., when a latent vector representation is modified and that modified latent vector is reconstructed using a decoder, the resulting reconstruction is more likely to be a valid instance of the inputted representation).
  • There are other techniques for computing losses which may be described elsewhere in this disclosure. Such losses may be based on quantifying the difference between two or more 3D representations.
  • MSE loss calculation may involve the calculation of an average squared distance between two sets, vectors or datasets. MSE may be generally minimized. MSE may be applicable to a regression problem, where the prediction generated by the neural network or other machine learning model may be a real number.
  • a neural network may be equipped with one or more linear activation units on the output to generate an MSE prediction.
  • Mean absolute error (MAE) loss and mean absolute percentage error (MAPE) loss can also be used in accordance with the techniques of this disclosure.
  • Cross entropy may, in some implementations, be used to quantify the difference between two or more distributions. Cross entropy loss may, in some implementations, be used to train the neural networks of the present disclosure.
  • Cross entropy loss may, in some implementations, involve comparing a predicted probability to a ground truth probability. Other names of cross entropy loss include “logarithmic loss,” “logistic loss,” and “log loss”. A small cross entropy loss may indicate a better (e.g., more accurate) model. Cross entropy loss may be logarithmic. Cross entropy loss may, in some implementations, be applied to binary classification problems. In some implementations, a neural network may be equipped with a sigmoid activation unit at the output to generate a probability prediction. In the case of multi-class classifications, cross entropy may also be used.
  • a neural network trained to make multi-class predictions may, in some implementations, be equipped with one or more softmax activation functions at the output (e.g., where there is one output node for class that is to be predicted).
  • Other loss calculation techniques which may be applied in the training of the neural networks of this disclosure include one or more of: Huber loss, Hinge loss, Categorical hinge loss, cosine similarity, Poisson loss, Logcosh loss, or mean squared logarithmic error loss (MSLE). Other loss calculation methods are described herein and may be applied to the training of any of the neural networks described in the present disclosure.
  • One or more of the neural networks of the present disclosure may, in some implementations, be trained, at least in part by a loss which is based on at least one of: a Point-wise Mesh Euclidean Distance (PMD) and an Earth Mover’s Distance (EMD).
  • PMD Point-wise Mesh Euclidean Distance
  • EMD Earth Mover’s Distance
  • Some implementations may incorporate a Hausdorff Distance (HD) calculation into the loss calculation.
  • HD Hausdorff Distance
  • Computing the Hausdorff distance between two or more 3D representations may provide one or more technical improvements, in that the HD not only accounts for the distances between two meshes, but also accounts for the way that those meshes are oriented, and the relationship between the mesh shapes in those orientations (or positions or poses).
  • Hausdorff distance may improve the comparison of two or more tooth meshes, such as two or more instances of a tooth mesh which are in different poses (e.g., such as the comparison of predicted setup to ground truth setup which may be performed in the course of computing a loss value for training a setups prediction neural network).
  • Reconstruction loss may compare a predicted output to a ground truth (or reference) output.
  • all_points_target is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to ground tmth data (e.g., a ground truth tooth restoration design, or a ground tmth example of some other 3D oral care representation).
  • all_points_predicted is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to generated or predicted data (e.g., a generated tooth restoration design, or a generated example of some other kind of 3D oral care representation).
  • Other implementations of reconstruction loss may additionally (or alternatively) involve L2 loss, mean absolute error (MAE) loss or Huber loss terms.
  • NLP natural language processing
  • One example application of NLP is the generation of new text based upon prior words or text.
  • Transformers have in turn provided significant improvements over GRU, LSTM and other such RNN-based NLP techniques due to an important attribute of the transformer model, which has the property of multi-headed attention.
  • the NLP concept of multi-headed attention may describe the relationship between each word in a sentence (or paragraph or document or corpus of documents) and each other word in that sentence (or paragraph or document or corpus of documents). These relationships may be generated by a multiheaded attention module, and may be encoded in vector form.
  • This vector may describe how each word in a sentence (or paragraph or document or corpus of documents) should attend to each other word in that sentence (or paragraph or document or corpus of documents).
  • RNN, LSTM and GRU models process a sequence, such a sentence, one word at a time from the start to the end of the sequence. Furthermore, the model may only account for a given subset (called a window) of the sentence when making a prediction.
  • transformer-based models may, in some instances, account for the entirety of the preceding text by processing the sequence in its entirety in a single step.
  • Transformer, RNN, LSTM, and GRU models can all be adapted for use in predictive models in digital dentistry and digital orthodontics, particularly for the setup prediction task.
  • an exemplary transformer model for use with 3D meshes and 3D transforms in setups prediction may be adapted from the Bidirectional Encoder Representation from Transformers (BERT) and/or Generative Pre-Training (GPT) models.
  • a GPT (or BERT) model may first be trained on other data, such as text or documents data, and then be used in transfer learning. Such a transfer learning process may receive a previously trained GPT or BERT model, and then do further training using data comprising 3D oral care representations.
  • Such transfer learning may be performed to train oral care models such as: segmentation, mesh cleanup, coordinate system prediction, setups prediction, validation of 3D oral care representations, transform prediction for placement of oral care meshes (e.g., teeth, hardware, appliance components, fixture model components), tooth restoration design generation (or generation of other 3D oral care representations - such as appliance components, fixture models or archforms), classification of 3D oral care representations, imputation of missing oral care parameters, clustering of clinicians or clustering of clinician preferences, or the like.
  • oral care models such as: segmentation, mesh cleanup, coordinate system prediction, setups prediction, validation of 3D oral care representations, transform prediction for placement of oral care meshes (e.g., teeth, hardware, appliance components, fixture model components), tooth restoration design generation (or generation of other 3D oral care representations - such as appliance components, fixture models or archforms), classification of 3D oral care representations, imputation of missing oral care parameters, clustering of clinicians or clustering of clinician preferences, or the like.
  • Oral care data may comprise one or more of (or combinations of): 3D representations of tooth (e.g., meshes, point clouds or voxels), sections of tooth meshes (such as subsets of mesh elements), tooth transforms (such as in matrix, vector and/or quaternion form, or combinations thereof), transforms for appliance components, transforms for fixture model components, and mesh coordinate system definitions (such as represented by transforms, for example, transformation matrices) and/or other 3D oral care representations described herein.
  • 3D representations of tooth e.g., meshes, point clouds or voxels
  • sections of tooth meshes such as subsets of mesh elements
  • tooth transforms such as in matrix, vector and/or quaternion form, or combinations thereof
  • transforms for appliance components transforms for fixture model components
  • mesh coordinate system definitions such as represented by transforms, for example, transformation matrices
  • Transformers may be trained for generating transforms to position teeth into setups poses (or to place appliance components for use in appliance generation or to place fixture model components for use in fixture model generation). Some implementations may operate in an offline prediction context, and some implementations operation in an online reinforcement learning (RL) context.
  • RL online reinforcement learning
  • a transformer may be initially trained in an offline context and then undergo further fine-tuning training in the online context.
  • the transformer may be trained from a dataset of cohort patient case data.
  • the transformer may be trained from either a physics model, or a CAD model, for example.
  • the transformer may learn from static data, such as transformations (e.g., trajectory transformer).
  • the transform may provide a mapping from malocclusion to setup (e.g., receiving transformation matrices as input and generating transformation matrices as ouput).
  • Some implementations of transformers may be trained to process 3D representations, such as 3D meshes, 3D point clouds or voxels (e.g., using a decision transformer) takes as input geometry (e.g., mesh, point cloud, voxels etc.), outputs transformations.
  • the decision transformer may be coupled with a representation generation module that encodes representation of the patient’s dentition (e.g., teeth), such as a VAE, a U-Net, an encoder, a transformer encoder, a pyramid encoder-decoder or a simple dense or fully connected network, or a combination thereof.
  • a representation generation module e.g., VAE, the U-Net, the encoder, the pyramid encoder-decoder or the dense network for generating the tooth representation
  • VAE the U-Net
  • the representation generation module may be trained on all teeth in both arches, only the teeth within the same arch (either upper or lower), only anterior teeth, only posterior teeth, or some other subset of teeth.
  • such a model may be trained on each individual tooth (e.g., an upper right cuspid), so that the model is trained or otherwise configured togenerate highly accurate representations for an individual tooth.
  • an encoder structure may encode such a representation.
  • a decision transformer may learn in an online context, in an offline context or both.
  • An online decision transformer may be trained (e.g., using RL techniques) to output action, state, and/or reward.
  • transformations may be discretized, to allow for piecewise or stepwise actions.
  • a transformer may be trained to process an embedding of the arch (i.e., to predict transforms for multiple teeth concurrently), to predict a setup.
  • embeddings of individual teeth may be concatenated into a sequence, and then input into the transformer.
  • a VAE may be trained to perform this embedding operation
  • a U-Net may be trained to perform such an embedding
  • a simple dense or fully connected network may be trained, or a combination thereof.
  • the transformer-based techniques of this disclosure may predict an action for an individual tooth, or may predict actions for multiple teeth (e.g., predict transformations for each of multiple teeth).
  • a 3D mesh transformer may include a transformer encoder structure (which may encode oral care data), and may be followed by a transformer decoder structure.
  • the 3D mesh transformer encoder may encode oral care data into a latent representation, which may be combined with attention information (e.g., to concatenate a vector of attention information to the latent representation).
  • the attention information may help the decoder focus on the relevant oral care data during the decoding process (e.g., to focus on tooth order or mesh element connectivity), so that the transformer decoder can generate a useful output for the 3D mesh transformer (e.g., an output which may be used in the generation of an oral care appliance).
  • Either or both of the transformer encoder or transformer decoder may generate a latent representation.
  • the output of the transformer decoder may be reconstructed using a decoder into, for example, one or more tooth transforms for a setup, one or more mesh element labels for segmentation, coordinate systems transforms for use in coordinate system generation, or one or more points of a point cloud or voxels or other mesh elements for another 3D representation).
  • a transformer may include modules such as one or more of: multi-headed attention modules, feed forward modules, normalization modules, linear modules, and softmax modules, and convolution models for latent vector compression, and/or representation.
  • the encoder may be stacked one or more times, thereby further encoding the oral care data, and enabling different representations of the oral care data to be learned (e.g., different latent representations). These representations may be embedded with attention information (which may influence the decoder’s focus to the relevant portions of the latent representation of the oral care data) and may be provided to the decoder in continuous form (e.g., as a concatenation of latent representations - such as latent vectors). In some implementations, the encoded output of the encoder (e.g., latent representations) may be used by downstream processing steps in the generation of oral care appliances.
  • the generated latent representation may be reconstructed into transforms (e.g., for the placement of teeth in setups, or the placement of appliance components or fixture model components), or may be reconstructed into 3D representations (e.g., 3D point clouds, 3D meshes or others disclosed herein).
  • the latent representation which is generated by the transformer e.g., containing continuously encoded attention information
  • Continuously encoded attention information may include attention information which has undergone processing by multiple multi-headed attention modules within the transformer encoder or transformer decoder, to name one example.
  • a loss may be computed for a particular domain using data from that domain. The loss calculation may train the transformer decoder to accurately reconstruct the latent representation into the output data structure pertaining to a particular domain.
  • the decoder when the decoder generates a transform for an orthodontic setup, the decoder may be configured with outputs that describe, for example, the 16 real values which comprise a 4x4 transformation matrix (other data structures for describing transforms are possible). Stated a different way, the latent output generated by the transformer encoder (or transformer decoder) may be used to predict setups tooth transforms for one or more teeth, to place those teeth in setup positions (e.g., either final setups or intermediate stages). Such a transformer encoder (or transformer decoder) may be trained, at least in part using a reconstruction loss (or a representation loss, among others described herein) function, which may compare predicted transforms to ground truth (or reference) transforms.
  • a reconstruction loss or a representation loss, among others described herein
  • the decoder when the decoder generates a transform for a tooth coordinate system, the decoder may be configured with outputs that describe, for example, the 16 real values which comprise a 4x4 transformation matrix (other data structures for describing transforms are possible). Stated a different way, the latent output generated by the transformer encoder (or transformer decoder) may be used to predict local coordinate systems for one or more teeth. Such a transformer encoder (or transformer decoder) may be trained, at least in part using a representation loss (or a reconstruction loss, among others described herein) function, which may compare predicted coordinate systems to ground truth (or reference) coordinate systems.
  • a representation loss or a reconstruction loss, among others described herein
  • the decoder when the decoder generates a 3D point cloud (or other 3D representation - such as 3D mesh, voxelized representation, or the like), the decoder may be configured with outputs that describe, for example, one or more 3D points (e.g., comprising XYZ coordinates). Stated a different way, the latent output generated by the transformer encoder (or transformer decoder) may be used to predict mesh elements for a generated (or modified) 3D representation.
  • Such a transformer encoder may be trained, at least in part using a reconstruction loss (or an LI, L2 or MSE loss, among others described herein) function, which may compare predicted 3D representations to ground truth (or reference) 3D representations.
  • a reconstruction loss or an LI, L2 or MSE loss, among others described herein
  • the decoder when the decoder generates mesh element labels for 3D representation segmentation or 3D representation cleanup, the decoder may be configured with outputs that describe, for example, labels for one or more mesh elements. Stated a different way, the latent output generated by the transformer encoder (or transformer decoder) may be used to predict mesh element labels for mesh segmentation or mesh cleanup. Such a transformer encoder (or transformer decoder) may be trained, at least in part using a cross entropy loss (or others described herein) function, which may compare predicted mesh element labels to ground truth (or reference) mesh element labels.
  • a cross entropy loss or others described herein
  • Multi-headed attention and transformers may be advantageously applied to the setups- generation problem.
  • Multi-headed attention is a module in a 3D transformer encoder network which computes the attention weights for the provided oral care data and produces an output vector with encoded information on how each example of oral care data should attend to each other oral care data in an arch.
  • An attention weight is a quantification of the relationship between pairs of oral care data.
  • a 3D representation of oral care data (e.g., comprising voxels, a point cloud, or a 3D mesh composed of vertices, faces or edges) may be provided to the transformer.
  • the 3D representation may describe the patient's dentition, a fixture model (or components of a fixture model), an appliance (or components of an appliance), or the like.
  • a transformer decoder (or a transformer encoder) may be equipped with multi-head attention. Multi -headed attention may enable the transformer decoder (or transformer encoder) to attend to different portions of the 3D representation of oral care data.
  • multi-headed attention may enable the transformer to attend to mesh elements within local neighborhoods (or cliques), or to attend to global dependencies between mesh elements (or cliques).
  • multi-headed attention may enable a transformer for setups prediction (e.g., a setups prediction model which is based on a transformer) to generate a transform for a tooth, and to substantially concurrently attend to each of the other teeth in the arch while that transform is generated.
  • the transform for each tooth may be generated in light of the poses of one or more other teeth in the arch, leading to a more accurate transform (e.g., a transform which conforms more closely to the ground truth or reference transform).
  • a transformer model may be trained to generate a tooth restoration design.
  • Multi-headed attention may enable the transformer to attend to multiple portions of the tooth (or to the surfaces of the adjacent teeth) while the tooth undergoes the generative process.
  • the transformer for restoration design generation may generate the mesh elements for the incisal edge of an incisor while, at least substantially concurrently, attending to the mesh elements of the mesial, distal, facial or lingual surfaces of the incisor.
  • the result may be the generation of mesh elements to form an incisal edge for the tooth which merges seamlessly with the adjacent surfaces of the tooth.
  • one or more attention vectors may be generated which describe how aspects of the oral care data interacts with other aspects of the oral care data associated with the arch.
  • the one or more attention vectors may be generated to describe how one or more portions of a tooth T1 interact with one or more portions of a tooth T2, a tooth T3, a tooth T4, and so one.
  • a portion of a mesh may be described as a set of mesh elements, as defined herein.
  • the interacting portions of tooth T1 and tooth T2 may be determined, in part, through the calculation of mesh correspondences, as described herein.
  • any of these models may be advantageously applied to the task of setups transform prediction, such as in the models described herein.
  • a transformer may be particularly advantageous in that a transformer may enable the transforms for multiple teeth, or even an entire arch to be generated at once, rather than individually, as may be the case with some other models, such as an encoder structure.
  • attention-free transformers may be used to make predictions based on oral care data.
  • One implementation of the GDL Setups neural network model may include a representation generation module (e.g., containing a U-Net structure, an autoencoder encoder, a transformer encoder, another type of encoder-decoder structure, or an encoder, etc.) which may provide its output to a module which is trained to generate tooth transformers (e.g., a set of fully connected layers with optional skip connections, or an encoder structure) to generate the prediction of a transform for each individual tooth.
  • Skip connections may, in some implementations, connect the outputs of a particular layer in a neural network to the inputs of another later in the neural network (e.g., a layer which is not immediately adjacent to the originating layer).
  • the transform-generation module may handle the transform prediction one tooth at a time.
  • Other implementations may replace this encoder structure with a transformer (e.g., transformer encoder or transformer decoder), which may handle all the predictions for all teeth substantially concurrently.
  • a transformer may be configured to receive a large number of input values, larger than some other neural network models (e.g., than a typical MLP). This is because an increased number of inputs may be accommodated by the transformer, the predictions corresponding to those inputs may be generated substantially concurrently.
  • the representation generation module may provide its output to the transformer, and the transformer may generate the setups transforms for all of the several teeth at once, with the technical advantage of improved accuracy (because the transforms for each tooth is generated in light of the transform for each of the adjacent or nearby teeth - leading to fewer collisions and better conformance with the goals of treatment).
  • a transformer may be trained to output a transformation, such as a transform encoded by a 4x4 matrix (or some other size), a quaternion, a translation vector, Euler angles or some other form.
  • the transformation may place a tooth into a setups pose, may place a fixture model component into a pose suitable for fixture model generation, or may place an appliance component into a pose suitable for appliance generation (e.g., dental restoration appliance, clear tray aligner, etc.).
  • the transform may define a coordinate system for aspects of the patient’s dentition, such as a tooth mesh (e.g., a local coordinate system for a tooth).
  • the inputs to the transformer may first be encoded using a neural network (e.g., a latent representation or embedding may be generated), such as one or more linear layers, and/or one or more convolutional layers.
  • the transformer may first be trained on an offline dataset, and subsequently be trained using a secondary actor-critic network, which may enable online reinforcement learning.
  • Transformers may, in some implementations, enable large model capacity and/or enable an attention mechanism (e.g., the capability to pay attention and respond to certain inputs).
  • the attention mechanisms e.g., multi-headed attention
  • the attention mechanisms that are found within transformers may enable intra-sequence relationships to be encoded into neural network features.
  • Intra-sequence relationships may be encoded, for example, by associating an order number (e.g., 1, 2, 3, etc.) with each tooth in an arch, or by associating an order number with each mesh element in a 3D representation (e.g., of a tooth).
  • intra-sequence relationships may be encoded, for example, by associating an order number (e.g., 1, 2, 3, etc.) with each element in the latent vector.
  • Transformers may be scaled by increasing the number of attention heads and/or by increasing the number of transformer layers. Stated differently, one or more aspects of a transformer may be independently trained to handle discrete tasks, and later combined to allow the resulting transformer to perform all of the tasks for which the individual components had been trained, without degrading the predictive accuracy of the neural network. Scaling a convolutional network may be more difficult, because the models may be less malleable or may be less interchangeable.
  • Convolution has an ability to be rotation and translation invariant, which leads to improved generalization, because a convolution model may not need to account for the manner in which the input data in rotated or translated.
  • Transformers have an ability to be permutation invariant, since intrasequence relationships may be encoded into neural network features.
  • transformers may be combined with convolution-based neural networks, such as by vertically stacking convolution layers and attention layers.
  • Stacking transformer blocks with convolutional blocks enables the resulting structure to have the translation invariance of convolution, and also the permutation invariance of a transformer.
  • Such stacking may improve model capacity and/or model generalization.
  • CoAtNet is an example of a network architecture which combines convolutional and attention-based elements and may be applied to the processing of oral care data.
  • a network for the modification or generation of 3D oral care representations may be trained, at least in part, from CoAtNet (or another model that combines convolution and self-attention/transformers) using transfer learning.
  • the techniques of this disclosure may include operations such as 3D convolution, 3D pooling, 3D unconvolution and 3D unpooling.
  • 3D convolution may aid segmentation processing, for example in down sampling a 3D mesh.
  • 3D un-convolution undoes 3D convolution for example, in a U- Net.
  • 3D pooling may aid the segmentation processing, for example in summarized neural network feature maps.
  • 3D un-pooling undoes 3D pooling for example in a U-Net.
  • These operations may be implemented by way of one or more layers in the predictive or generative neural networks described herein. These operations may be applied directly on mesh elements, such as mesh edges or mesh faces. These operations provide for technical improvements over other approaches because the operations are invariant to mesh rotation, scale, and translation changes. In general, these operations depend on edge (or face) connectivity, therefore these operations remain invariant to mesh changes in 3D space as long as edge (or face) connectivity is preserved. That is, the operations may be applied to an oral care mesh and produce the same output regardless of the orientation, position or scale of that oral care mesh, which may lead to data precision improvement.
  • MeshCNN is a general-purpose deep neural network library for 3D triangular meshes, which can be used for tasks such as 3D shape classification or mesh element labelling (e.g., for segmentation or mesh cleanup). MeshCNN implements these operations on mesh edges. Other toolkits and implementations may operate on edges or faces.
  • neural networks may be trained to operate on 2D representations (such as images). In some implementations of the techniques of this disclosure, neural networks may be trained to operate on 3D representations (such as meshes or point clouds).
  • An intraoral scanner may capture 2D images of the patient's dentition from various views. An intraoral scanner may also (or alternatively) capture 3D mesh or 3D point cloud data which describes the patient's dentition.
  • autoencoders (or other neural networks described herein) may be trained to operate on either or both of 2D representations and 3D representations.
  • a 2D autoencoder (comprising a 2D encoder and a 2D decoder) may be trained on 2D image data to encode an input 2D image into a latent form (such as a latent vector or a latent capsule) using the 2D encoder, and then reconstruct a facsimile of the input 2D image using the 2D decoder.
  • a latent form such as a latent vector or a latent capsule
  • 2D images may be readily captured using one or more of the onboard cameras.
  • 2D images may be captured using an intraoral scanner which is configmed for such a function.
  • 2D image convolution may involve the "sliding" of a kernel across a 2D image and the calculation of elementwise multiplications and the summing of those elementwise multiplications into an output pixel.
  • the output pixel that results from each new position of the kernel is saved into an output 2D feature matrix.
  • neighboring elements e.g., pixels
  • may be in well-defined locations e.g., above, below, left and right
  • a 2D pooling layer may be used to down sample a feature map and summarize the presence of certain features in that feature map.
  • 2D reconstruction error may be computed between the pixels of the input and reconstmcted images.
  • the mapping between pixels may be well understood (e.g., the upper pixel [23, 134] of the input image is directly compared to pixel [23,134] of the reconstructed image, assuming both images have the same dimensions).
  • 2D autoencoder-based techniques of this disclosure is the ease of capturing 2D image data with a handheld device. In some instances, where outside data sources provide the data for analysis, there may be instances where only 2D image data are available. When only 2D image data are available, then analysis using a 2D autoencoder.
  • Modem mobile devices may also have the capability of generating 3D data (e.g., using multiple cameras and stereophotogrammetry, or one camera which is moved around the subject to capture multiple images from different views, or both), which in some implementations, may be arranged into 3D representations such as 3D meshes, 3D point clouds and/or 3D voxelized representations.
  • 3D representations such as 3D meshes, 3D point clouds and/or 3D voxelized representations.
  • the analysis of a 3D representation of the subject may in some instances provide technical improvements over 2D analysis of the same subject.
  • a 3D representation may describe the geometry and/or structure of the subject with less ambiguity than a 2D representation (which may contain shadows and other artifacts which complicate the depiction of depth from the subject and texture of the subject).
  • 3D processing may enable technical improvements because of the inverse optics problem which may, in some instances, affect 2D representations.
  • the inverse optics problem refers to the phenomenon where, in some instances, the size of a subject, the orientation of the subject and the distance between the subject and the imaging device may be conflated in a 2D image of that subject. Any given projection of the subject on the imaging sensor could map to an infinite count of ⁇ size, orientation, distance ⁇ pairings.
  • 3D representations enable the technical improvement in that 3D representations remove the ambiguities introduced by the inverse optics problem.
  • a device that is configmed with the dedicated purpose of 3D scanning such as a 3D intraoral scanner (or a CT scanner or MRI scanner), may generate 3D representations of the subject (e.g., the patient's dentition) which have significantly higher fidelity and precision than is possible with a handheld device.
  • 3D intraoral scanner or a CT scanner or MRI scanner
  • 3D representations of the subject e.g., the patient's dentition
  • the use of a 3D autoencoder is offers technical improvements (such as increased data precision), to extract the best possible signal out of those 3D data (i.e., to get the signal out of the 3D crown meshes used in tooth classification or setups classification).
  • a 3D autoencoder (comprising a 3D encoder and a 3D decoder) may be trained on 3D data representations to encode an input 3D representation into a latent form (such as a latent vector or a latent capsule) using the 3D encoder, and then reconstruct a facsimile of the input 3D representation using the 3D decoder.
  • a latent form such as a latent vector or a latent capsule
  • a 3D convolution may be performed to aggregate local features from nearby mesh elements. Processing may be performed above and beyond the techniques for 2D convolution, to account for the differing count and locations of neighboring mesh elements (relative to a particular mesh element).
  • a particular 3D mesh element may have a variable count of neighbors and those neighbors may not be found in expected locations (as opposed to a pixel in 2D convolution which may have a fixed count of neighboring pixels which may be found in known or expected locations).
  • the order of neighboring mesh elements may be relevant to 3D convolution.
  • a 3D pooling operation may enable the combining of features from a 3D mesh (or other 3D representation) at multiple scales.
  • 3D pooling may iteratively reduce a 3D mesh into mesh elements which are most highly relevant to a given application (e.g., for which a neural network has been trained).
  • 3D pooling may benefit from processing beyond that entailed in 2D convolution, to account for the differing count and locations of neighboring mesh elements (relative to a particular mesh element).
  • the order of neighboring mesh elements may be less relevant to 3D pooling than to 3D convolution.
  • 3D reconstruction error may be computed using one or more of the techniques described herein, such as computing Euclidean distances between corresponding mesh elements, the two meshes. Other techniques are possible in accordance with aspects of this disclosure. 3D reconstruction error may generally be computed on 3D mesh elements, rather than the 2D pixels of 2D reconstruction error. 3D reconstruction error may enable technical improvements over 2D reconstruction error, because a 3D representation may, in some instances, have less ambiguity than a 2D representation (i.e., have less ambiguity in form, shape and/or structure).
  • a 3D representation may be produced using a 3D scanner, such as an intraoral scanner, a computerized tomography (CT) scanner, ultrasound scanner, a magnetic resonance imaging (MRI) machine or a mobile device which is enabled to perform stereophotogrammetry.
  • a 3D representation may describe the shape and/or structure of a subject.
  • a 3D representation may include one or more 3D mesh, 3D point cloud, and/or a 3D voxelized representation, among others.
  • a 3D mesh includes edges, vertices, or faces. Though interrelated in some instances, these three types of data are distinct.
  • the vertices are the points in 3D space that define the boundaries of the mesh. These points would alternatively be described as a point cloud but for the additional information about how the points are connected to each other, as described by the edges.
  • An edge is described by two points and can also be referred to as a line segment
  • a face is described by a number of edges and vertices.
  • a face comprises three vertices, where the vertices are interconnected to form three contiguous edges.
  • Some meshes may contain degenerate elements, such as non-manifold mesh elements, which may be removed, to benefit Other mesh pre-processing operations are possible in accordance with aspects of this disclosure.
  • 3D meshes are commonly formed using triangles, but may in other implementations be formed using quadrilaterals, pentagons, or some other n-sided polygon.
  • a 3D mesh may be converted to one or more voxelized geometries (i.e., comprising voxels), such as in the case that sparse processing is performed.
  • the techniques of this disclosure which operate on 3D meshes may receive as input one or more tooth meshes (e.g., arranged in one or more dental arches). Each of these meshes may undergo pre-processing before being input to the predictive architecture (e.g., including at least one of an encoder, decoder, pyramid encoder-decoder and U-Net). This pre-processing may include the conversion of the mesh into lists of mesh elements, such as vertices, edges, faces or in the case of sparse processing - voxels. For the chosen mesh element type or types (e.g., vertices), feature vectors may be generated. In some examples, one feature vector is generated per vertex of the mesh. Each feature vector may contain a combination of spatial and/or structural features, as specified in the following table:
  • Table 1 discloses non-limiting examples of mesh element features.
  • color or other visual cues/identifiers
  • a mesh element feature in addition to the spatial or structural mesh element features described in Table 1.
  • a point differs from a vertex in that a point is part of a 3D point cloud, whereas a vertex is part of a 3D mesh and may have incident faces or edges.
  • a dihedral angle (which may be expressed in either radians or degrees) may be computed as the angle (e.g., a signed angle) between two connected faces (e.g., two faces which are connected along an edge).
  • a sign on a dihedral angle may reveal information about the convexity or concavity of a mesh surface.
  • a positively signed angle may, in some implementations, indicate a convex surface.
  • a negatively signed angle may, in some implementations, indicate a concave surface.
  • directional curvatures may first be calculated to each adjacent vertex around the vertex. These directional curvatures may be sorted in circular order (e.g., 0, 49, 127, 210, 305 degrees) in proximity to the vertex normal vector and may comprise a subsampled version of the complete curvature tensor. Circular order means: sorted in by angle around an axis.
  • the sorted directional curvatures may contribute to a linear system of equations amenable to a closed form solution which may estimate the two principal curvatures and directions, which may characterize the complete curvature tensor.
  • a voxel may also have features which are computed as the aggregates of the other mesh elements (e.g., vertices, edges and faces) which either intersect the voxel or, in some implementations, are predominantly or fully contained within the voxel. Rotating the mesh may not change structural features but may change spatial features.
  • the term “mesh” should be considered in a nonlimiting sense to be inclusive of 3D mesh, 3D point cloud and 3D voxelized representation.
  • mesh element features apart from mesh element features, there are alternative methods of describing the geometry of a mesh, such as 3D keypoints and 3D descriptors. Examples of such 3D keypoints and 3D descriptors are found in “TONIONI A, et al. in ‘Learning to detect good 3D keypoints.’, Int J Comput. Vis. 2018 Vol .126, pages 1-20 3D keypoints and 3D descriptors may, in some implementations, describe extrema (either minima or maxima) of the surface of a 3D representation.
  • one or more mesh element features may be computed, at least in part, via deep feature synthesis (DFS), e.g. as described in: J. M. Kanter and K. Veeramachaneni, "Deep feature synthesis: Towards automating data science endeavors," 2015 IEEE International Conference on Data Science and Advanced Analytics (DSAA), 2015, pp.
  • DFS deep feature synthesis
  • Predictive models which may operate on feature vectors of the aforementioned features include but are not limited to: GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, Tooth Classification, Setups Classification, Setups Comparison, VAE Mesh Element Labeling, MAE Mesh In-filling, Mesh Reconstruction Autoencoder, Validation Using Autoencoders, Mesh Segmentation, Coordinate System Prediction, Mesh Cleanup, Restoration Design Generation, Appliance Component Generation and/or Placement, and Archform Prediction, may input.
  • such feature vectors may be presented to one or more internal layers of a neural network which is part of one or more of those predictive models.
  • mesh element features may convey aspects of a 3D representation’s surface shape and/or structure to the neural network models of this disclosure.
  • Each mesh element feature describes distinct information about the 3D representation that may not be redundantly present in other input data that are provided to the neural network. For example, a vertex curvature may quantify aspects of the concavity or convexity of the surface of a 3D representation which would not otherwise be understood by the network.
  • mesh element features may provide a processed version of the structure and/or shape of the 3D representation; data that would not otherwise be available to the neural network. This processed information is often more accessible, or more amenable for encoding by the neural network.
  • a system implementing the techniques disclosed herein has been utilized to mn a number of experiments on 3D representations of teeth. For example, mesh element features have been provided to a representation generation neural network which is based on a U-Net model, and also to a representation generation model based on a variational autoencoder with continuous normalizing flows.
  • tooth movements specify one or more transformations that can be encoded in various ways to specify tooth positions and orientations within the setup and are applied to 3D representations of teeth.
  • the tooth positions can be cartesian coordinates of a tooth's canonical origin location which is defined in some semantic context.
  • Tooth orientations can be represented as rotation matrices, unit quaternions, or other 3D rotation representations such as Euler angles with respect to a frame of reference (either global or local).
  • Dimensions are real valued 3D spatial extents and gaps can be binary presence indicators or real valued gap sizes between teeth especially in instances when certain teeth are missing.
  • tooth rotations may be described by 3x3 matrices (or by matrices of other Tooth position and rotation information may, in some implementations, be combined into the same transform matrix, for example, as a 4x4 matrix, which may reflect homogenous coordinates, in some instances, affine spatial transformation matrices may be used to describe tooth transformations, for example, the transformations which describe the maloccluded pose of a tooth, an intermediate pose of a tooth and/or a final setup pose of a tooth. Some implementations may use relative coordinates, where setup transformations are predicted relative to malocclusion coordinate systems (e.g., a malocclusion-to-setup transformation is predicted instead of a setup coordinate system directly).
  • Other implementations may use absolute coordinates, where setup coordinate systems are predicted directly for each tooth.
  • transforms can be computed with respect to the centroid of each tooth mesh (vs the global origin), which is termed “relative local.”
  • relative local coordinates Some of the advantages of using relative local coordinates include eliminating the need for malocclusion coordinate systems (landmarking data) which may not be available for all patient case datasets.
  • absolute coordinates Some of the advantages of using absolute coordinates include simplifying the data preprocessing as mesh data are originally represented as relative to the global origin.
  • tooth position encoding and tooth orientation encoding may, in some implementations, also apply one or more of the neural networks models of the present disclosure, including but not limited to: GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, FDG Setups, Setups Classification, Setups Comparison, VAE Mesh Element Labeling, MAE Mesh Infilling, Mesh Reconstruction VAE, and Validation Using Autoencoders.
  • convolution layers in the various 3D neural networks described herein may use edge data to perform mesh convolution.
  • edge information reduces or potentially eliminates the model’s sensitivity to different input orders of 3D elements.
  • the convolution layers may use vertex data to perform mesh convolution.
  • vertex information is advantageous in that there are typically fewer vertices than edges or faces, thereby enabling vertex-oriented processing to function at a lower processing overhead and lower computational resource cost.
  • the convolution layers may use face data to perform mesh convolution.
  • the convolution layers may use voxel data to perform mesh convolution.
  • voxel information provides a technical improvement in that, depending on the granularity chosen, there may be significantly fewer voxels to process compared to the vertices, edges or faces in the mesh. Sparse processing (with voxels) may lead to a lower processing overhead and lower computational cost (especially in terms of computer memory or RAM usage).
  • the orthodontic metrics may be used to quantify the physical arrangement of an arch of teeth for the purpose of orthodontic treatment (as opposed to restoration design metrics - which pertain to dentistry and describe the shape and/or form of one or more pre-restoration teeth, for the purpose of supporting dental restoration). These orthodontic metrics can measure how badly maloccluded the arch is, or conversely the metrics can measure how correctly arranged the teeth are.
  • the GDL Setups model RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity SetupsFDG Setups may incorporate one or more of these orthodontic metrics, or other similar or related orthodontic metrics.
  • such orthodontic metrics may be incorporated into the feature vector for a mesh element, where these perelement feature vectors are provided to the setups prediction network as inputs.
  • such orthodontic metrics may be directly consumed by a generator, an MLP, a transformer, or other neural network as direct inputs (such as presented in one or more input vectors of real numbers S, such as described elsewhere in this disclosure). The use of such orthodontic metrics in the training of the generator may improve the performance (e., correctness) of the , resulting in predicted transforms which place teeth more nearly in the correct final setups poses than would otherwise be possible.
  • Such orthodontic metrics may be consumed by an encoder structure or by a U-Net structure (in the case of GDL Setups).
  • Such orthodontic metrics may be consumed by an autoencoder, variational autoencoder, masked autoencoder or regularized autoencoder (in the case of the VAE Setups, VAE Mesh Element Labelling, MAE Mesh In-Filling).
  • Such orthodontic metrics may be provided to a neural network which generates action predictions as a part of a reinforcement learning RL Setups model, orthodontic metrics a label to a setup arch (e.g., labels such as mal, staging or final setup). This description is non-limiting, as the orthodontic metrics may also be incorporated in other ways into the various techniques of this disclosure.
  • the various loss calculations of the present disclosure may, in some examples, incorporate one or more orthodontic metrics, with the of the correctness of the resulting neural
  • An orthodontic metric may be used to directly compare a predicted example to the corresponding ground truth example (such as is with metrics)
  • one or more orthodontic metrics may be incorporated into a loss computation.
  • Such an orthodontic metric may be computed on the predicted example, and then the orthodontic metric would also be computed on the ground tmth example. These two orthodontic metrics results would then be consumed by the loss computation, with the advantage of improving the performance of the resulting neural network.
  • one or more orthodontic metrics pertaining to the alignment of two or more adjacent teeth may be computed and incorporated into a loss function, for example, to train, at least in part, a setups prediction neural network
  • such an orthodontic metric may influence the network to align the mesial surface of a tooth with the distal surface of an adjacent tooth.
  • Backpropagation is an example algorithm by which a neural network may be trained using one or more loss values.
  • one or more orthodontic metrics may be used to evaluate the predicted output of a neural network, such as a setups prediction Such metric(s) may enable the training algorithm to determine how close the predicted output is to an acceptable output, for example, in a quantified sense.
  • this use of an orthodontic metric may enable a loss value to be computed which does not depend entirely on a comparison to a ground truth
  • such a use of an orthodontic metric may enable loss calculation and network training to proceed without the need for a comparison against a ground truth example.
  • loss may be computed based on a general principle or specification for the predicted output (such as a setup) rather than tying loss calculation to a specific ground tmth example (which may have been defined by a particular doctor, clinician, or technician, whose treatment may differ from that of other technicians or doctors).
  • a specific ground tmth example which may have been defined by a particular doctor, clinician, or technician, whose treatment may differ from that of other technicians or doctors.
  • such an orthodontic metric may be defined based on a Frechet Inception Distance) score.
  • An orthodontic metric that can be computed using tensors may be especially advantageous when training one of the neural networks of the present disclosure, because tensor operations may promote efficient computations. The more efficient (and faster) the computation, the faster the rate at which training can proceed.
  • an error pattern may be identified in one or more predicted outputs of an ML model (e.g., a transformation matrix for a predicted tooth setup, a labelling of mesh elements for mesh cleanup, an addition of mesh elements to a mesh for the purpose of mesh in-filling, a classification label for a setup, a classification label for a tooth mesh, etc.).
  • One or more orthodontic metrics may be selected to become an input to the next round of ML model training, to address any pattern of errors or deficiencies which may be identified in the one or more predicted outputs.
  • [00120] may be defined relative to an archfrom coordinate frame, the LDE coordinate system.
  • a point may be described using an LDE coordinate frame relative to an archform, where L, D and E correspond to: 1) Length along the curve of the archform, 2) Distance away from the archform, and 3) distance in the direction perpendicular to the L and D axes (which may be termed Eminence), respectively.
  • L, D and E correspond to: 1) Length along the curve of the archform, 2) Distance away from the archform, and 3) distance in the direction perpendicular to the L and D axes (which may be termed Eminence), respectively.
  • Various of the OM and other techniques of the present disclosure may compute collisions between 3D representations (e.g., of oral care objects, such as teeth).
  • Such collisions may be computed as at least one of: 1) penetration distance between 3D tooth representations, 2) count of overlapping mesh elements between 3D tooth representations, and 3) volume of overlap between 3D tooth representations.
  • an OM may be defined to quantify the collision of two or more 3D representations of oral care structures, such as teeth.
  • Some optimization algorithms, such as setups prediction techniques, may seek to minimize collisions between oral care structures (such as teeth). Between-arch orthodontic metrics are described as follows.
  • a 3D tooth orientation vector may be calculated using the tooth's mesial-distal axis.
  • a 3D vector which may be tangent vector to the archform at the position of the tooth may also be calculated.
  • the XY components i.e., which may be 2D vectors
  • Cosine similarity may be used to calculate the 2D orientation difference (angle) between the archform tangent and the tooth's mesial-distal axis.
  • the absolute difference may be calculated between each tooth’s X-coordinate and the global coordinate reference frame’s X-axis, delta may indicate the arch asymmetry for a given tooth pair.
  • the result of such a calculation may be the mean X-axis delta of one or more tooth-pairs from the arch. This calculation may, in some implementations, be performed relative to the Y-axis withy-coordinates (and/or relative to the Z axis with Z-coordinates).
  • Archform D-axis Differences dimension difference (e., the positional difference in the facial -lingual direction) between two arch states, for one or more teeth, some implementations, a dictionary of the D- direction tooth movement for each tooth, with tooth UNS number as the key. May use the LDE coordinate system relative to an archform.
  • Archform (Lower) Length Ratio the ratio between the current lower arch length and the arch length as it was in the original maloccluded lower arch.
  • Archform (Upper) Length Ratio the ratio between the current upper arch length and the arch length as it was in the original maloccluded upper arch.
  • Archform Parallelism (Full arch) - For at least one local tooth coordinate system origin in the upper arch, the one or more nearest origins (e.g., tooth local coordinate system origins) in the lower arch.
  • the two nearest origins may be used, the straight line distance from the upper arch point to the line formed between the origins of the two teeth in the opposing (lower) arch. May return the standard deviation of the set of “point-to-line” distances mentioned above, where the set may be composed of the point-to-line distances for each tooth in the arch.
  • This metric may share some computational with the archform_parallelism_global orthodontic metric, except that this metric may input the mean distance from a tooth origin to the line formed by the neighboring teeth in opposing arches (e.g., a tooth in the upper arch and the corresponding tooth in the lower arch). The mean distance may be computed for one or more such pairs of teeth. In some implementations, this may be computed for all pairs of teeth. Then the mean distance may be subtracted from the distance that is computed for each tooth pair. This OM may yield the deviation of a tooth from a “typical” tooth parallelism in the arch.
  • n-element list for each tooth (e.g. n may equal 2).
  • Such an n-element vector may be computed for each molar and each premolar in the upper and lower arches.
  • the lingual cusps may be projected onto the plane (i.e., at this point the angle of inclination may be determined). By performing an additional projection, the approximate vertical distance between the lingual cusps and the buccal cusps may be computed. This distance may be used as the buccolingual inclination OM.
  • Canine Overbite The upper and lower canines may be identified.
  • the first premolar for the given side of the mouth may be identified.
  • a distance may be computed between the upper canine and the lower canine, and also between the upper pre-molar and the lower pre-molar.
  • the average ( median, or mode or some other statistic) may be computed for the measured distances.
  • the z- component of this result indicates the degree of overbite.
  • Overbite may be computed between any tooth in one arch and the corresponding tooth in the other arch.
  • Canine Overjet Contact - May calculate the collisions (e.g., collision distances) between pairs of canines on opposing arches.
  • Canine Overjet Contact KDE - May take an orthodontic metric score for the current patient case as input, and may convert that score into to a log-likelihood using a previously trained kernel density estimation (KDE) model or distribution. This operation may yield information about where in the distribution of "typical" values this patient case lies.
  • KDE kernel density estimation
  • Canine Overjet - This OM may share some computational steps with the canine overbite OM.
  • average distances may be computed.
  • the distance calculation may compute the Euclidean distance of the XY components of a tooth in the upper arch and a tooth in the lower arch, to yield oveget (i.e., as opposed to computing the difference in Z-components, as may be performed for canine overbite).
  • Oveget may be computed between any tooth in one arch and the corresponding tooth in the other arch.
  • Canine Class Relationship (also applies to first, second and third molars) -
  • This OM may, in some implementations comprise two functions (e.g., written in Python).
  • get_canine_landmarks() Get landmarks for each tooth which may be used to compute the class relationship, and then, in some implementations, map those landmarks onto the global coordinate space so that measurements may be made between teeth.
  • class_relationship_score_by_side() May compute the average position of at least one landmark on at least one tooth in the lower arch, and may compute the same for the upper arch.
  • This OM may compute how far forward or behind the tooth is positioned on the 1-axis relative to the tooth or teeth of interest in the opposing arch.
  • Crossbite - Fossa in at least one upper molar may be located by finding the halfway point between distal and mesial marginal ridge saddles of the tooth.
  • a lower molar cusp may lie between the marginal ridges of the corresponding upper molar.
  • This OM may compute a vector from the upper molar fossa midpoint to the lower molar cusp. This vector may be projected onto the d-axis of the archform, yielding a lateral measure of distance from the cusp to the fossa. This distance may define the crossbite magnitude.
  • This OM may identify the leftmost and rightmost edges of a tooth, and may identify the for that tooth’s neighbor.
  • the OM may then draw a vector from the leftmost edge of the tooth to the leftmost edge of the tooth’s neighbor.
  • the OM may then draw a vector from the rightmost edge of the tooth to the rightmost edge of the tooth’s neighbor.
  • the OM may then calculates the linear fit error between the two vectors.
  • Such a calculation may involve making two vectors:
  • Vec tooth right tooths leftside to left tooths leftside
  • Vec neighbor right tooths rightside to left tooths leftside
  • EdgeAlignment score 1 - abs(dot(Vec_tooth, Vec neighbor)) ).
  • a score of 0 may indicate perfect alignment.
  • a score of 1 may mean perpendicular alignment.
  • Incisor Interarch Contact KDE - May identify the deviation of the IncisorlnterarchContact from the mean of a modeled distribution of such statistics across a dataset of one or more other patient cases.
  • Leveling - May compute a measure of leveling between a tooth and its neighbor.
  • This OM may calculate the difference in height between two or more neighboring teeth. For molars, this OM may use the midpoint between the mesial and distal saddle ridges as the height of the molar. For non-molar teeth, this OM may use the length of the crown from gums to tip. In some implementations, the tip may be the origin of the local coordinate space of the tooth. Other implementations may place the origin in other locations. A simple subtraction between the heights of neighboring teeth may yield the leveling delta between the teeth (e.g., by comparing Z components).
  • Midline - May compute the position of the midline for the upper incisors and/or the lower incisors, and then may compute the distance between them.
  • Molar Interarch Contact KDE - May compute a molar interarch contact score (i.e., a collision depth or other type of collision), and then may identify where that score lies in a pre-defined KDE (distribution) built from representative cases.
  • a molar interarch contact score i.e., a collision depth or other type of collision
  • this OM may identify one or more landmarks (e.g., mesial cusp, or central cusp, etc.). Get the tooth transform for that tooth. For each cusp on the current tooth, the cusp may be scored according to how well the cusp contacts the neighboring (corresponding) tooth in the opposite arch. A vector may be found from the cusp of the tooth in question to the vertical intersection point in the corresponding tooth of the opposing arch. The distance and/or direction (i.e., up or down) to the opposing arch may be computed. A list may be returned that contains the resulting signed distances, one for each cusp on the tooth in question.
  • landmarks e.g., mesial cusp, or central cusp, etc.
  • Overbite The upper and lower central incisors may be compared along the z-axis. The difference along the z-axis may be used as the overbite score.
  • Overjet The upper and lower central incisors may be compared along the y-axis. The difference along the y-axis may be used as the oveijet score.
  • Molar Interarch Contact - May calculate the contact score between molars, and may use collision measurement(s) (such as collision depth).
  • Root Movement d The tooth transforms for an initial state and a next state may be recieved.
  • the archform axes at a point L along the archform may be computed.
  • This OM may return a distance moved along the d-axis. This may be accomplished by projecting the root pivot point onto the d-axis.
  • Root Movement 1 The tooth transforms for an initial state and a next state may be received.
  • the archform axes at a point L along the archform may be computed.
  • This OM may return a distance moved along the 1-axis. This may be accomplished by projecting the root pivot point onto the 1-axis.
  • Spacing May compute the spacing between each tooth and its neighbor.
  • the transforms and meshes for the arch may be received.
  • the left and right edges of each tooth mesh may be computed.
  • One or more points of interest may be transformed from local coordinates into the global arch coordinate frame.
  • the spacing may be computed in a plane (e.g., the XY plane) between each tooth and its neighbor to the "left”.
  • Torque - May compute torque (i.e., rotation around and axis, such as the x-axis). For one or more teeth, one or more rotations may be converted from Euler angles into one or more rotation matrices. A component (such as a x-component) of the rotations may be extracted and converted back into Euler angles. This x- component may be interpreted as the torque for a tooth. A list maybe returned which contains the torque for one or more teeth, and may be indexed by the UNS number of the tooth.
  • the neural networks of this disclosure may exploit one or more benefits of the operation of parameter tuning, whereby the inputs and parameters of a neural network are optimized to produce more data-precise results.
  • One parameter which may be tuned is neural network learning rate (e.g., which may have values such as 0.1, 0.01, 0.001, etc.).
  • Data augmentation schemes may also be tuned or optimized, such as schemes where “shiver” is added to the tooth meshes before being input to the neural network (i.e., small random rotations, translations and/or scaling may be applied to vary the dataset and make the neural network robust to variations in data).
  • FIG. 1 shows a data augmentation method that systems of this disclosure may apply to 3D oral care representations.
  • a non-limiting example of a 3D oral care representation is a tooth mesh or a set of tooth meshes.
  • Tooth data 100 e.g., 3D meshes
  • the systems of this disclosure may generate copies of the tooth data 100 (102).
  • the systems of this disclosure may apply one or more stochastic rotations to the tooth data 100 (104).
  • the systems of this disclosure may apply stochastic translations to the tooth data 100 (106).
  • the systems of this disclosure may apply stochastic scaling operations to the tooth data 100 (108).
  • the systems of this disclosure may apply stochastic perturbations to one or more mesh elements of the tooth data 100 (110).
  • the systems of this disclosure may output augmented tooth data 112 that are formed by way of the method of FIG. 1.
  • a linear activation function may be well suited to some regression applications (among other applications), in an output layer.
  • a sigmoid/logistic activation function may be well suited to some binary classification applications (among other applications), in an output layer.
  • a softmax activation function may be well suited to some multiclass classification applications (among other applications), in an output layer.
  • a sigmoid activation function may be well suited to some multilabel classification applications (among other applications), in an output layer.
  • a ReLU activation function may be well suited in some convolutional neural network (CNN) applications (among other applications), in a hidden layer.
  • CNN convolutional neural network
  • additional methods may be employed to update weights, in addition to or in place of the techniques described above. These additional methods include the Levenberg-Marquardt method and/or simulated annealing.
  • the backpropagation algorithm is used to assign the results of loss calculation back into the network so that network weights can be adjusted, and learning can progress.
  • the neural networks of the present disclosure may embody part or all of a variety of different neural network models. Examples include the U-Net architecture, multi-later perceptron (MLP), transformer, pyramid architecture, recurrent neural network (RNN), autoencoder, variational autoencoder, regularized autoencoder, conditional autoencoder, capsule network, capsule autoencoder, stacked capsule autoencoder, denoising autoencoder, sparse autoencoder, conditional autoencoder, long/short term memory (LSTM), gated recurrent unit (GRU), deep belief network (DBN), deep convolutional network (DCN), deep convolutional inverse graphics network (DCIGN), liquid state machine (LSM), extreme learning machine (ELM), echo state network (ESN), deep residual network (DRN), Kohonen network (KN), neural Turing machine (NTM), or generative adversarial network (GAN).
  • U-Net architecture multi-later perceptron (MLP), transformer, pyramid architecture, recurrent
  • Autoencoders that can be used in accordance with aspects of this disclosure include but are not limited to: AtlasNet, FoldingNet and 3D-PointCapsNet. Some autoencoders may be implemented based on PointNet.
  • a transform may be described by a 9x1 transformation vector (e.g., that specifies a translation vector and a quaternion).
  • a transform may be described by a transformation matrix (e.g., a 4x4 affine transformation matrix).
  • systems of this disclosure may implement a principal components analysis (PCA) on an oral care mesh, and use the resulting principal components as at least a portion of the representation of the oral care mesh in subsequent machine learning and/or other predictive or generative processing.
  • PCA principal components analysis
  • Systems of this disclosure may implement end-to-end training.
  • Some of the end-to-end training-based techniques of this disclosure may involve two or more neural networks, where the two or more neural networks are trained together (i.e., the weights are updated concurrently during the processing of each batch of input oral care data).
  • End-to-end training may, in some implementations, be applied to setups prediction by concurrently training a neural network which leams a representation of the teeth, along with a neural network which generates the tooth transforms.
  • a neural network (e.g., a U-Net) may be trained on a first task (e.g., such as coordinate system prediction).
  • the neural network trained on the first task may be executed to provide one or more of the starting neural network weights for the training of another neural network that is trained to perform a second task (e.g., setups prediction).
  • the first network may learn the low-level neural network features of oral care meshes and be shown to work well at the first task.
  • the second network may exhibit faster training and/or improved performance by using the first network as a starting point in training.
  • Certain layers may be trained to encode neural network features for the oral care meshes that were in the training dataset.
  • These layers may thereafter be fixed (or be subjected to minor changes over the course of training) and be combined with other neural network components, such as additional layers, which are trained for one or more oral care tasks (such as setups prediction).
  • additional layers which are trained for one or more oral care tasks (such as setups prediction).
  • a portion of a neural network for one or more of the techniques of the present disclosure may receive initial training on another task, which may yield important learning in the trained network layers. This encoded learning may then be built upon with further task-specific training of another network.
  • transfer learning may be used for setups prediction, as well as for other oral care applications, such as mesh classification (e.g., tooth or setups classification), mesh element labeling, mesh element in-filling, procedure parameter imputation, mesh segmentation, coordinate system prediction, restoration design generation, mesh validation (for any of the applications disclosed herein).
  • mesh classification e.g., tooth or setups classification
  • mesh element labeling e.g., mesh element in-filling
  • procedure parameter imputation e.g., mesh element in-filling
  • mesh segmentation e.g., procedure parameter imputation
  • coordinate system prediction e.g., coordinate system prediction
  • restoration design generation for any of the applications disclosed herein.
  • a neural network trained to output predictions based on oral care meshes may first be partially trained on one of the following publicly available datasets, before being further trained on oral care data: Google PartNet dataset, ShapeNet dataset, ShapeNetCore dataset, Princeton Shape Benchmark dataset, ModelNet dataset, ObjectNet3D dataset, ThingilOK dataset (which is especially relevant to 3D printed parts validation), ABC: A Big CAD Model Dataset For Geometric Deep Learning, ScanObjectNN, VOCASET, 3D-FUTURE, MCB: Mechanical Components Benchmark, PoseNet dataset, PointCNN dataset, MeshNet dataset, MeshCNN dataset, PointNet++ dataset, PointNet dataset, or PointCNN dataset.
  • a neural network which was previously trained on a first dataset may subsequently receive further training on oral care data and be applied to oral care applications (such as setups prediction).
  • Transfer learning maybe employed to further train any of the following networks: GCN (Graph Convolutional Networks), PointNet, ResNet or any of the other neural networks from the published literature which are listed above.
  • a first neural network may be trained to predict coordinate systems for teeth (such as by using the techniques described in WO2022123402A1 or US Provisional Application No. US63/366492).
  • a second neural network may be trained for setups prediction, according to any of the setups prediction techniques of the present disclosure (or a combination of any two or more of the techniques described herein).
  • Transfer learning may assign at least a portion of the knowledge or capability of the first neural network to the second neural network. As such, transfer learning may provide the second neural network an accelerated training phase to reach convergence.
  • the training of the second network may, after being augmented with the transferred learning, then be completed using one or more of the techniques of this disclosure.
  • Systems of this disclosure may train ML models with representation learning.
  • representation learning e.g., neural network that predicts a transform for use in setups prediction
  • the generative network e.g., neural network that predicts a transform for use in setups prediction
  • the representation generation model extracts hierarchical neural network features and/or reconstruction characteristics of an inputted representation (e.g., a mesh or point cloud) through loss calculations or network architectures chosen for that purpose).
  • Reconstruction characteristics may comprise values in of a latent representation (e.g., a latent vector) that describe aspects of the shape and/or structure of the 3D representation that was provided to the representation generation module that generated the latent representation.
  • the weights of the encoder module of a reconstruction autoencoder may be trained to encode a 3D representation (e.g., a 3D mesh, or others described herein) into a latent vector representation (e.g., a latent vector).
  • the capability to encode a large set (e.g., hundreds, thousands or millions) of mesh elements into a latent vector may be learned by the weights of the encoder.
  • Each dimension of that latent vector may contain a real number which describes some aspect of the shape and/or structure of the original 3D representation.
  • the weights of the decoder module of the reconstruction autoencoder may be trained to reconstruct the latent vector into a close facsimile of the original 3D representation.
  • the capability to interpret the dimensions of the latent vector, and to decode the values within those dimensions may be learned by the decoder.
  • the encoder and decoder neural network modules are trained to perform the mapping of a 3D representation into a latent vector, which may then be mapped back (or otherwise reconstructed) into a 3D representation that is substantially similar to an original 3D representation for which the latent vector was generated.
  • examples of loss calculation may include KL-divergence loss, reconstruction loss or other losses disclosed herein.
  • Representation learning may reduce the size of the dataset required for training a model, because the representation model learns the representation, enabling the generative network to focus on learning the generative task.
  • the result may be improved model generalization because meaningful neural network features of the input data (e.g., local and/or global features) are made available to the generative network.
  • a first network may learn the representation, and a second network may make the predictive decision.
  • each of the networks may generate more accurate results for their respective tasks than with a single network which is trained to both learn a representation and make a decision.
  • transfer learning may first train a representation generation model. That representation generation model (in whole or in part) may then be used to pre-train a subsequent model, such as a generative model (e.g., that generates transform predictions).
  • a representation generation model may benefit from taking mesh element features as input, to improve the capability of a second ML module to encode the structure and/or shape of the inputted 3D oral care representations in the training dataset.
  • One or more of the neural networks models of this disclosure may have attention gates integrated within. Attention gate integration provides the advantage of enabling the associated neural network architecture to focus resources on one or more input values.
  • an attention gate may be integrated with a U-Net architecture, with the advantage of enabling the U-Net to focus on certain inputs, such as input flags which correspond to teeth which are meant to be fixed (e.g.,. prevented from moving) timing orthodontic treatment (or which require other special handling).
  • An attention gate may also be integrated with an encoder or with an autoencoder (such as VAE or capsule autoencoder) to improve predictive accuracy, in accordance with aspects of this disclosure.
  • attention gates can be used to configure a machine learning model to give higher weight to aspects of the data which are more likely to be relevant to correctly generated outputs.
  • attention gates or mechanisms
  • the quality and makeup of the training dataset for a neural network can impact the performance of the neural network in its execution phase.
  • Dataset filtering and outlier removal can be advantageously applied to the training of the neural networks for the various techniques of the present disclosure (e.g., for the prediction of final setups or intermediate staging, for mesh element labeling or a neural network for mesh in-filling, for tooth reconstruction, for 3D mesh classification, etc.), because dataset filtering and outlier removal may remove noise from the dataset.
  • dataset filtering and outlier removal may remove noise from the dataset.
  • the mechanism for realizing an improvement is different than using attention gates, that ultimate outcome is that this approach allows for the machine learning model to focus on relevant aspects of the dataset, and may lead to improvements in accuracy similar to improvements in accuracy realized vis-a-vis attention gates.
  • a patient case may contain at least one of a set of segmented tooth meshes for that patient, a mal transform for each tooth, and/or a ground tmth setup transform for each tooth.
  • a patient case may contain at least one of a set of segmented tooth meshes for that patient, a mal transform for each tooth, and/or a set of ground truth intermediate stage transforms for each tooth.
  • a training dataset may exclude patient cases which contact passive stages (i.e., stages where the teeth of an arch do not move).
  • the dataset may exclude cases where passive stages exist at the end of treatment.
  • a dataset may exclude cases where overcrowding is present at the end of treatment (i.e., where the oral care provider, such as an orthodontist or dentist) has chosen a final setup where the tooth meshes overlap to some degree.
  • the dataset may exclude cases of a certain level (or levels) of difficulty (e.g., easy, medium and hard).
  • the dataset may include cases with zero pinned teeth (or may include cases where at least one tooth is pinned).
  • a pinned tooth may be designated by a technician as they design the treatment to stop the various tools from moving that particular tooth.
  • a dataset may exclude cases without any fixed teeth (conversely, where at least one tooth is fixed).
  • a fixed tooth may be defined as a tooth that shall not move in the course of treatment.
  • a dataset may exclude cases without any pontic teeth (conversely, cases in which at least one tooth is pontic).
  • a pontic tooth may be described as a “ghost” tooth that is represented in the digital model of the arch but is either not actually present in the patient’ s dentition or where there may be a small or partial tooth that may benefit from future work (such as the addition of composite material through a dental restoration appliance).
  • the advantage of including a pontic tooth in a patient case is to leave space in the arch as a part of a plan for the movements of other teeth, in the course of orthodontic treatment.
  • a pontic tooth may save space in the patient’s dentition for future dental or orthodontic work, such as the installation of an implant or crown, or the application of a dental restoration appliance, such as to add composite material to an existing tooth that is too small or has an undesired shape.
  • the dataset may exclude cases where the patient does not meet an age requirement (e.g., younger than 12). In some implementations, the dataset may exclude cases with interproximal reduction (IPR) beyond a certain threshold amount (e.g., more than 1.0 mm).
  • IPR interproximal reduction
  • the dataset to train a neural network to predict setups for clear tray aligners (CTA) may exclude patient cases which are not related to CTA treatment.
  • the dataset to train a neural network to predict setups for an indirect bonding tray product may exclude cases which are not related to indirect bonding tray treatment. In some implementations, the dataset may exclude cases where only certain teeth are treated.
  • a dataset may comprise of only cases where at least one of the following are treated: anterior teeth, posterior teeth, bicuspids, molars, incisors, and/or cuspids.
  • the mesh comparison module may compare two or more meshes, for example for the computation of a loss function or for the computation of a reconstruction error. Some implementations may involve a comparison of the volume and/or area of the two meshes. Some implementations may involve the computation of a minimum distance between corresponding vertices/faces/edges/voxels of two meshes. For a point in one mesh (vertex point, mid-point on edge, or triangle center, for example) compute the minimum distance between that point and the corresponding point in the other mesh.
  • the open-source software packages CloudCompare and MeshLab each have mesh comparison tools which may play a role in the mesh comparison module for the present disclosure.
  • a Hausdorff Distance may be computed to quantify the difference in shape between two meshes.
  • the open-source software tool Metro developed by the Visual Computing Lab, can also play a role in quantifying the difference between two meshes.
  • the following paper describes the approach taken by Metro, which may be adapted by the neural networks applications of the present disclosure for use in mesh comparison and difference quantification: "Metro: measuring error on simplified surfaces" by P. Cignoni, C. Rocchini and R. Scopigno, Computer Graphics Forum, Blackwell Publishers, vol. 17(2), June 1998, pp 167-174.
  • Some techniques of this disclosure may incorporate the operation of, for one or more points on the first mesh, projecting a ray normal to the mesh surface and calculating the distance before that ray is incident upon the second mesh.
  • the lengths of the resulting line segments may be used to quantify the distance between the meshes.
  • the distance may be assigned a color based on the magnitude of that distance and that color may be applied to the first mesh, by way of visualization.
  • the setups prediction techniques described herein may generate a transform to place a tooth in a setup pose.
  • a predicted transform may entail both the position and the orientation of the tooth, which is a significant improvement over existing techniques which use one neural network to generate a position prediction and another neural network to generate a pose prediction.
  • the predicted position and the predicted orientation affect each other. Generating the predicted position and the predicted orientation substantially concurrently offers improvements in predictive accuracy relative to generating predicted position and predicted orientation separately (e.g., predicting one without the benefit of the other).
  • the MLP Setups, VAE Setups, and Capsule Setups models of the present disclosure improve upon existing techniques with the addition of (among other things) a latent space input: either the latent space vector A of an oral care mesh or the latent capsule T of an oral care mesh.
  • a latent space input either the latent space vector A of an oral care mesh or the latent capsule T of an oral care mesh.
  • Prior setups prediction techniques did not train a reconstruction autoencoder to generate representations of teeth, and therefore could not verify the correctness of their outputs.
  • the advantage of using a reconstruction autoencoder to generate tooth representations is that the latent representation (e.g., A or T) may be reconstructed by the reconstruction autoencoder.
  • Reconstruction error (as described herein) may be computed, to demonstrate the correctness of the latent encoding (e.g., to demonstration that the latent representation correctly describes the shape and/or structure of the tooth). Results with a high reconstruction error may be excluded from downstream (e.g., further or additional) processing, which leads to a more accurate system as a whole. Either or both of A and T may be reconstructed (via a decoder) into a facsimile of an inputted oral care 3D representation (e.g., an inputted tooth mesh). One or more latent space vectors A (or latent capsules T) may be provided to the MLP Setups model.
  • One or more latent space vectors A may also be provided to the VAE Setups model.
  • One or more latent capsules T may also be provided to the Capsule Autoencoder Setups model.
  • This latent space vector A (or latent capsule T) may be reconstmcted into a close facsimile of the input tooth mesh through the operation of a decoder that has been trained for that task.
  • the latent space vector A (or latent capsule T) is powerful because, although A (or T) is relatively extremely compact, A (or T) describes sufficient characteristics of the inputted oral care mesh (e.g., tooth mesh) to enable such a reconstruction of that oral care mesh (e.g., tooth mesh).
  • the latent space vector A (or latent capsule T) can be used as an additional input to predictive or generative models of this disclosure.
  • the latent space vector A (or latent capsule T) can be used as an additional input to at least one of an MLP, an encoder, a transformer, a regularized autoencoder, or a VAE of this disclosure.
  • the latent space vector A (or latent capsule T) can be used as an input to the GDL Setups model described in the present disclosure. Furthermore, the latent space vector A (or latent capsule T) can be used as an input to the RL Setups model described in the present disclosure.
  • the advantage of training a setups prediction neural network to take a latent space vector A (or latent capsule T) as an input is to provide information about the reconstruction characteristics of the tooth mesh to the network.
  • Reconstruction characteristics may contain information about local and/or global attributes of the mesh.
  • Reconstruction characteristics may include information about mesh structure. Information about shape may, in some instances, be included. An awareness of these reconstruction characteristics may better enable the trained setups prediction model to predict a final setup or intermediate staging, thereby providing the technical improvement of improved data precision.
  • a further advantage of using the latent space vector A is the vector’s size.
  • a neural network may encode an understanding of the input mesh and pose data more resource-efficiently if those data are presented in a compact form (such as a vector of 128 real values), as opposed to inputting the full mesh (which may contain thousands of mesh elements).
  • the latent representation of a mesh may provide a more favorable signal-to-noise ratio than the original form of that mesh or those meshes, thereby improving the capability of a subsequent ML model (such as a neural network or SVM) to form predictions, draw inferences, and/or otherwise generate outputs (such as transforms or meshes) based on the input mesh(es).
  • FIG. 2 shows how some of various setups prediction models can take as input either 1) tooth meshes or 2) latent space vectors (or latent capsules) which represent tooth meshes in reduced- dimensionality form.
  • This portion of the disclosure is directed to a reinforcement learning machine learning model to place an oral care mesh (e.g., such as placing one or more teeth to predict a final setup or to predict a series of intermediate stages that implement orthodontic treatment).
  • the soft actor critic (SAC) model may be applied to the creation of final setups.
  • the machine learning technique of reinforcement learning may train a decision-making agent through a system of rewarding desired behavior and/or punishing undesirable behavior.
  • the agent may be able to generate actions and learn through a process of interacting with its environment.
  • the RL Setups model is a novel adaptation of the Soft Actor Critic (SAC) reinforcement learning algorithm, which is trained to generate final setups and/or intermediate stages for orthodontic treatment.
  • the RL Setups model has at least 3 main components: the RL environment, the RL neural networks (the decision-making entities, of which there are 5 or more) and the replay buffer (which stores the accumulated results of past executions of the algorithm).
  • FIG. 6 describes five of the neural networks which may contribute to the training method of FIG. 3.
  • Other implementations of SAC may be trained on other kinds of 3D oral care representations (e.g., hardware or appliance components), to place those oral care representations relative to one or more other oral care representations.
  • a transform may be predicted for a tooth to place that tooth in a pose relative to at least one other tooth.
  • a transform may be predicted for a hardware element to place that hardware element relative to a tooth.
  • a transform may be predicted to place an appliance or an appliance component (such as a library component) relative to one or more teeth or one or more other appliance components.
  • the Soft Actor Critic (SAC) network architecture may contain 5 (or more) neural networks, a replay buffer function 300 and an environment 336.
  • SAC Neural Network Architecture The remainder of the specification in this section (“SAC Neural Network Architecture”) applies to some implementations.
  • SAC is a form of Q-Leaming.
  • Actor Network a policy network which may predict an action based on a state.
  • Critic Networks two q-value networks - a minimum value predicted by these network may be used to update, at least in part, the policy and value functions.
  • Value Networks two or more networks which may predict how valuable is a state - the target value network may be softly updated based on the main value network.
  • ADAM (alternatives are described in “Neural network architecture”)
  • Feedforward ReLu nonlinear function (alternatives are described in “Neural network architecture”)
  • Activation Function Tanh (alternatives are described in “Neural network architecture”)
  • Replay Buffer may save tuples containing at least one of state, action, reward, and next state.
  • Maximum Entropy may encourage exploration by randomly selecting actions.
  • Reparameterization may be used to update the Policy Actor network and may reduce problems with backpropagating errors.
  • Loss simple mean square error (MSELoss).
  • Network weights may be modified in small steps (e.g., ⁇ 3e-3).
  • Policy Network When training the policy actor network, a standard deviation with minimum of -20 and maximum of 2 is defined. Other implementations may follow other minimums and maximums, for example to avoid negative values and/or reduce the range of values (e.g., [-5,2] and [0.000001,1] respectively).
  • the environment may be considered to be the universe in which the agent operates (e.g., an orthodontic setup).
  • the environment may contain a module that applies movement (rotation, translation), may compute one or more rewards, may provide action samples, may define data dimensions and, may flag if/when the setup/goal may be achieved during training.
  • Each of the 5 neural networks may, in some implementations, be implemented using one or more MLPs consisting of linear layers and ReLu activation functions. Other possible neural networks and/or neural network components are disclosed elsewhere in this disclosure.
  • Agent The agent may be interpreted to be who/what performs/applies actions to a state and then who may receive the response (e.g., transformations to be applied to teeth, such as translations and/or rotations, after one or more actions are applied and one or more rewards are issued) from the environment.
  • the actor’s experience may be recoded in a loop of steps taken, where each step may correspond to an action which may be applied to one or more teeth.
  • a reward may be assigned after each action is applied.
  • An action sample may, in some implementations, be a 232-dimensional vector (other dimensions are possible).
  • An example action sample is: [upper arch form, upper arch teeth translation and rotation, lower arch-form, lower arch teeth translation and rotation]
  • a state may in some implementations be a 232-dimensional vector (or a vector of another size), and may have some of the same characteristics as an action, [upper arch form, upper arch teeth translation and rotation, lower arch-form, lower arch teeth translation and rotation]
  • Action applied to State Applying an action to a state may be interpreted to mean that each tooth is transformed, based, at least in part, on the current state of the tooth.
  • SAC may seek to maximize a policy’s entropy (e.g., which may involve a stochastic element to the moving of teeth in order to achieve a setup).
  • SAC may, in some implementations, be an off-policy algorithm.
  • SAC may in some implementations be used for intermediate staging, in which case motion limits may be advantageous (i.e., limits on rotation magnitude and/or translation magnitude).
  • FIG. 3 shows an example training method for a reinforcement learning model to predict transforms for 3D oral care representations (e.g., such as placing teeth for setups prediction).
  • a tuple 302 may be received by the agent 304.
  • a tuple 302 may comprise ⁇ a current state, an action, a next state, a reward, and a ‘done’ flag ⁇ , though other tuples are possible in other implementations.
  • the state may have one or more vectors which represent each tooth in each of the two arches.
  • a vector element may contain real numbers which may describe the translation and rotation transformations for one or more teeth (e.g., as translation vector and/or quaternion, or as transformation matrices).
  • the state may also, in some implementations, have an archform (e.g., as encoded by a set of control points or otherwise described elsewhere in this disclosure).
  • the Value Network 328 and the Target Value Network 308 comprise two of the five (or more) neural networks.
  • the 2 value networks may compute values for the current state 310 and next state (aka new state) 312 that may be used in updating at least one of the critic loss 316 or the value loss 320.
  • State data e.g., current state 310 and next state 312 from the tuples may be provided to these two neural networks, and the networks may provide the corresponding outputs (e.g., the state value and the target value tuple 318) to the calculation of losses: the value loss (used in updating the Value Network 328 and Target Value Network 308) and the critic loss (used in updating the 2 Critic Networks 334 and 306).
  • the purpose of the two critic networks may be to compute predicted Q-values 314 which may be used in updating the policy loss and the critic loss.
  • a tuple 338 comprising the current state and the predicted action (provided by the Actor Policy Network) may be provided to each of the two critic networks, and each critic network may generate a predicted Q-value 314. The minimum of these two predicted Q-values may be provided to the calculation of the critic loss 316.
  • a Q-value may map state and action information to one or more values which represent an expected longterm reward.
  • the RL techniques described herein may, in some implementations, learn an optimal policy which optimizes (e.g., maximizes) long-term cumulative rewards.
  • the purpose of the Actor Policy Network 330 may be to generate a predicted action 324.
  • the Actor Policy Network (APN) 330 may also generate a log probability 322 (i.e., the log probability of a state vector, after the state vector may have undergone a ReLU operation), which may be used in the calculation of the policy loss 326, which may, in turn, be used in training the Actor Policy Network 330.
  • a new tuple may be outputted to the environment 336.
  • This new tuple 332 may encapsulate the learning which took place during that round of training and may include any reward which accrued to the agent from this latest round of execution.
  • the new tuple may be received by the replay buffer 300, which may store one or more tuples and which may provide samples of tuples to the neural networks during future rounds of training.
  • the Actor Policy Network 330 comprises the neural network that may ultimately be trained by the training process and may get deployed for use in an operational RL Setups prediction model.
  • the APN may take the current state as input and may output an action (i.e., vector of transformations for the teeth to place the teeth into their setup poses). This action may be provided to the Environment, which may then proceed to generate a new state (i.e., which apply the transformations to the one or more teeth of the one or more aches). This process may iterate until a 'done' criterion is reached. In some implementations, a 'done' criterion may be reached when one or more metrics (as described elsewhere in this disclosure) show that the one or more arches reflect tooth positions which are acceptable for use in a setup (such as a final setup or intermediate stage). In some implementations, one or more of the iterations of the training procedure described in FIG.
  • motion limits may be imposed, so that a given tooth does not move too far (in terms of translation and/or rotation) for a given stage.
  • Each patient case may be represented by a vector (the size of which may vary depending on the contents of the vector).
  • the vector therefore may contain one or more arch forms, one or more teeth (for example 28 teeth - in the case that upper and lower 8s, namely, wisdom teeth, are excluded), translation information and/or rotation information.
  • One example of such a vector may be of 232 dimensions: [upper arch form, upper arch teeth translation and rotation, lower arch-form, lower arch teeth translation and rotation]
  • the environment may represent the ‘world’ of the task of transform generation (e.g., for orthodontic setups transform generation, or appliance component transform generation).
  • the environment may represent the arches of teeth in orthodontic treatment, where every predicted action may be applied to one or more tooth, which may result in the returning of one or more new states, rewards and/or done flags (i.e., to signify that the algorithm has reached completion).
  • FIG. 6 provides an illustration of various neural networks trained for use in the reinforcement learning methods of this disclosure (e.g., RL Setups). Some implementations may incorporate other neural networks in place of or in addition to the neural networks shown in FIG. 6, such as one or more of the neural networks listed elsewhere in this disclosure.
  • Actions The Environment may be interpreted as the place where actions are applied to the arch state (i.e., each tooth has a state, and translation and/or rotation may be applied to a tooth state).
  • An RL arch vector may comprise one or more (e.g., 28) tooth transforms, and a list of archform control points (or some other representation of the archform). There may be lower and upper arches. Actions may be applied to one or both arches.
  • An action may be a tooth movement (e.g., involving at least one of a translation and a rotation). As an action is applied, the environment may return at least three data points: an arch state, a reward, and a flag indicating whether the final state has been reached.
  • An action may comprise of a tooth transformation, which may have translation and/or rotation components.
  • Action samples may be used to train the neural network models.
  • An action sample may, for example, be a set of tooth movements between the states of a given patient case (e.g., embodied as a vector that represents translation and rotation for each tooth in each arch - using quaternions and translation vectors). Multiple teeth may move in correspondence to each action sample.
  • Action samples may be provided to the training environment to drive learning.
  • One or both of these values may also used in computing the critic loss.
  • the current state output value may be used to update the target value network.
  • the reward, in the form of a “done” flag, and the discount factor (gamma) may be used to transform the target value into the final target value.
  • the final target value may be used to compute one or both of the critic loss and the value loss.
  • the purpose of the policy loss may be to update the actor policy network.
  • the purpose of the actor policy network may be to predict an action.
  • One or more vectors that defines the current case state may be provided, and one or more action samples may be provided (e.g., to train the five or more neural networks on how to receive input in the form of a case state and move the teeth into setup poses).
  • the goal state may be defined, at least in part, by an example of a ground truth setup.
  • loss may be computed as a difference between a predicted action and a ground tmth action (e.g., a quantification of the difference in tooth transformations). Such a loss may, in some implementations, be used to train one or more neural networks in the RL Setups model.
  • the five (or more) neural network models may be trained via backpropagation.
  • Loss may, in some implementations, be computed as a function of rewards received. Loss may be a real number.
  • the reward may be computed by the environment as a mean squared error (MSE) between the generated state and the goal state (i.e., by way of a vector comparison). MSE may be computed on two vectors.
  • a state may contain the translation and/or rotation information for one or more teeth for one or both arches, alongside the archform information.
  • the rewards (and associated neural network loss values) may be used to train, at least in part, the five (or more) neural network models via backpropagation.
  • the replay buffer may store tuples having a ⁇ state, action, reward, flag to indicate whether goal was achieved, new predicted state ⁇ structure. Such tuples may comprise samples of what has been accomplished via past iterations of learning (training passes). Every time an action is applied, a tuple may be created and stored in the replay buffer. When the replay buffer is full (e.g., at 200 samples), a portion of the samples may be used to update one or more of the neural networks. The replay buffer may be used indirectly in loss calculation. The replay buffer may provide samples that may be used as inputs to the neural networks. The frequency of updates to the neural networks may be dependent on the size of the replay buffer. From time during training, the agent may request tuples from the replay buffer to drive learning.
  • Techniques of this disclosure may provide additional inputs to RL setups as well.
  • RL Setups such as in implementations that use the SAC algorithm
  • one or more of the input values from one or more of the following list of input vector may be introduced to the RL Setups model: B, K, L, M, N, O, R, S, P, Q, U, V.
  • one or more of the inputs B, K, L, M, N, O, R, S, P, Q, U, V may be provided to one or more of the neural networks of RL Setups: Actor Policy Network, Value Network, Target Value Network and both Critic Networks (i.e., in addition to the state information and/or predicted action information which various of these networks may also receive as input, according to various implementations).
  • one or more of the inputs B, K, L, M, N, O, R, S, P, Q, U, V may be received by the Actor Policy Network (APN), along with the input of the current state, with the advantage of better enabling the APN to generate actions (e.g., incremental tooth transforms) which produce effective setups predictions which meet the needs of patients (e.g., to generate an effective final setup or intermediate stage).
  • APN Actor Policy Network
  • Some implementations may adapt the RL Setups technique described above to perform operations on 3D oral care representations, such as tooth meshes, gums, hardware, appliances and appliance components (e.g., prefab library components). Such appliances and appliance components may pertain to dental restoration, and may be used to shape dental composites to produce veneers.
  • representation learning may be used in conjunction with reinforcement learning to generate transforms. Such transform may, in some implementations, be applied to tooth for setups prediction.
  • the APN may contain at least a first machine learning model and a second machine learning model (e.g., potentially in addition to other machine learning models).
  • the current state and/or the next state may maintain information about the tooth meshes.
  • a first ML model e.g., a U-Net, autoencoder, a pyramid encoder-decoder, a 3D encoder, or a series of convolution and/or pooling layers
  • a first ML model may be used to generate representations of the teeth.
  • a second ML model (e.g., an encoder, a transformer, an autoencoder, an MLP or a network comprising one or more global average pooling layers) may be trained to map those representations of the teeth (e.g., which may contain a representation vector corresponding to each mesh element) to one or more vectors of transformation values for one or more teeth (e.g., the APN may map the tooth representations into transforms which may place those teeth into setups poses).
  • Such implementations may generate tooth transforms which may be applied to corresponding tooth meshes.
  • collision detection may be performed on the tooth meshes through the course of training and deployment.
  • These techniques of the present disclosure may incorporate such collision detection between meshes into the calculation of loss values, which may be used to train, at least in part, the neural networks of the RL Setups technique, such as the APN.
  • the resulting APN may generate tooth transforms which seek to mitigate or reduce or potentially even avoid collisions between transformed tooth meshes.
  • Such an RL model may incorporate mesh collision mitigate or collision avoidance measures as described in US patent application with Publication No. US20220054234A1.
  • various RL-based techniques of this disclosure may implement one or more of a state-action-reward-state-action (S ARSA), which is a policy-based training method.
  • S ARSA state-action-reward-state-action
  • the S ARSA policy gives information about the probability that a certain tooth transformation or pose may be advantageous.
  • Examples of policy -free training methods include Q-learning (in which the agent for adjusting the poses of teeth explores the environment in a self-directed method) and/or Deep Q-learning (which is an implementation of Q-Leaming using neural networks).
  • RL-based approaches that are in accordance with aspects of this disclosure are described below.
  • the following RL algorithms may be trained to be applied in whole, in part or in modified form to the setups generation task.
  • Some implementations of these RL models for use in RL Setups generation may incorporate one or more of the neural networks described elsewhere in this disclosure.
  • Multiobjective RL for taking into account the refinement penalty, extended treatment duration penalty, and biological constraints
  • Safe RL utilizing collisions as a risk metric and using one of the methods that balances risk and reward to identify high reward staging that simultaneously limits collisions, thus increasing speed of staging.
  • DDPG Deterministic Policy Gradient
  • the RL techniques described herein may, in some implementations, use imitation learning.
  • an expert may provide input to help train the model in the form of sparse rewards (e.g., a manually specified reward function).
  • An agent may, in some implementations, attempt to learn an optimal policy by following the decisions of the expert.
  • An example of Imitation Learning that techniques of this disclosure may incorporate is described as follows. Generative approaches using Imitation learning may be used in the generation of intermediate staging.
  • GAIL Generative Adversarial Imitation Learning
  • GAIL is an effective method to improve the quality of setups (final and/or staging), reduce the expected number of required refinement stages (i.e., stages requiring rework after aligners have already been made); and decrease the count of stages required for treatment (leading to fewer aligners).
  • GAIL is an algorithm which may learn to “mimic” an expert’s behavior using one or more exemplary ground truth demonstrations (e.g., from an expert).
  • GAIL may harness the so-called “generative adversarial training” concept to fit a distribution over states and actions visited by the expert.
  • Ground truth setups (whether for final setups or for intermediate staging) and/or the generated (predicted) setups (final or intermediate staging) from a policy network (the generator network) may be provided to a discriminator network.
  • the discriminator may be implemented as a binary -classifier which classifies each input as 0 or 1 (i.e., whether the input is from a ground truth source or from an artificial intelligence-based algorithm).
  • the generator Upon successfully training both the generator and discriminator networks, the generator (after achieving convergence) will eventually be able how to generate setups (final or intermediate staging) that are indistinguishable from those provided as ground truth, even by a highly trained discriminator network.
  • Oral care arguments may include oral care parameters as disclosed herein, or other real-value, text-based or categorical inputs which specify intended aspects of the one or more 3D oral care representations which are to be generated.
  • oral care arguments may include oral care metrics, which may describe intended aspects of the one or more 3D oral care representations which are to be generated. Oral care arguments are specifically adapted to the implementations described herein.
  • the oral care arguments may specify the intended designs (e.g., including shape and/or structure) of 3D oral care representations which may be generated (or modified) according to techniques described herein.
  • implementations using the specific oral care arguments disclosed herein generate more accurate 3D oral care representations than implementations that do not use the specific oral care arguments.
  • a text encoder may encode a set of natural language instructions from the clinician (e.g., generate a text embedding).
  • a text string may comprise tokens.
  • An encoder for generating text embeddings may, in some implementations, apply either mean-pooling or max-pooling between the token vectors.
  • a transformer e.g., BERT or Siamese BERT
  • may be trained to extract embeddings of text for use in digital oral care e.g., by training the transformer on examples of clinical text, such as those given below).
  • such a model for generating text embeddings may be trained using transfer learning (e.g., initially trained on another corpus of text, and then receive further training on text related to digital oral care). Some text embeddings may encode text at the word level. Some text embeddings may encode text at the token level.
  • a transformer for generating a text embedding may, in some implementations, be trained, at least in part, with a loss calculation which compares predicted outputs to ground truth outputs (e.g., softmax loss, multiple negatives ranking loss, MSE margin loss, cross-entropy loss or the like).
  • the non-text arguments such as real values or categorical values, may be converted to text, and subsequently embedded using the techniques described herein.
  • a local coordinate system for a 3D oral care representation such as a tooth, may be described by one or more transforms (e.g., an affine transformation matrix, translation vector or quaternion).
  • Systems of this disclosure may be trained for coordinate system prediction using past cohort patient case data.
  • the past patient data may include at least: one or more tooth meshes or one or more ground truth tooth coordinate systems.
  • Machine learning models such as: U-Nets, encoders, autoencoders, pyramid encoder-decoders, transformers, or convolution and/or pooling layers, may be trained for coordinate system prediction.
  • Representation learning may determine a representation of a tooth (e.g., encoding a mesh or point cloud into a latent representation, for example, using a U-Net, encoder, transformer, convolution and/or pooling layers or the like), and then predict a transform for that representation (e.g., using a trained multilayer perceptron, transformer, encoder, transformer, or the like) that defines a local coordinate system for that representation (e.g., comprising one or more coordinate axes).
  • a representation of a tooth e.g., encoding a mesh or point cloud into a latent representation, for example, using a U-Net, encoder, transformer, convolution and/or pooling layers or the like
  • a transform for that representation e.g., using a trained multilayer perceptron, transformer, encoder, transformer, or the like
  • a local coordinate system for that representation e.g., comprising one or more coordinate axes.
  • the mesh convolutional techniques described herein can leverage invariance to rotations, translations, and/or scaling of that tooth mesh to generate predications that techniques that are not invariant to the rotations, translations, and/or scaling of that tooth mesh cannot generate.
  • Pose transfer techniques may be trained for coordinate system prediction, in the form of predicting a transform for a tooth.
  • Reinforcement learning techniques may be trained for coordinate system prediction, in the form of predicting a transform for a tooth.
  • Machine learning models such as: U-Nets, encoders, autoencoders, pyramid encoderdecoders, transformers, or convolution and/or pooling layers, may be trained as a part of a method for hardware (or appliance component) placement.
  • Representation learning may train a first module to determine an embedded representation of a 3D oral care representation (e.g., encoding a mesh or point cloud into a latent form using an autoencoder, or using a U-Net, encoder, transformer, block of convolution and/or pooling layers or the like). That representation may comprise a reduced dimensionality form and/or information-rich version of the inputted 3D oral care representation.
  • a representation may be aided by the calculation of a mesh element feature vector for one or more mesh elements (e.g., each mesh element).
  • a representation may be computed for a hardware element (or appliance component).
  • Such representations are suitable to be provided to a second module, which may perform a generative task, such as transform prediction (e.g., a transform to place a 3D oral care representation relative to another 3D oral care representation, such as to place a hardware element or appliance component relative to one or more teeth) or 3D point cloud generation.
  • transform prediction e.g., a transform to place a 3D oral care representation relative to another 3D oral care representation, such as to place a hardware element or appliance component relative to one or more teeth
  • 3D point cloud generation e.g., a transform to place a 3D oral care representation relative to another 3D oral care representation, such as to place a hardware element or appliance component relative to one or more teeth
  • Such a transform may comprise an affine transformation matrix, translation vector or quatern
  • Machine learning models which may be trained to predict a transform to place a hardware element (or appliance component) relative to elements of patient dentition include: MLP, transformer, encoder, or the like.
  • Systems of this disclosure may be trained for 3D oral care appliance placement using past cohort patient case data.
  • the past patient data may include at least: one or more ground truth transforms and one or more 3D oral care representations (such as tooth meshes, or other elements of patient dentition).
  • the mesh convolution and/or mesh pooling techniques described herein leverage invariance to rotations, translations, and/or scaling of that tooth mesh to generate predications that techniques that are not invariant to the rotations, translations, and/or scaling of that tooth mesh cannot generate.
  • Reinforcement learning techniques may be trained for hardware or appliance component placement.
  • Techniques of this disclosure may, in some implementations, use PointNet, PointNet++, or derivative neural networks (e.g., networks trained via transfer learning using either PointNet or PointNet++ as a basis for training) to extract local or global neural network features from a 3D point cloud or other 3D representation (e.g., a 3D point cloud describing aspects of the patient’s dentition - such as teeth or gums).
  • Techniques of this disclosure may, in some implementations, use U-Nets to extract local or global neural network features from a 3D point cloud or other 3D representation.
  • 3D oral care representations are described herein as such because 3-dimensional representations are currently state of the art.
  • 3D oral care representations are intended to be used in a non-limiting fashion to encompass any representations of 3 -dimensions or higher orders of dimensionality (e.g., 4D, 5D, etc.), and it should be appreciated that machine learning models can be trained using the techniques disclosed herein to operate on representations of higher orders of dimensionality.
  • input data may comprise 3D mesh data, 3D point cloud data, 3D surface data, 3D polyline data, 3D voxel data, or data pertaining to a spline (e.g., control points).
  • An encoderdecoder structure may comprise one or more encoders, or one or more decoders.
  • the encoder may take as input mesh element feature vectors for one or more of the inputted mesh elements. By processing mesh element feature vectors, the encoder is trained in a manner to generate more accurate representations of the input data.
  • the mesh element feature vectors may provide the encoder with more information about the shape and/or structure of the mesh, and therefore the additional information provided allows the encoder to make better-informed decisions and/or generate more-accurate latent representations of the mesh.
  • encoder-decoder structures include U-Nets, autoencoders or transformers (among others).
  • a representation generation module may comprise one or more encoder-decoder structures (or portions of encoders-decoder structures - such as individual encoders or individual decoders).
  • a representation generation module may generate an information-rich (optionally reduced-dimensionality) representation of the input data, which may be more easily consumed by other generative or discriminative machine learning models.
  • a U-Net may comprise an encoder, followed by a decoder.
  • the architecture of a U-Net may resemble a U shape.
  • the encoder may extract one or more global neural network features from the input 3D representation, zero or more intermediate-level neural network features, or one or more local neural network features (at the most local level as contrasted with the most global level).
  • the output from each level of the encoder may be passed along to the input of corresponding levels of a decoder (e.g., by way of skip connections).
  • the decoder may operate on multiple levels of global-to-local neural network features. For instance, the decoder may output a representation of the input data which may contain global, intermediate or local information about the input data.
  • the U-Net may, in some implementations, generate an information-rich (optionally reduced-dimensionality) representation of the input data, which may be more easily consumed by other generative or discriminative machine learning models.
  • An autoencoder may be configured to encode the input data into a latent form.
  • An autoencoder may train an encoder to reformat the input data into a reduced-dimensionality latent form in between the encoder and the decoder, and then train a decoder to reconstruct the input data from that latent form of the data.
  • a reconstruction error may be computed to quantify the extent to which the reconstructed form of the data differs from the input data.
  • the latent form may, in some implementations, be used as an information-rich reduced-dimensionality representation of the input data which may be more easily consumed by other generative or discriminative machine learning models.
  • an autoencoder may be trained to input a 3D representation, encode that 3D representation into a latent form (e.g., a latent embedding), and then reconstruct a close facsimile of that input 3D representation at the output.
  • a latent form e.g., a latent embedding
  • a transformer may be trained to use self-attention to generate, at least in part, representations of its input.
  • a transformer may encode long-range dependencies (e.g., encode relationships between a large number of inputs).
  • a transformer may comprise an encoder or a decoder. Such an encoder may, in some implementations, operate in a bi-directional fashion or may operate a self-attention mechanism. Such a decoder may: operate a masked self-attention mechanism; operate a cross-attention mechanism; or operate in an auto-regressive manner, according to particular implementations.
  • the self-attention operations of the transformers described herein may, in some implementations, relate different positions or aspects of an individual 3D oral care representation in order to compute a reduced-dimensionality representation of that 3D oral care representation.
  • the cross-attention operations of the transformers described herein may, in some implementations, mix or combine aspects of two (or more) different 3D oral care representations.
  • the auto-regressive operations of the transformers described herein may, in some implementations, consume previously generated aspects of 3D oral care representations (e.g., previously generated points, point clouds, transforms, etc.) as additional input when generating a new or modified 3D oral care representation.
  • the transformer may, in some implementations, generate a latent form of the input data, which may be used as an information-rich reduced-dimensionality representation of the input data, which may be more easily consumed by other generative or discriminative machine learning models.
  • an encoder-decoder structure may first be trained as an autoencoder. In deployment, one or more modifications may be made to the latent form of the input data. This modified latent form may then proceed to be reconstructed by the decoder, yielding a reconstructed form of the input data which differs from the input data in one or more intended aspects. Oral care arguments, such as oral care parameters or oral care metrics may be provided to the encoder, the decoder, or may be used in the modification of the latent form, to guide the encoder-decoder structure in generating a reconstructed form that has desired characteristics (e.g., characteristics which may differ from that of the input data).
  • Federated learning may enable multiple remote clinicians to iteratively improve a machine learning model (e.g., validation of 3D oral care representations, mesh segmentation, mesh cleanup, other techniques which involve labeling mesh elements, coordinate system prediction, non-organic object placement on teeth, appliance component generation, tooth restoration design generation, techniques for placing 3D oral care representations, setups prediction, generation or modification of 3D oral care representations using autoencoders, generation or modification of 3D oral care representations using transformers, generation or modification of 3D oral care representations using diffusion models, 3D oral care representation classification, imputation of missing values), while protecting data privacy (e.g., the clinical data may not need to be sent “over the wire” to a third party).
  • a machine learning model e.g., validation of 3D oral care representations, mesh segmentation, mesh cleanup, other techniques which involve labeling mesh elements, coordinate system prediction, non-organic object placement on teeth, appliance component generation, tooth restoration design generation, techniques for placing 3D oral care representations, setups prediction, generation or modification of
  • a clinician may receive a copy of a machine learning model, use a local machine learning program to further train that ML model using locally available data from the local clinic, and then send the updated ML model back to the central hub or third party.
  • the central hub or third party may integrate the updated ML models from multiple clinicians into a single updated ML model which benefits from the learnings of recently collected patient data at the various clinical sites. In this way, a new ML model may be trained which benefits from additional and updated patient data (possibly from multiple clinical sites), while those patient data are never actually sent to the 3rd party.
  • Training on a local in-clinic device may, in some instances, be performed when the device is idle or otherwise be performed during off-hours (e.g., when patients are not being treated in the clinic).
  • Devices in the clinical environment for the collection of data and/or the training of ML models for techniques described herein may include intra-oral scanners, CT scanners, X- ray machines, laptop computers, servers, desktop computers or handheld devices (such as smart phones with image collection capability).
  • contrastive learning may be used to train, at least in part, the ML models described herein. Contrastive learning may, in some instances, augment samples in a training dataset to accentuate the differences in samples from difference classes and/or increase the similarity of samples of the same class.
  • a reward model (RM) for setups prediction also referred to as a preference model
  • the preference model may be trained to output the term r omega, which may be a scalar value which may describe at least one aspect of preferability of a setup.
  • a policy may be a setups prediction model (or alternatively, a model that predicts a transform for another type of oral care representation or a model which generates an oral care representation) which receives input pertaining to the ranking or quality of one or more setups from a clinician or technician (e.g., received by way of a prompt).
  • the policy may output one or more transforms for one or more teeth, or in some implementations, may output transforms for other types of oral care representation or may output 3D oral care representations themselves.
  • the action space of this policy may pertain to the space of possible transforms and the observation space may pertain to the distribution of possible transforms, which may be large.
  • the reward function may pertain to a combination of the preference model and a constraint on policy shift.
  • FIG. 8 shows an example training method.
  • Patient case data 802 including 3D representations of the patient’s teeth, malocclusion tooth transforms, and/or ground truth setup tooth transforms, may be provided to the model Fl 804 and model F2 806.
  • a setups prediction model Fl 804 may be trained on a dataset of cohort patient cases in a supervised/semi-supervised fashion (e.g., GDL Setups, Transformer Setups, Autoencoder Setups, Pose Transfer Setups and the like).
  • Such models may include ML components such as: U-Nets, encoders, decoders, pyramid encoder-decoders, transformers, autoencoders, multi-layer perceptrons (MLP), or the like.
  • MLP multi-layer perceptrons
  • a U-Net may be used to extract hierarchical neural network features from the tooth meshes of the patient’s arch and generate latent representations as embedding vectors.
  • the embedding vectors may be provided to an MLP (e.g., four fully connected layers with optional skip connections), which may generate the setups transforms for the teeth.
  • the performance of setup prediction model Fl 804 may be improved upon using reinforcement learning from human feedback (RLHF).
  • RLHF may entail the providing of feedback from a human to train a reward model. Stated another way, RLHF may incorporate feedback from a human user into the data that are used to train a reward model.
  • the trained reward model may be used as a reward function to optimize the policy of the agent, using reinforcement learning.
  • Setup prediction model Fl 804 may be trained for tooth placement in the setups application or may also be trained more broadly to place 3D oral care representations other than teeth.
  • model Fl may place a 3D oral care representation relative to at least one other 3D oral care representation or relative to at least one global coordinate axis (e.g., to place a fixture model component or an appliance component relative to one or more teeth).
  • a copy may be made of model Fl, referred to as setups prediction model F2 806.
  • the training of Fl may be frozen to prevent further learning by setups prediction model Fl 804, whereas the further learning may be applied to setups prediction model F2 806 through RL optimizations (e.g., referred to herein as 'fine-tuning').
  • a software application such as implemented using a web interface on a PC or mobile device, may enable feedback to be collected from a clinician or technician.
  • the following process may undergo multiple iterations in the course of training model F2 through RLHF.
  • a maloccluded arch of teeth 802 from the training dataset may be provided to model Fl 804, yielding a predicted setup Y1 812 as output.
  • the same maloccluded arch of teeth may be provided to model F2 806, yielding a predicted setup Y2 814 as output.
  • Model Fl 804 represents the original setups model, while model F2 806 may undergo “fine-tuning” or improvement with each iteration.
  • F pre ferencemodei 816 is defined which may enable grading/ranking/scoring of the predicted setup Y2.
  • F pre ferencemodei 816 may in some instances have already undergone training. 816 may in some instances have undergone human-in-the-loop training or otherwise been trained (at least in part) on at least one aspect of a dataset of human-in-the-loop responses.
  • Predicted setup Y2 may be provided to F pre f ere ncemodei 816, which may then output a notion of “preferability” (e.g., as a scalar or categorical value), called r omega.
  • r omega may, in some instances, describe the fitness, suitability, or quality of the predicted setup predicted setup Y2, as assigned by a clinician or technician (e.g., using the software interface described above, or via other channels).
  • Predicted setup Y1 and predicted setup Y2 may then be compared using a technique such a scaled KL-Divergence loss or another of the loss calculation methods described herein.
  • the output of this loss calculation is denoted by r_KL 818.
  • the KL-divergence loss may penalize the RL policy from moving substantially away from the initial pretrained model with each iteration of the RLHF training. This penalty may benefit training, by biasing the model to output setups which are coherent, valid setups, at least within an acceptable margin of error.
  • Model F2 may, in some instances, continue to be trained until the performance of model F2 reaches a threshold, the performance of model F2 supersedes model Fl in some metric and/or until no further learning can otherwise be achieved.
  • Imitation Learning (which is one example of Reinforcement Learning) may be used to train an ML model to place a 3D oral care representation (e.g., a tooth for a clear tray aligner, an appliance component for a dental restoration appliance or a component for a fixture model).
  • a neural network may be trained using Imitation Learning to generate transforms to place 3D oral care representations (e.g., teeth for restoration, pontics, appliance components or fixture model components) into poses which are suitable for oral care appliance generation.
  • the training may be performed using cohort patient case data.
  • an RL model may use a reward function to learn, at least in part, a policy.
  • the trained policy may be used by a deployed RL model to make decisions.
  • anRL model may use reference or ground truth data (e.g., a reference or ground truth transformation matrix for each tooth) to learn, at least in part, a policy.
  • a policy may represent actions for the agent to taken in a given state.
  • An example of an “action” is a transform that may be applied to a tooth mesh, an appliance, an appliance component, or other 3D representation.
  • some implementations may be trained on reference (or ground truth) patient data which includes state-action pairs, such as [current state tooth transforms, action transforms]. For example, such a pair may take the form of: [[4x4 tooth transform], [4x4 tooth transform]]).
  • An Imitation Learning model may use supervised learning to learn a policy, which may minimize loss.
  • the supervised learning model may comprise one or more neural networks (e.g., an MLP comprising an input layer, one or more hidden layers and an output layer).
  • the training data may include tooth movement data (e.g., quaternions, or 4x4 transformation matrices).
  • the training data may include latent representations of the patient’s dentition, such as latent representations of the patient’s teeth generated using an autoencoder, a transformer or the other representation generation methods described herein.
  • transforms or latent representations may be provided to the training module (e.g., a neural network) which generates tooth transforms for setups predictions.
  • Some implementations may train a module which combine reference behaviors (actions, such as reference tooth movements) with a reinforcement learning reward function to train an algorithm that learns to replicate the optimal rewards of reference behaviors (state-actions pairs).
  • Some Reinforcement Learning models may train a Multi-armed Bandit model to generate transforms to place 3D representation of teeth, appliance components, fixture model components (among other items) into poses which are suitable for appliance generation.
  • Multiple candidate actions e.g., data pertaining to tooth movements - such as tooth transforms for use in setups
  • the neural network module of the Multi-armed Bandit model may be trained to ‘select’ one of the provided candidate actions, an action which may optimize a reward.
  • the reward may quantify the results of an action taken by the agent. A higher reward may correspond to an action that more-highly benefits the agent.
  • a Multi-armed Bandit model may comprise at least: one or more reinforcement learning agents, one or more environments, or one or more reward functions.
  • a Multi-armed Bandit model may then compute a reward (e.g., mean square error between the expected tooth state and the actual tooth state) based on the chosen action (e.g., a predicted tooth transform) and train (e.g., via penalizing or rewarding) the neural network module through a loss function (e.g., vertex to vertex Euclidean distance between expect position and actual tooth position after applied the action).
  • a reward e.g., mean square error between the expected tooth state and the actual tooth state
  • train e.g., via penalizing or rewarding
  • the neural network module e.g., via penalizing or rewarding
  • a loss function e.g., vertex to vertex Euclidean distance between expect position and actual tooth position after applied the action.
  • Such implementations may be trained on both arches and all teeth in a patient’s mouth, or the implementation may be trained on the social six (e.g., top front six teeth), or be trained on some other specific set of teeth or malocclusion. Examples of malocclusion include open bite or deep bite cases.
  • the training dataset may contain state-action pair vectors (e.g., [[4x4 tooth transform], [4x4 tooth transform]]), each of which may contain information about the pose of the 3D representation which is to be placed (e.g., such as tooth position, rotation or translation).
  • the rewards may be computed only based on the actions, and independently of the predicted next state of the teeth.
  • the rewards may take into consideration information about neighboring teeth, collisions between teeth or the distance between a tooth’s current position and the final setup position of that tooth. For example, such rewards may be computed position adjacent teeth in a manner so as to minimize collisions.
  • rewards may be subject to constraints.
  • constraints may include information pertaining to tooth movement limits (e.g., max clinical rotation allowed, or max translation allowed), space between teeth (e.g., max or min), collisions between adjacent teeth (e.g., maximum allowed collisions between the mesh elements of adjacent teeth), or root movement (e.g., max or min).
  • reinforcement learning may be used to train an ML model to label mesh elements (e.g., points, edges/faces/vertices, voxels), for example, for mesh segmentation or mesh cleanup.
  • a segmentation parameters module may be trained using RL techniques to generate segmentation parameters. Such parameters may identify portions of a mesh which may be subject to more fine-grained segmentation than other portions of the mesh. For example, a portion of the mesh with a higher density (e.g., point cloud density, number of points, point distance, etc.) than other portions of the mesh may receive a finer-grained segmentation than those other portions of the mesh.
  • a 3D representation e.g., 3D point cloud or 3D mesh, etc.
  • segmentation parameters module may be provided to the segmentation parameters module.
  • the generated parameters may be provided to a subsequent ML model for segmentation (e.g., a U-Net, pyramid encoder-decoder, 3D SWIN, MASK- CNN or CNN) that then generates mesh element labels for the 3D representation based at least in part on the segmentation parameters.
  • a subsequent ML model for segmentation e.g., a U-Net, pyramid encoder-decoder, 3D SWIN, MASK- CNN or CNN
  • mesh element labels e.g., labels on faces, vertices, edges, voxels or points.
  • the newly segmented meshes may be used to compute an RL reward.
  • Examples of the reward include Euclidean distance between the labelled mesh elements of a newly segmented mesh and a ground truth mesh (e.g., pairwise Euclidean distances between the mesh of the same mesh element label between the two meshes).
  • the newly segmented meshes e.g., the mesh element labels
  • a representation generation ML model e.g., an encoder from an autoencoder, an encoder from a transformer, or a stand-alone encoder
  • the latent representation may comprise a reduced-dimensionality form of the 3D representation.
  • the latent representation may describe aspects of dental anatomy (e.g., fossae, tips, ridges, and the like). Denoising, downsampling or filtering operations may be applied to improve the fidelity of the 3D representation (e.g., prior to segmentation). Data augmentation of the training dataset may enable better performance and reduce overfitting of the segmentation ML model.
  • the mesh element labeling techniques of this disclosure may be performed for the purpose of mesh cleanup (e.g., to label mesh elements corresponding to hardware that is to be removed from a tooth crown or extraneous material that is to be removed from a scanned dental arch).
  • an automated setups prediction model may be trained to generate a setup with a customized curve-of-spee (e.g., a curve-of-spee which conforms to the intended outcome of the treatment of the patient).
  • a customized curve-of-spee e.g., a curve-of-spee which conforms to the intended outcome of the treatment of the patient.
  • Such a model may be trained on cohort patient case data.
  • One or more oral care metrics may be computed on each case to quantify or measure aspects of that case's curve-of- spee.
  • one or more of such metrics may be provided to the setups prediction model, for example, to influence the model regarding the geometry and/or structure of each case's curve-of- spee.
  • That same input pathway to the trained neural network may be configured with one or more values as instructions to the model about an intended curve- of-spee. Such values may automatically generate a setup with a curve-of-spee which meets the aesthetic and/or medical treatment needs of the particular patient case.
  • a curve-of-spee metric may measure the curvature of the occlusal or incisal surfaces of the teeth on either the left or right sides of the arch, with respect to the occlusal plane.
  • the occlusal plane may, in some instances, be computed as a surface which averages the incisal or occlusal surfaces of the teeth (for one or both arches).
  • a curvature metric may be computed along a normal vector, such as a vector which is normal to the occlusal plane.
  • a curvature metric may be computed along the normal vector of another plane.
  • an XY plane may be defined to correspond to the occlusal plane.
  • An orthogonal plane may be defined as the plane that is orthogonal to the occlusal plane, which also passes through a curve- of-spee line segment, where the curve-of-spee line segment is defined by a first endpoint which is a landmarking point on a first tooth (e.g., canine) and a second endpoint which is a landmarking point on the most-posterior tooth of the same side of the arch.
  • a landmarking point can in some implementations be located along the incisal edge of a tooth or on the cusp of a tooth.
  • the landmarking points for the intermediate teeth may form a curved path, such as may be described by a polyline.
  • the following is a non-limiting list of curve-of-spee oral care metrics.
  • the line segment is defined by joining the highest cusp of the most-posterior tooth (in the lower arch) and the cusp of the first tooth on that side (in the lower arch). Given the subset of teeth between the first tooth and the most-posterior tooth, the point is defined by the highest cusp of the lowest tooth of this subset.
  • a curve-of-spee metric may be computed using the following 4 steps, i) Line: Form a line between the highest cusp on the most posterior tooth and the cusp of the first tooth, ii) Curve Point A: Given the set of teeth between the most posterior tooth and the first tooth, find the highest point of the lowest tooth, iii) Curve Point B: Project Curve Point A onto the Line to find a point (Curve Point B) along the line that is closest to Curve Point A. iv) Curve-Of-Spee: Find the height difference between Curve Point B and Curve Point A.
  • [00210] 2 Project one or more intermediate landmark points (e.g., points on the teeth which lie between the first tooth and the most-posterior tooth on that side of the arch) and the curve-of-spee line segment onto the orthogonal plane. Compute the curve-of-spee metric by measuring the distance between the farthest of the projected intermediate points to the projected curve-of-spee line segment. This yields a measure for the curvature of the arch relative to the orthogonal plane.
  • intermediate landmark points e.g., points on the teeth which lie between the first tooth and the most-posterior tooth on that side of the arch
  • Curve of Spee by measuring the distance between the farthest of the intermediate points to the curve-of- spee line segment. This yields a measure for the curvature of the arch in 3D space.
  • Curve-of-spee metrics 5 and 6 may help the network to reduce some more degrees of freedom in defining how the patient’s arch is curved in the posterior of the mouth.
  • the RL techniques of this disclosure may generate transforms to place appliance components (e.g., as described herein) or fixture model components (e.g., as described herein) into poses which are suitable for oral care appliance generation (e.g., dental restoration appliances or orthodontic appliances - such as CTA).
  • appliance components e.g., as described herein
  • fixture model components e.g., as described herein
  • the RL models for transform generation of this disclosure may be trained to generate transforms for one or more oral care appliance components (e.g., to generate transforms to place the one or more appliance components relative to one or more teeth of the patient).
  • Losses e.g., LI, L2, or reconstruction loss, among others described herein
  • Such losses may be used to train, at least in part, the second ML module.
  • pre-defined (or library) appliance components which may be placed using techniques of this disclosure include: vents, rear snap clamps, door hinges, door snaps, an incisal registration feature, center clips, custom labels, a manufacturing case frame, a diastema matrix handle, among others.
  • a digital fixture model may comprise 3D representations of the patient’s dentition, with optional fixture model components attached to that dentition or placed in relation to the dentition.
  • the RL models for transform generation described herein may be trained to generate transforms to place digital fixture model components into poses that are suitable for oral care appliance generation.
  • Fixture model components may include 3D representations (e.g., 3D point clouds, 3D meshes, or voxelized representations) of one or more of the following non-limiting items: 1) interproximal webbing - which may fill-in space or smooth-out the gaps between teeth to ensure aligner removability.
  • interproximal reinforcement - a structure on the exterior of an oral care appliance (e.g., an aligner tray), which may extend from a first gingival edge of the appliance body on a labial side of the appliance body along an interproximal region between the first tooth and the second tooth to a second gingival edge of the appliance body on a lingual side of the appliance body.
  • the effect of the interproximal reinforcement on the appliance body at the interproximal region may be stiffer than a labial face and a lingual face of the first shell. This may allow the aligner to grasp the teeth on either side of the reinforcement more firmly.
  • gingival ridge structure - a structure which may extend along the gingival edge of a tooth in the mesial-distal direction for the purpose of enhancing engagement between the aligner and a given tooth.
  • torque points - structures which may enhance force delivered to a given tooth at specified locations.
  • power ridges - structures which may enhance force delivered to a given tooth at a specified location.
  • dimples - structures which may enhance force delivered to a given tooth at specified locations.
  • a physical pontic is a tooth pocket that does not cover a tooth when the aligner is installed on the teeth.
  • the tooth pocket may be filled with tooth-colored wax, silicone, or composite to provide a more aesthetic appearance.
  • power bars - blockout added in an edentulous space to provide strength and support to the tray.
  • a power bar may fill- in voids.
  • Abutments or healing caps may be blocked-out with a power bar.
  • the trimline may define the path along which a clear aligner may be cut or separated from a physical fixture model, after 3D printing. 13) undercut fill - material which is added to the fixture model to avoid the formation of cavities between the fixture model’s height of contour and another boundary (e.g., the gingiva or the plane that the plane that undergirds the physical fixture model after 3D printing).
  • a digital pontic tooth may be placed using the RL techniques of this disclosure.
  • a pontic tooth is a digital 3D representation of a tooth which may act as a placeholder in an arch.
  • a pontic may function as a tooth pocket that may be filled with a tooth-colored material (e.g., wax) to improve aesthetics.
  • a pontic tooth may hold a space in the arch open during orthodontic setups generation. When automated setups prediction is performed, transforms for intermediate stages may be generated.
  • One or more pontic teeth may be defined to hold open space for missing teeth or act as a placeholder as space closes or opens in the arch during setups predicted staging or an unerupted tooth may erupt into a space where a pontic is present. Creating a pocket for the tooth to erupt into.
  • Digital pontic teeth may be placed in (or generated within) an arch (e.g., during setups generation or fixture model generation) to reserve space in an arch for missing or extracted teeth (e.g., so that adjacent teeth do not encroach upon that space over the course of successive intermediate stages of orthodontic treatment).
  • digital pontic tooth may be used when the space (e.g., the space that is to be held open) is at least a threshold dimension (e.g., a width of 4mm, among others).
  • digital pontic teeth for UL4-UR4 or LL4-LR4 may be placed (or generated or modified) when space is available or when there is a partially erupted tooth within the space.
  • a digital pontic tooth When there is a partially erupted tooth in the space, a digital pontic tooth may be placed over the empting tooth to maintain a space for the erupting tooth to erupt into. Pontic teeth may be placed to be inside the gingiva, to minimize (or avoid) heavy occlusal contacts (e.g., contacts between the chewing surfaces of the upper or lower arches), or to cover an erupting tooth (when present), among other conditions.
  • An RL-based ML model may be trained to generate (or modify) oral care data describing a current state of a treatment.
  • examples of such data include 3D representations of oral care data (e.g., teeth, appliance components, fixture model components, etc.), or transforms which may be applied to those 3D representations of oral care data.
  • a pre-restoration tooth mesh (or point cloud) may be provided to an RL-based ML model for 3D representation modification.
  • the RL-based ML model may modify the pre-restoration tooth mesh, according to one or more oral care arguments (e.g., restoration design metrics) which may influence the RL-base ML model to modify the shape and/or structure of the pre-restoration tooth mesh in a manner so that the generated post-restoration tooth mesh is suitable for use in treatment of the patient.
  • oral care arguments e.g., restoration design metrics
  • Other 3D representations of oral care data described herein may also be modified using RL: appliance components (e.g., parting surfaces, etc.), fixture model components (e.g., digital pontic teeth, or interproximal webbing, etc.), archforms, among other examples described herein.
  • RL techniques may be used to train ML models described herein (e.g., techniques based on encoder-decoder structures, such as autoencoders or transformers) for the generation or modification of 3D representations (e.g., 3D point clouds, voxels, etc.).
  • An RL method of training an ML model to generate a 3D representation may involve elements such as: state, actions, rewards, and policy.
  • the state of a 3D representation may be described by the structures, positions, and/or mesh element feature vectors of a set of mesh elements.
  • a mesh element feature vector may be computed, and be provided to the ML models described herein.
  • the state of a 3D representation e.g., point cloud
  • the state of a 3D mesh may comprise the XYZ coordinates of the vertices, and/or the edge or face lists which describe the structure of the connections between the vertices.
  • positional and/or structural state data may be augmented by one or more mesh element features (e.g., vertex normals, face normals, edge dihedral angles, vertex curvatures, edge curvatures, vertex concavity /convexity measurements, distance to nearest opposing surface), as described herein.
  • the state of a 3D representation may comprise, at least in part, the set of mesh element feature vectors associated with the 3D representation’s mesh elements.
  • Each mesh element feature vector may be modified by actions, explicitly or implicitly, and/or may be used to inform actions and/or rewards.
  • An agent i.e., a trained MLP or other of the neural networks described herein may generate one or more actions which modify the state of the 3D representation which is being generated.
  • the actions may be operable to modify the configuration of mesh elements (e.g., points of a point cloud, etc.) into a shape and/or structure which is suitable for use in generating an oral care appliance.
  • One or more oral care arguments e.g., oral care procedure parameters or oral care metrics
  • Actions may include modifying the positions and/or orientations of mesh elements.
  • Nonlimiting examples of actions include:
  • a morphing action may comprise modifying the shape and/or structure of an initial 3D representation to more nearly reflect the shape and/or structure of a target 3D representation, for example, over a series of evolving steps.
  • Actions samples may be provided as training data. Actions could also be computed based on the mesh-morphing technique. The same technique can be used to compute rewards when using 3D representations.
  • the RL environment may comprise one or more constraints (e.g., on motion, rotation, translation, width, length, etc.), a function that applies actions or a reward function.
  • the reward may quantify the magnitude of success of a new state (e.g., after application of the predicted action), for example, a modified shape and/or structure of the 3D representation (e.g., in shaping one or more teeth for use in restoring the patients’ dentition, in shaping a fixture model component, etc.).
  • the reward may impact loss computation (e.g., comparison between predicted action and the ground truth action, such using reconstruction loss or mean square error of a list of KPIs (e.g., mesh elements, positions, etc.) or another of the losses described herein).
  • the reward may, in some implementations, involve the calculation of oral care metrics (e.g., “Bilateral Symmetry & Ratios,” “Proximal Contacts,” or “Tooth Morphology” when a point cloud for a tooth restoration design is generated).
  • oral care metrics e.g., “Bilateral Symmetry & Ratios,” “Proximal Contacts,” or “Tooth Morphology” when a point cloud for a tooth restoration design is generated.
  • Such metrics may contribute to assessing the progress of the 3D representation towards taking-on a shape and/or structure that is indicated by the one or more oral care arguments.
  • the metrics may measure how well the point cloud describes a tooth restoration design, a fixture model component, a generated appliance component, or another 3D oral care representation that is suitable for use in appliance generation.
  • the generated (or modified) point cloud may then be used as part of the generation of an oral care appliance.
  • the dataset may comprise 3D oral care representations that are divided into train, test and validation sets (e.g., 75%, 15%, 10%).
  • An RL model may be trained to generate (or modify) 3D representations, in some implementations, the RL model may function alongside another ML model (e.g., a transformer or an autoencoder) to generate (or modify) a 3D representation.
  • another ML model e.g., a transformer or an autoencoder
  • an RL model e.g., an MLP with optional skip connections
  • the encoder-decoder structure may generate a 3D representation of the patient’s dentition (e.g., one or more tooth crowns), which may be provided to the RL model.
  • the RL model may generate one or more actions to be taken to improve the provided 3D representation of patient's dentition.
  • the agent may comprise one or more neural networks (e.g., an encoder-decoder structure such as an autoencoder or transformer, or an MLP) which are trained with multiple sample actions (e.g., from past data from cohort patient cases) so the agent leams how to generate actions.
  • the actions generated by the RL model may be applied to a partially generated or formed 3D representation (e.g., a 3D representation of the patient’s dentition that is under modification - such as to incorporate fixture model components into the dentition as a part of fixture model generation), resulting in a modified 3D representation of the patient’s dentition with evolving fixture model components attached.
  • Ground truth (or reference) cohort patient case data may be provided to the training workflow.
  • One or more loss values may be computed by comparing the generated (or modified) 3D representation to the corresponding ground truth data 3D representation.
  • the one or more loss values (e.g., reconstruction loss, chamfer loss or other losses described herein) may be used to train, at least in part, the one or more neural networks which comprise the agent of the RL model.
  • a reward may, in some implementations, be computed by comparing a modified 3D representation (e.g., new state or next state) to a corresponding ground truth 3D representation.
  • the reward function may influence the agent to execute actions which change the state of the 3D representation, and to change the state of the 3D representation in a manner which causes the 3D representation to assume a shape and/or structure which is more nearly suitable for use in generating an oral care appliance.
  • This computed reward may be used to optimize a agent or policy neural network, enabling that the agent or policy neural network to predict or generate better actions.
  • loss values may be computed based on a comparison between predicted actions and corresponding ground truth actions.
  • Rewards may be computed based upon a comparison between predicted next states and corresponding ground truth next states. Both losses and rewards may, in some implementations, be computed upon vectors of real values. In loss calculation, those vectors describe actions. In reward calculation, those vectors describe states. While losses and rewards both reflect differences, losses reflect differences in actions, and rewards reflect differences in states.
  • an action may, in some implementations, comprise a transform that describes a transition of a tooth from a current pose to a next pose.
  • a current state may, in some implementations, comprise a transform that describes a current pose of a tooth, and a next state may comprise a different transform that describes a next pose of a tooth.
  • an action may, in some implementations, comprise an adjustment to the concavity of a surface (among others described herein).
  • a current state of a 3D representation may, in some implementations, comprise the set of mesh elements (and optionally the associated mesh element feature vectors).
  • a next current state of a 3D representation may, in some implementations, comprise the set of mesh elements (and optionally the associated mesh element feature vectors) after the application of at least one action.
  • the next state may comprise the set of mesh elements, where at least one mesh element (e.g., the position of a vertex, among others) has been modified by the application of an action.
  • the loss calculation methods described herein may compare two or more vectors of real values, such as MSE, LI, L2, reconstruction loss, among others. Rewards and losses may, in some examples, be computed based upon these method of comparing vectors of real values. For example, in some implementations, loss may be computed (at least in part) using reconstruction loss, and the reward may be computed using MSE. [00236] In some implementations, a loss may be computed, at least in part, based upon a reward.
  • a reward value may contribute to the calculation of a loss value, in some implementations.
  • RL reinforcement learning
  • an agent neural network goal is to learn a policy.
  • a policy maps a state to an action.
  • a machine learning model is trained to learn actions from states.
  • the reinforcement environment represents the common processes of applying actions to states.
  • the RL environment may also apply constraints to the actions or to the new states. Constraints may be represented by motions limits (e.g. : translation, rotation) or shape limits (e.g. : magnitude of shape change or deformation).

Landscapes

  • Health & Medical Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Medical Informatics (AREA)
  • Public Health (AREA)
  • Biomedical Technology (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Pathology (AREA)
  • Epidemiology (AREA)
  • General Health & Medical Sciences (AREA)
  • Primary Health Care (AREA)
  • Dental Tools And Instruments Or Auxiliary Dental Instruments (AREA)

Abstract

Systems and techniques for generating a transformation for a three-dimensional (3D) representation of oral care data are disclosed. The method involves receiving oral care data describing the current state of a treatment and providing it as input to a reinforcement learning (RL) model. The processing circuitry executes one or more trained neural networks using the oral care data to generate a representation that specifies a predicted action associated with the current state of the treatment. By utilizing the current state and the predicted action, the processing circuitry generates a predicted next state of the treatment. These systems and techniques enable the generation of accurate and personalized transformations for 3D representations of oral care data, facilitating improved treatment planning and decision-making processes.

Description

REINFORCEMENT LEARNING FOR FINAL SETUPS AND INTERMEDIATE STAGING IN CLEAR TRAY ALIGNERS
Related Documents
[0001] The entire disclosure of PCT Application No. PCT/IB2022/057373 is incorporated herein by reference. The entire disclosures of each of PCT Applications with Publication Nos. WO2022123402A1, WO2021245480A1, and W02020026117A1 are incorporated herein by reference. The entire disclosure of each of the following Provisional U.S. Patent Applications is incorporated herein by reference: 63/432,627; 63/366,492; 63/366,495; 63/352,850; 63/366,490; 63/366,494; 63/370,160; 63/366,507; 63/352,877; 63/366,514; 63/366,498; 63/366,514; and 63/264,914.
Technical Field
[0002] This disclosure relates to configurations and training of neural networks to improve the accuracy of automatically generated clear tray aligner (CT A) devices used in orthodontic treatments.
Summary
[0003] Some existing techniques have attempted to use machine learning to generate the CTA devices, but with mixed results. As a result, there is a need for better machine learning models and training approaches to improve the systems that automate the production of CTAs.
[0004] The present disclosure describes systems and techniques for training and using one or more machine learning models, such as neural networks, to produce intermediate stages and final setups for CTAs, in a manner which is customized to the treatment needs of the patient. Such a neural network is termed herein as a “setups prediction neural network” or simply a “setups prediction model.” One or more neural networks may be trained using reinforcement learning (RL) to implement a decision-making agent which may generate setups predictions. These predictions may be generated through a system of rewarding desired behavior and/or punishing undesirable behavior. The agent may become trained to generate actions and learn through a process of interacting with its environment, which has the advantage of enabling the setups prediction model to learn about that environment. The Soft Actor Critic (SAC) and Generative Adversarial Imitation Learning (GAIL) reinforcement learning algorithms, among other RL models disclosed herein, may be advantageously applied to setups prediction. RL models for the placement of 3D oral care representations may encode learning from past rounds of training in tuples, which may be stored in a replay buffer. RL is particularly well suited to the prediction of transforms for 3D oral care representations, because RL may learn from experience from a minimal dataset, and RL may correct errors that occur though training. Once an error is corrected, the error may be less likely to recur, because the knowledge of the error is captured in the replay buffer for later use. RL may maintain a balance between exploitation and exploration, which is especially beneficial to the task of arranging teeth for an orthodontic setup (e.g., which may explore a space of many possible optimal configurations of teeth mixed in with many possible non-optimal configurations of teeth). A final setup (also referred to as final setups) is a target configuration of 3D tooth representations (such as 3D tooth meshes) such as the teeth appear at the end of treatment. An intermediate setup (also referred to as an “intermediate stage” or as “intermediate staging”) describes a configuration of teeth during one of the several stages of treatment, after the teeth leave their maloccluded poses (e.g., positions and/or orientations) and before the teeth reach their final setup poses. In some implementations, a final setup may be used to generate, at least in part, one or more intermediate stages. Each stage may be used in the generation of a clear tray aligner. Such aligners may incrementally move the patient's teeth from the initial or maloccluded poses to the final poses represented by the final setup.
[0005] In a first aspect, a first computer-implemented method for generating setups for orthodontic alignment treatment is described including the steps of receiving, by one or more computer processors, a first digital representation of a patient’s teeth, using, by the one or more computer processors and to determine a prediction for one or more tooth movements for a final setup, a generator that is a machine learning model, such as comprising one or more neural networks (e.g., a 3D encoder, 3D decoder, an MLP, an encoder-decoder structure, a neural network with an attention layer or other neural networks disclosed herein) that has been initially trained to predict one or more tooth movements for a final setup, further training, by the one or more computer processors, the setups prediction model based on the using, and where the training of the setups prediction model is modified by performing operations including predicting, by the generator, one or more tooth movements for a final setup based on the first digital representation of the patient’s teeth, computing a loss function which quantifies the difference between predicted tooth movements and reference tooth movements, and modifying the setups prediction model using that loss. An encoder-decoder structure may comprise at least one encoder or at least one decoder. Non-limiting examples of an encoder-decoder structure include a 3D U-Net, a transformer, a pyramid encoder-decoder or an autoencoder, among others.
[0006] The first aspect can optionally include additional features. For instance, the method can produce, by the one or more processors, an output state for the final setup. The method can determine, by the one or more computer processors, a difference between the one or more predicted tooth movements and the one or more reference tooth movements. The determined difference between the one or more predicted tooth movements and the one or more reference tooth movements can be used to modify the training of the generator. Modifying the training of the generator can include adjusting one or more weights of the generator’s neural network. The method can generate, by the one or more computer processors, one or more lists specifying mesh elements of the first digital representation of the patient’s teeth. At least one of the one or more lists can specify one or more edges in the first digital representation of the patient’s teeth. At least one of the one or more lists can specify one or more polygonal faces in the digital representation of the patient’ s teeth. At least one of the one or more lists can specify one or more vertices in the first digital representation of the patient’s teeth (e.g., such as derived from a 3D mesh). At least one of the one or more lists can specify one or more points in the first digital representation of the patient’s teeth (e.g., such as derived from a 3D point cloud). A 3D point cloud may, in some instances, comprise the plurality of vertices extracted from a 3D mesh. At least one of the one or more lists can specify one or more voxels in the first digital representation of the patient’s teeth (e.g., such as derived from a sparse representation). The method can compute, by the one or more computer processors, one or more mesh element features. In the case of edges, the one or more mesh element features can include edge endpoints, edge curvatures, edge normal vectors, edges movement vectors, edge normalized lengths, vertices, faces of associated three-dimensional representations, voxels, and combinations thereof. Other mesh element features for edges are disclosed herein. Mesh element features for each of vertices, points, faces and voxels are also disclosed herein. The method can generate, by the one or more computer processors, a digital representation predicting the position and orientation of the patient’s teeth based on the one or more predicted tooth movements. A prediction for the movement of a tooth may comprise a transform (e.g., such as one or more of an affine transformation matrix, a translation vector, a quaternion, or one or more Euler angles). The setups prediction model may predict each of tooth position and tooth orientation information. In some non-limiting examples, the network may predict the orientation and position information substantially concurrently. The setups prediction model may predict a setup transform for each tooth in the arch, to place each tooth in the final setup pose. The method can generate, by the one or more computer processors, a digital representation of the patient’s teeth based on the one or more reference tooth movements. In some non-limiting implementations, the generator of a setups prediction model may be trained, at least in part, with the assistance of a discriminator. The discriminator may determine whether a representation of the one or more tooth movements predicted by the generator is distinguishable from a representation of one or more reference tooth movements can include the steps of receiving the representation of the one or more tooth movements predicted by the generator, the representation of the one or more reference tooth movements, and the first digital representation of the patient’s teeth, comparing the representation of the one or more tooth movements predicted by the generator, the representation of the one or more reference tooth movements, wherein the comparison is based at least in part on the first digital representation of the patient’s teeth, and determining, by the one or more computer processors, a probability that the representation of the one or more tooth movements predicted by the generator is the same as the representation of one or more reference tooth movements.
[0007] In a second aspect, a second computer-implemented method for generating setups for orthodontic alignment treatment pertains to intermediate staging prediction. Intermediate staging of teeth from a malocclusion stage to a final stage requires determining accurate individual teeth movements in a way that teeth are not colliding with each other, the teeth move toward their final state, and the teeth follow optimal and preferably short trajectories. Because each tooth has six degrees-of-freedom and an average arch has about fourteen teeth, finding the optimal teeth trajectory from initial to final stage is a large and complex problem.
[0008] The second computer-implemented method is customized to the treatment needs of the patient (e.g., as specified by a clinician, which may include technician or healthcare professional) and is described including the steps of receiving, by one or more computer processors, a first digital representation of a patient’s teeth, and a representation of a final setup, using, by the one or more computer processors and to determine a prediction for one or more tooth movements for one or more intermediate stages, a generator that is a machine learning model, such as a neural network, included in a setups prediction machine learning model, such as comprising one or more neural networks (e.g., a 3D encoder, 3D decoder, a 3D U-Net, a multilayer perceptron (MLP), a transformer, an autoencoder, a pyramid encoder-decoder, and other neural networks disclosed herein), and that has been initially trained to predict one or more tooth movements for one or more intermediate stages, further training, by the one or more computer processors, the setups prediction model based on the using, wherein the training of the setups prediction model is modified by performing operations including predicting, by the generator, one or more tooth movements for at least one intermediate stage based on the first digital representation of the patient’s teeth, computing a loss function which quantifies the difference between predicted tooth movements and reference tooth movements and modifying the setups prediction model using that loss. The second aspect can also include one or more of the optional features described above in reference to the first aspect.
[0009] Techniques of this disclosure relate to the automatic generation of transformation for a three- dimensional (3D) representations of oral care data (e.g., teeth, etc.). The methods involve receiving oral care data describing the current state of a treatment and providing that data as input to a reinforcement learning (RL) model. Trained neural networks are then executed using the oral care data to generate a representation that specifies a predicted action associated with the treatment of the patient. The methods may generate a predicted next state of the treatment based on the current state and the action.
[0010] The oral care data describing the current state of the treatment can include 3D representations of the oral care data (e.g., tooth meshes, appliance components, fixture model components, etc.) or transforms that modify the poses of the 3D representations of oral care data (e.g,. tooth transforms). Reward values may be computed to quantify the difference between a predicted next state and a corresponding ground truth next state. Loss values may be computed to quantify the difference between a predicted action and a corresponding ground truth action, in some implementations, a loss may be computed, at least in part, by a reward (i.e., a reward may be considered by a loss calculation).
[0011] A 3D representation of oral care data can include a 3D mesh, a 3D point cloud, a 3D voxelized representation, or a 3D surface. A 3D representation of oral care data may represent teeth, appliance components, or fixture model components, among other 3D oral care representations described herein.
[0012] The methods can be used in digital oral care, such as orthodontic alignment treatment. The trained neural network can be executed according to a policy -based or policy -free methodology, or the neural network can be trained using various paradigms such as Q-leaming, soft actor critic (SAC) learning, or generative adversarial imitation learning (GAIL).
[0013] The computing device can be deployed in a clinical context and perform the method in near real-time during a patient encounter. The RL model may, in some non-limiting implementations, be trained through reinforcement learning by human feedback (RLHF). Additionally, the RL model can be used to generate designs for orthodontic appliances, or dental restoration appliances. [0014] The method can also involve providing additional input data to the RL model, such as 3D geometries describing teeth, vectors containing values for computing dimensions and distances between teeth, latent vector information, position and orientation vectors, or tooth-related information.
[0015] The computing device includes interface hardware for receiving the 3D representation of oral care data and processing circuitry for executing the method steps.
[0016] Overall, this disclosure describes efficient and effective methods for generating transformations for 3D representations of oral care data, particularly in orthodontic alignment treatment. [0017] Stated another way, methods of this disclosure may generate transformations for a three- dimensional (3D) representations of oral care data (e.g., 3D meshes of teeth, or appliance components). The methods may be trained, at least in part, using reinforcement learning (RL). The methods may operate, in deployment, through the use of the reinforcement learning architecture described herein. Such RL models (e.g., RL Setups for orthodontic setup generation), may generate transforms to place teeth, appliance components or other 3D representations of oral care data into poses relative to other 3D representations of oral care data. A 3D representation of oral care data is at least one of a 3D mesh, a 3D point cloud, a 3D voxelized representation, or a 3D surface. For example, a 3D mesh may represent a tooth in a dental arch, contained within a patient case. In some implementations, an actor of the RL model may be configmed to issue one or more actions associated with moving at least one tooth represented in the 3D oral care representation to generate an orthodontic setup transform. In some implementations, when the one or more actions is issued, the actor of the RL model may be configmed to predict a transform associated with moving a tooth (e.g., for orthodontic setups generation), an appliance component (e.g., dental restoration appliance generation, etc.), or a fixtme model component (e.g., for fixture model generation). In some implementations, the actor may contain one or more nemal networks. In some implementations, the actor may be trained using on at least one of: (i) positive rewards that are issued to reward desired behaviors, or (ii) negative rewards that me issued to penalize undesired behaviors. The RL models of this disclosure may, in some implementations, be executed according to a policy -based methodology, or according to a policy -free methodology. The RL models may, in some implementations, be trained according to a Q-leaming paradigm. The RL models may, in some implementations, be trained according to a soft actor critic (SAC) learning paradigm. The RL models may, in some implementations, be trained according to a generative adversarial imitation learning (GAIL) paradigm. In some implementations, the transformations predicted by the methods may be used for orthodontic setups generation (e.g., for alignment treatment). A setup transformation may be associated with transforming one or more teeth of the patient into a final setup (or intermediate stage). A final setup may describe the poses of the patient’s teeth after completion of an orthodontic treatment. An intermediate stage may describe the poses of the patient’s teeth during the course of an orthodontic treatment of a patient. The methods may, in some instances, be deployed at a clinical context. The RL model may be trained, at least in part, through reinforcement learning by human feedback (RLHF). The RL model may, in some implementations, be used to generate a design for a dental restoration appliance, or an orthodontic appliance (e.g., a clear tray aligner), among others. Inputs which may be provided to the RL models of this disclosure, in some implementations, include one or more of: (i) one or more 3D geometries describing one or more teeth, (ii) one or more vectors P containing at least one value pertaining to at least one method of computing a dimension of at least one tooth, (iii) one or more vectors Q containing at least one value pertaining to at least one method of computing a distance between adjacent teeth, (iv) one or more vectors B containing latent vector information about one or more teeth, (v) one or more vectors N containing at least one value pertaining to the position of at least one tooth, (vi) one or more vectors O containing at least one value pertaining to the orientation of at least one tooth, (vii) one or more vectors R at least one of tooth name, designation, tooth type and tooth classification.
Brief Description of Drawings
[0018] FIG. 1 shows a method of augmenting training data for use in training machine learning (ML) models of this disclosure.
[0005] FIG. 2 shows a summary of some of the setups prediction methods described herein.
[0006] FIG. 3 shows a method of using reinforcement learning to train an ML model to generate transforms for 3D representations (e.g., tooth transforms, or appliance component transforms).
[0007] FIG. 4 shows an example environment.
[0008] FIG. 5 summarizes a method of training an ML model using reinforcement learning, according to techniques of this disclosure.
[0009] FIG. 6 various neural networks trained for use in the reinforcement learning methods of this disclosure.
[0010] FIG. 7 shows transformer which may be configured to generate orthodontic setups transforms.
[0011] FIG. 8 shows an example method of using reinforcement learning with human feedback to train ML models of this disclosure.
Detailed Description
[0019] Described herein are techniques for the automatic prediction of setups, which may provide the advantage of improving accuracy in comparison to existing techniques, enable new clinicians to be trained in the generation of effective setups, enable customized setups to be produced (e.g., which align with the specifications of clinicians), and provide the technical improvement of enhanced data precision in the formulation of these setups.
[0020] A setups prediction model of this disclosure may receive a variety of input data, which, as described herein, may include tooth meshes representing one or both arches of the patient. The tooth data may be presented in the form of 3D representations, such as meshes or point clouds. These data may be preprocessed, for example, by arranging the constituent mesh elements into lists and computing an optional mesh element feature vector for each mesh element. Such vectors may impart valuable information of the shape and/or structure of the tooth to the setups prediction neural network. Additional inputs may enable the setups prediction neural network to better understand the distribution of the inputted data (e.g., tooth meshes), which provides the technical improvement of enabling customization to the specific medical/dental needs of the patient when the setups prediction model is deployed. For example, one or more oral care metrics may be computed. Oral care metrics may be used for measuring one or more physical aspects of a setup (e.g., physical relationships within a tooth or between teeth). In some instances, an orthodontic metric may be computed for a ground truth setup which is then used in the training of a machine learning model (e.g., a setups prediction model). The metric value may be received at the input of the setups prediction model, as a way of training the model to encode a distribution of such a metric over the several examples of the training dataset. For example, an “overbiteleff ’ metric may be computed for a setup which is received by the setups prediction model (e.g., at least one of mal and approved setup). During training, the network may then receive this metric value as an input, to assist in training the network to link that inputted metric value to the physical aspects of the received setup (e.g., to learn a distribution over the possible values of that metric across the examples of the training dataset). The metric may be computed for the mal setup, and that metric value be providedas an input the network during training, alongside the malocclusion transforms and/or tooth meshes. The metric may also (or alternatively) be computed for the approved setup, and that metric be provided as an input to the network during training, alongside the approved setup transforms and/or tooth meshes (e.g., for application during loss calculation time). Such a loss calculation may quantify the difference between a prediction and a ground truth example (e.g., between a predicted setup and a ground truth setup). By providing the network a metric value at training time, the network may, through the course of loss calculation and subsequent backpropagation, learn to encode a distribution of that metric. A technical improvement provided by the setups prediction techniques described herein is the customization of orthodontic treatment to the patient. Oral care parameters may enable a clinician to customize specific desired aspects of the dimensions, proportions and other physical aspects of a predicted setup. For example, in deployment, one or more oral care parameters (procedure parameters or restoration design parameters) may be defined and provided to the trained setups prediction model as part of the execution-phase input to specify one or more aspects of an intended setup upon an execution run. in some implementations, a procedure parameter may be defined which corresponds to an oral care metric (e.g., such as the overbiteleft metric described above), which may be received at the input to a deployed setups prediction neural network and be taken as an instruction to the setups prediction neural network to generate a setup with the specified quantity of the metric (e.g., overbiteleft). The setups prediction model may be especially suited to generating a setup with a prescribed value of a procedure parameter in the circumstance where that prescribed value falls within the distribution of the corresponding metric value that appeared in the training dataset. Other procedure parameters may also be defined corresponding to other orthodontic metrics and be taken as instructions to the setups prediction model for the quantity of the relevant metric that is to be imparted to the predicted setup. This interplay between oral care metrics and oral care parameters may also apply to the training and deployment of other predictive models in oral care as well.
[0021] To train the setups prediction neural network effectively, aspects of this disclosure are directed to forming training data that have a distribution which describes the kind of setup that the setups prediction neural network is configmed to produce. For example, to produce a final setup with an overbite of approximately 2.0 mm, one approach is to use ground truth training data with an overbite of approximately 2.0 mm. This approach may lead to a clean training signal and may produce useful results, and an alternative method may enable the network to learn to account for differences in overbite among the various ground truth training samples in the training dataset. An overbite metric may be computed for the malocclusion arches of a training sample (a patient case). This overbite value may be received as an input to the setups prediction neural network at training time, along with the maloccluded tooth data, and serve as a signal to the neural network regarding the magnitude of overbite present in that mal arch. The network thereby learns that different cases have different overbite magnitudes and can encode a distribution of possible overbite magnitudes, which can then be imparted to the predicted setup. Upon deployment, the trained neural network may receive the maloccluded tooth data as input and may also receive an input to indicate a magnitude of the overbite (e.g., or some other oral care metric) that is desired in the predicted setup (e.g., in the form of a procedure parameter which has been defined for the purpose). This approach may enable the setups prediction neural network to account for differences in the distribution of the training dataset without excluding patient cases from the training dataset (e.g., as may be done in the case of filtering the training dataset), with the added benefit of enabling the deployed setups prediction neural network to customize the predicted setup, according to the specification of the clinician who uses the setups prediction model. Other orthodontic metrics (e.g., those disclosed herein) may also be computed in keeping with this technique. Corresponding procedure parameters (e.g., those disclosed herein or those defined to correspond to specific metrics) may be provided to the trained network to effect the customization of the outputted setups prediction. Other techniques disclosed herein, besides setups prediction, may also be trained with this use of oral care metrics and procedure parameters being received as inputs to a predictive model.
[0022] A setups prediction neural network of this disclosure may be trained, at least in part, by the calculation of one or more loss values (e.g., reconstruction loss or other loss values described herein). Such loss values may quantify the difference between a predicted setup and a corresponding ground truth setup. In some instances, these setups may be registered with each other (e.g., using iterative closest point (ICP) or singular value decomposition (SVD)) before the loss is computed, to reduce noise and improve the accuracy of the resulting trained setups prediction neural network. Such a registration may alternatively or additionally be performed between the maloccluded setup and the corresponding ground truth setup, with the advantage of reducing noise in the loss measurement and improving the accuracy of the trained network. [0023] The setups prediction neural network may compute a transform for each tooth, to move that tooth into a pose which is suitable for the end of orthodontic treatment (e.g., the final setup). The pose of the tooth may include a change in position in 3D space and may also include a change in orientation (e.g., with respect to one or more coordinate axes - e.g., local coordinate axes with origin at the crown centroid). The transform may effect the change in orientation by pivoting the tooth mesh relative to a pivot point or tooth origin. This pivot point may be chosen to lie within the crown centroid. Alternatives include at the apex of the root tip, origin of malocclusion transform or at a point along an archform.
[0024] In some implementations, the setups prediction neural network may be trained conditionally on interproximal reduction (IPR) information. IPR may be applied to the teeth, to enable greater packing of teeth a in final setup. The setups model may be trained to account to IPR quantities (e.g., millimeters of offset in from either or both of the mesial and distal sides of a tooth) and/or IPR cut planes (which may be used in conjunction with mesh Boolean operations to remove material on either or both of the mesial and distal sides of a tooth). For example, IPR cut planes may be used to modify one or more tooth meshes for one or more patient cases which are used to train the setups prediction model. This step improves the accuracy of setups prediction model training by improving data precision, because material is removed from the teeth which may otherwise lead to collisions between teeth in the final setup (and result in noise in the training data). After the trained setups prediction model is deployed, IPR may be applied to a trial patient case, to modify the shapes of the teeth before the case is received as input to the setups prediction model. In some instances, IPR may be applied to one or more tooth meshes of a patient case before the computation of orthodontic metrics.
[0025] In orthodontics, an anterior posterior (AP) shift may involve a sagittal shift of the mandible (lower arch), moving the mandible either forward or backwards. The application of the AP Shift may improve the class relationship of the teeth. Class may describe the patient’s malocclusion. Possible classes include: class 1, class 2 or class 3. Elastics may aid in the shift of the mandible. Such elastics may attach to hardware on the teeth, such as buttons. In some instances, the setups prediction model of this disclosure may directly receive an AP shift transform as an input, which may improve the data precision of the resulting model. In some instances, an AP shift transform may first be applied to the patient case data before the patient case data are received as input to the setups prediction model of this disclosure.
[0026] The predictive models of the present disclosure may, in some implementations, may produce more accurate results by the incorporation of one or more of the following inputs: archform information V, interproximal reduction (IPR) information U, tooth dimension information P, tooth gap information Q, latent capsule representations of oral care meshes T, latent vector representations of oral care meshes A, procedure parameters K (which may describe a clinician’s intended treatment of the patient), doctor preferences L (which may describe the typical procedure parameters chosen by a doctor), flags regarding tooth status M (such as for fixed or pinned teeth), tooth position information N, tooth orientation information O, tooth name/dental notation R, oral care metrics S (comprising at least one of oral care metrics and restoration design metrics).
[0027] Systems of this disclosure may, in some instances, be deployed at a clinical context (such as a dental or orthodontic office) for use by clinicians (e.g., doctors, dentists, orthodontists, nurses, hygienists, oral care technicians). Such systems which are deployed at a clinical context may enable clinicians to process oral care data (such as dental scans) in the clinic environment, or in some instances, in a "chairside" context (where the patient is present in the clinical environment). A non-limiting list of examples of techniques may include: segmentation, mesh cleanup, coordinate system prediction, CTA trimline generation, restoration design generation, appliance component generation or placement or assembly, generation of other oral care meshes, the validation of oral care meshes, setups prediction, removal of hardware from tooth meshes, hardware placement on teeth, imputation of missing values, clustering on oral care data, oral care mesh classification, setups comparison, metrics calculation, or metrics visualization. The execution of these techniques may, in some instances, enable patient data to be processed, analyzed and used in appliance generate by the clinician before the patient leaves the clinical environment (which may facilitate treatment planning because feedback may be received from the patient during the treatment planning process).
[0028] Systems of this disclosure may automate operations in digital orthodontics (e.g., setups prediction, hardware placement, setups comparison), in digital dentistry (e.g., restoration design generation) or in combinations thereof. Some techniques may apply to either or both of digital orthodontics and digital dentistry. A non-limiting list of examples is as follows: segmentation, mesh cleanup, coordinate system prediction, oral care mesh validation, imputation of oral care parameters, oral care mesh generation or modification (e.g., using autoencoders, transformers, continuous normalizing flows or denoising diffusion models), metrics visualization, appliance component placement or appliance component generation or the like. In some instances, systems of this disclosure may enable a clinician or technician to process oral care data (such as scanned dental arches). In addition to segmentation, mesh cleanup, coordinate system prediction or validation operations, the systems of this disclosure may enable orthodontic treatment planning, which may involve setups prediction as at least one operation. Systems of this disclosure may also enable restoration design generation, where one or more restored tooth designs are generated and processed in the course of creating oral care appliances. Systems of this disclosure may enable either or both of orthodontic or dental treatment planning, or may enable automation steps in the generation of either or both of orthodontic or dental appliances. Some appliances may enable both of dental and orthodontic treatment, while other appliances may enable one or the other.
[0029] Techniques of this disclosure may require a training dataset of hundreds or thousands of cohort patient cases, to ensure that the neural network is able to encode the distribution of patient cases which are likely to be encountered in clinical treatment. A cohort patient case may include a set of tooth crown meshes, a set of tooth root meshes, or a data file containing attributes of the case (e.g., a JSON file). A typical example of a cohort patient case may contain up to 32 crown meshes (e.g., which may each contain tens of thousands of vertices or tens of thousands of faces), up to 32 root meshes (e.g., which may each contain tens of thousands of vertices or tens of thousands of faces), multiple gingiva mesh (e.g., which may each contain tens of thousands of vertices or tens of thousands of faces) or one or more JSON files which may each contain tens of thousands of values (e.g., objects, arrays, strings, real values, Boolean values or Null values).
[0030] Aspects of the present disclosure can provide a technical solution to the technical problem of predicting, using a machine learning model which has been trained using reinforcement learning, orthodontic setups for use in oral care appliance generation (e.g., intermediate stages or final setups for the generation of aligner trays). Aspects of the present disclosure may need to be executed in a time- constrained manner, such as when an oral care appliance must be generated for a patient immediately after intraoral scanning (e.g., while the patient waits in the clinician’s office). As such, aspects of the present disclosure are necessarily rooted in the underlying computer technology of setups transform prediction for oral care appliance generation and cannot be performed by a human, even with the aid of pen and paper. For instance, implementations of the present disclosure must be capable of: 1) storing thousands or millions of mesh elements of the patient’s dentition in a manner that can be processed by a computer processor; 2) performing calculation on thousands or millions of mesh elements, e.g., to quantify aspects of the shape and or/structure of an individual tooth in the 3D representation of the patient’s dentition; and 3) predicting, based on a machine learning model which has been trained using reinforcement learning, orthodontic setups for use in oral care appliance generation), and do so during the course of a short office visit.
[0031] This disclosure pertains to digital oral care, which encompasses the fields of digital dentistry and digital orthodontics. This disclosure generally describes methods of processing three-dimensional (3D) representations of oral care data. It should be understood, without loss of generality, that there are various types of 3D representations. One type of 3D representation is a 3D geometry. A 3D representation may include, be, or be part of one or more of a 3D polygon mesh, a 3D point cloud (e.g., such as derived from a 3D mesh), a 3D voxelized representation (e.g., a collection of voxels - for sparse processing), or 3D representations which are described by mathematical equations. Although the term “mesh” is used frequently throughout this disclosure, the term should be understood, in some implementations, to be interchangeable with other types of 3D representations. A 3D representation may describe elements of the 3D geometry and/or 3D structure of an object.
[0032] Dental arches SI, S2, S3 and S4 all contain the exact same tooth meshes, but those tooth meshes are transformed differently, according to the following description. A first arch S 1 includes a set of tooth meshes arranged (e.g., using transforms) in their positions in the mouth, where the teeth are in the mal positions and orientations. A second arch S2 includes the same set of tooth meshes from SI arranged (e.g., using transforms) in their positions in the mouth, where the teeth are in the ground truth setup positions and orientations. A third arch S3 includes the same meshes as SI and S2, which are arranged (e.g., using transforms) in their positions in the mouth, where the teeth are in the predicted final setup poses (e.g., as predicted by one or more of the techniques of this disclosure). S4 is a counterpart to S3, where the teeth are in the poses corresponding to one of the several intermediate stages of orthodontic treatment with clear tray aligners.
[0033] It should be understood, without the loss of generality, that the techniques of this disclosure which apply to final setups are also applicable to intermediate staging in orthodontic treatment, particularly geometric deep learning (GDL) Setups, reinforcement learning (RL) Setups, variational autoencoder (VAE) Setups, Capsule Setups, multilayer perceptron (MLP) Setups, Diffusion Setups, pose transfer (PT) Setups, Similarity Setups, force directed graphs (FDG) Setups, Transformer Setups, Setups Comparison, or Setups Classification. The Metrics Visualization aspects of this disclosure may also be configured to visualize data from both final setups and intermediate stages. MLP Setups, VAE Setups and Capsule Setups each fall within the scope of Autoencoder Setups. Some implementations of MLP Setups may fall within the Scope of Transformer Setups. FIG. 2 shows a non-limiting selection of models which may be trained for setups prediction. Representation Setups refers to any of MLP Setups, VAE Setups, Capsule Setups and any other setups prediction machine learning model which uses an autoencoder to create the representation for at least one tooth. [0034] Each of the setups prediction techniques of this disclosure is applicable to the fabrication of clear tray aligners and/or indirect bonding trays. The setups predictions techniques may also be applicable to other products that involve final teeth poses, also. A pose may comprise a position (or location) and a rotation (or orientation).
[0035] A 3D mesh is a data structure which may describe the geometry or shape of an object related to oral care, including but not limited to a tooth, a hardware element, or a patient’s gum tissue. A 3D mesh may include one or more mesh elements such as one or more of vertices, edges, faces and combinations thereof. In some implementations, mesh element may include voxels, such as in the context of sparse mesh processing operations. Various spatial and structural features may be computed for these mesh elements and be provided to the predictive models of this disclosure, with the predictive models of this disclosure providing the technical advantage of improving data precision in the form of the models of this disclosure outputting more accurate predictions.
[0036] A patient’s dentition may include one or more 3D representations of the patient’s teeth (e.g., and/or associated transforms), gums and/or other oral anatomy. An orthodontic metric (OM) may, in some implementations, quantify the relative positions and/or orientations of at least one 3D representation of a tooth relative to at least one other 3D representation of a tooth. A restoration design metric (RDM) may, in some implementations, quantify the at least one aspect of the structure and/or shape of a 3D representation of a tooth. An orthodontic landmark (OL) may, in some implementations, locate one or more points or other structural regions of interest on a 3D representation of a tooth. An OL may, in some implementations, be used in the generation of an orthodontic or dental appliance, such as a clear tray aligner or a dental restoration appliance. A mesh element may, in some implementations, comprise at least one constituent element of a 3D representation of oral care data. For example, in the case of a tooth that is represented by a 3D mesh, mesh elements may include at least: vertices, edges, faces and voxels. A mesh element feature may, in some implementations, quantify some aspect of a 3D representation in proximity to or in relation with one or more mesh elements, as described elsewhere in this disclosure. Orthodontic procedure parameters (OPP) may, in some implementations, specify at least one value which defines at least one aspect of planned orthodontic treatment for the patient (e.g., specifying desired target attributes of a final setup in final setups prediction). Orthodontic Doctor preferences (ODP) may, in some implementations, specify at least one typical value for an OPP, which may, in some instances, be derived from past cases which have been treated by one or more oral care practitioners. Restoration Design Parameters (RDP) may, in some implementations, specify at least one value which defines at least one aspect of planned dental restoration treatment for the patient (e.g., specifying desired target attributes of a tooth which is to undergo treatment with a dental restoration appliance). Doctor Restoration Design Preferences (DRDP) may, in some implementations, specify at least one typical value for an RDP, which may, in some instances, be derived from past cases which have been treated by one or more oral care practitioners. 3D oral care representations may include, but are not limited to: 1) a set of mesh element labels which may be applied to the 3D mesh elements of teeth/gums/hardware/appliance meshes (or point clouds) in the course of mesh segmentation or mesh cleanup; 2) 3D representation(s) for one or more teeth/gums/hardware/appliances for which shapes have been modified (e.g., trimmed, distorted, or filled- in) in the course of mesh segmentation or mesh cleanup; 3) one or more coordinate systems (e.g., describing one, two, three or more coordinate axes) for a single tooth or a group of teeth (such as a full arch - as with the LDE coordinate system); 4) 3D representation(s) for one or more teeth for which shapes have been modified or otherwise made suitable for use in dental restoration; 5) 3D representation(s) for one or more dental restoration appliance components; 6) one or more transforms to be applied to one or more of: dental restoration appliance library component placement relative to one or more teeth, a tooth to be placed for an orthodontic setup (either final setup or intermediate stage), a hardware element to be placed relative to one or more teeth or the like; 7) an orthodontic setup; 8) a 3D representation of a hardware element (such as facial bracket, lingual bracket, orthodontic attachment, button, hook, bite ramp, etc.) to be placed relative to one or more teeth, etc.; 8) a 3D representation of a bonding pad for a hardware element (which may be generated for a specific tooth by outlining a perimeter on the tooth, specifying a thickness to form a shell, and then subtracting-out the tooth via a Boolean operation); 9) 3D representation of a clear tray aligner (CTA); 10) the location or shape of a CTA trimline (e.g., described as either a mesh or polyline); 11) archform that describes the contours or layout of an arch of teeth (e.g., described as a 3D polyline or as a 3D mesh or surface), which may follow the incisal edges one or more teeth, which may follow the facial surfaces of one or more teeth, which may in some implementations correspond to the maloccluded arch and in other implementations correspond to the final setup arch (the effects of malocclusion on the shape of the archform may be diminished by smoothing or averaging of the shape of the archform), which may be described by one or more control points and/or a spline; 12) 3D representation of a fixture models (e.g., depictions of teeth and gums for use in thermoforming clear tray aligners, or depictions of teeth/gums/hardware for use in thermoforming indirect bonding trays); 13) one or more latent space vectors (or latent capsules) produced by the 3D encoder stage of a 3D autoencoder which has been trained on the reconstruction of oral care meshes (e.g., a variational autoencoder which has been trained for tooth reconstruction); 14) one or more oral care metrics values (e.g., such as orthodontic metrics or restoration design generation metrics) for one or more teeth; 15) one or more landmarks (e.g., 3D points) which describe the shapes and/or geometrical attributes of one or more teeth, other dentition structures or hardware structures (e.g., to be used in orthodontic setups creation or restoration appliance component generation or placement); 16) 3D representation created by scanning (e.g., optically scanning, CT scanning or MRI scanning) a 3D printed part corresponding to one or more teeth/gums/hardware/appliances (e.g., a scanned fixture model); 17) 3D printed aligners (including optionally local thickness, reinforcing rib geometry, flap positioning, or the like) 18) 3D representation of the patient's dentition that was captured chairside by a clinician or medical practitioner (e.g., in a context where the 3D representation is validated chairside, before the patient leaves the clinic, so that errors can be detected and re-scans performed as necessary); 19) dental restoration tooth design (e.g., for veneers, crowns, bridges or dental restoration appliances); 20) 3D representations of one or more teeth for use in digital oral care treatment; 21) other 3D printed parts pertaining to oral care procedures or other fields; 22) IPR cut surfaces; 23) one or more orthodontic setups transforms associated with one or more IPR cut surfaces; 24) a (digital) pontic tooth design which may fill at least a portion of the space between teeth to allow room in an orthodontic setup for an erupting tooth to later emerge from the gums; or 25) a component of a fixture model (e.g., comprising fixture model components such as interproximal webbing, block-out, bite locks, bite ramps, interproximal reinforcement, gingival ridges, torque points, power ridges, pontic tooth or dimples, among others).
[0037] The techniques of this disclosure may be advantageously combined. For example, the Setups Comparison tool may be used to compare the output of the GDL Setups model against ground truth data, compare the output of the RL Setups model against ground truth data, compare the output of the VAE Setups model against ground truth data and compare the output of the MLP Setups model against ground truth data. With each of these setups prediction models compared against ground truth data, it may be possible to determine which model gives the best performance on a certain dataset or within a given problem domain. Furthermore, the Metrics Visualization tool can enable a global view of the final setups and intermediate stages produced by one or more of the setups prediction models, with the advantage of enabling the selection of the best setups prediction model. The Metrics Visualization tool, furthermore, enables the computation of metrics which have a global scope over a set of intermediate stages. These global metrics may, in some implementations, be consumed as inputs to the neural networks for predicting setups (e.g., GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, among others). The global metrics may also be consumed by FDG Setups. The local metrics from this disclosure (i.e., a local metric is a metric which may be computed for one stage or setup of treatment, rather than over several stages or setups) may, in some implementations, be consumed by the neural networks herein for predicting setups, with the advantage of improving predictive results. The metrics described in this disclosure may, in some implementations, be visualized using the Metric Visualization tool.
[0038] The VAE and MAE models for mesh element labelling and mesh in-filling can be advantageously combined with the setups prediction neural networks, for the purpose of mesh cleanup ahead of or during the prediction process. In some implementations, the VAE for mesh element labelling may be used to flag mesh elements for further processing, such as metrics calculation, removal or modification. In some instances, such flagged mesh elements may be used as inputs to a setups prediction neural network, to inform that neural network about important mesh features, attributes or geometries, with the advantage of improving the performance of the resulting setups prediction model. In some implementations, mesh in-filling may cause the geometry of a tooth to become more nearly complete, enabling the better functioning of a setups prediction model (i.e., improved correctness of prediction on account of better-formed geometry). In some instances, a neural network to classify a setup (i.e., the Setups Classifier) may aid in the functioning of a setups prediction neural network, because the setups classifier tells that setups prediction neural network when the predicted setup is acceptable for use and can be provided to a method for aligner tray generation. A Setups Classifier (e.g., GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups and FDG Setups, among others) may aid in the generation of final setups and also in the generation of intermediate stages. Furthermore, a Setups Classifier neural network may be combined with the Metrics Visualization tool. In other implementations, a Setups Classification neural network may be combined with the Setups Comparison tool (e.g., the Setup Comparison tool may output an indication of how a setup produced in part by the Setups Classifier compares to a setup produced by another setups prediction method). In some implementations, the VAE for mesh element labelling may identify one or more mesh elements for use in a metrics calculation. The resulting metrics outputs may be visualized by the Metrics Visualization tool.
[0039] In some examples, the Setups Classifier neural network may aid in the setups prediction technique described in U.S. Patent Application No. US20210259808A1 (which is incorporated herein by reference in its entirety) or the setups prediction technique described in PCT Application with Publication No. WO2021245480A1 (which is incorporated herein by reference in its entirety) or in PCT Application No. PCT/IB2022/057373 (which is incorporated herein by reference in its entirety). The Setups Classifier would help one or more of those techniques to know when the predicted final setup is most nearly correct. In some instances, the Setups Classifier neural network may output an indication of how far away from final setup a given setup is (i.e., a progress indicator).
[0040] In some implementations, the latent space embedding vector(s) from the reconstruction VAE can be concatenated with the inputs to the setups prediction neural network described in WO2021245480A1. The latent space vectors can also be incorporated as inputs to the other setups prediction models: GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups and Diffusion Setups, among others. The advantage is to impart the reconstruction characteristics (e.g., latent vector dimensions of a tooth mesh) to that neural network, hence improving the generated setups prediction.
[0041] In some examples, the various setups prediction neural networks of this disclosure may work together to produce the setups required for orthodontic treatment. For example, the GDL Setups model may produce a final setup, and the RL Setups model may use that final setup as input to produce a series of intermediate stages setups. Alternatively, the VAE Setups model (or the MLP Setups model) may create a final setup which may be used by an RL Setups model to produce a series of intermediate stages setups. In some implementations, a setup prediction may be produced by one setups prediction neural network, and then taken as input to another setups prediction neural network for further improvements and adjustments to be made, in some implementations, such improvements may be performed in iterative fashion.
[0042] In some implementations, a setups validation model, such as the model disclosed in US Provisional Application No. US63/366495, may be involved in this iterative setups prediction loop. First a setup may be generated (e.g., using a model trained for setups prediction, such as GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups and FDG Setups, among others), then the setup undergoes validation. If the setup passes validation, the setup may be outputted for use. If the setup fails validation, the setup may be sent back to one or more of the setups prediction models for corrections, improvements and/or adjustments. In some instances, the setups validation model may output an indication of what is wrong with the setup, enabling the setups generation model to make an improved version upon the next iteration. The process iterates until done. [0043] Generally speaking, in some implementations, two or more of the following techniques of the present disclosure may be combined in the course of orthodontic and/or dental treatment: GDL Setups, Setups Classification, Reinforcement Learning (RL) Setups, Setups Comparison, Autoencoder Setups (VAE Setups or Capsule Setups), VAE Mesh Element Labeling, Masked Autoencoder (MAE) Mesh Infilling, Multi-Layer Perceptron (MLP) Setups, Metrics Visualization, Imputation of Missing Oral Care Parameters Values, Tooth Classification Using Latent Vector, FDG Setups, Pose Transfer Setups, Restoration Design Metrics Calculation, Neural Network Techniques for Dental Restoration and/or Orthodontics (e.g., 3D Oral Care Representation Generation or Modification Using Transformers), Landmark-based (LB) Setups, Diffusion Setups, Imputation of Tooth Movement Procedures, Capsule Autoencoder Segmentation, Diffusion Segmentation, Similarity Setups, Validation of Oral Care Representations (e.g., using autoencoders), Coordinate System Prediction, Restoration Design Generation or 3D Oral Care Representation Generation or Modification Using Diffusion Models.
[0044] Oral care parameters may include one or more values that specify orthodontic procedure parameters, or restoration design parameters (RDP), as described herein. Oral care parameters may define one or more intended aspects of a 3D oral care representation, and may be provided to an ML model to promote that ML model to generate output which may be used in the generation of oral care appliances that are suitable for the treatment of a patient. Other types of values include doctor preferences and restoration design preferences, as described herein. Doctor preferences and restoration design preferences may define the typical treatment choices or practices of a particular clinician. Restoration design preferences are subjective to a particular clinician, and so differ from restoration design parameters. In some implementations, doctor preferences or restoration design preferences may be computed by unsupervised means, such as clustering, which may determine the typical values that a clinician uses in patient treatment. Those typical values may be stored in a datastore, and recalled to be provided to an automated ML model as default values (e.g., default values which may be modified before execution of the model).
[0045] For example, one clinician may prefer one value for a restoration design parameter (RDP), while another clinician may prefer a different value for that RDP, when faced with a similar diagnosis or treatment protocol. One example of such an RDP is dental restoration style. Procedure parameters and/or doctor preferences may, in some implementations, be provided to a setups prediction model for orthodontic treatment, for the purpose of improving the customization of the resulting orthodontic appliance. Restoration design parameters and doctor restoration preferences may in some implementations be used to design tooth geometry for use in the creation of a dental restoration appliance, for the purpose of improving the customization of that appliance. In addition to oral care parameters, doctor preferences, and doctor restoration preferences, some implementations of ML prediction models of this disclosure, in orthodontic treatment, may also take as input a setup (e.g., an arrangement of teeth). In some such implementations, an ML prediction model of this disclosure may take as input a final setup (i.e., final arrangement of teeth), such as in the case of a prediction model trained to generate intermediate stages. For simplicity, these preferences are referred to as doctor restoration preferences, but it is intended to be used in a non-limiting sense. Specifically, it should be appreciated that these preferences may be specified by any treating or otherwise appropriate medical professional and are not intended to be limited to doctor preferences per se (i.e., preferences from someone in possession of an M.D. or equivalent degree).
[0046] An oral care professional or clinician, such as a dentist or orthodontist, may specify information about patient treatment in the form of a patient-specific set of procedure parameters. In some instances, an oral care professional may specify a set of general preferences (aka doctor preferences) for use over a broad range of cases, to use as default values in the set of procedure parameters specification process. Oral care parameters may in some implementations be incorporated into the techniques described in this disclosure, such as one or more of GDL Setups, VAE Setups, RL Setups, Setups Comparison, Setups Classification, VAE Mesh Element Labelling, MAE Mesh In-Filling, Validation Using Autoencoders, Imputation of Missing Procedure Parameters Values, Metrics Visualization, or FDG Setups. One or more of these models may take as input one or more procedure parameters vector K and/or one or more doctor preference vectors L. In some implementations, one or more of these models may introduce to one or more of a neural network’s hidden layers one or more procedure parameters vector K and/or one or more doctor preferences vectors L. In some implementations, one or more of these models may introduce either or both of K and L to a mathematical calculation, such as a force calculation, for the purpose of improving that calculation and the ultimate customization of the resulting appliance to the patient.
[0047] Some implementations of a neural network for predicting a setup (such as GDL Setup, VAE Setup or RL Setup) may incorporate information from an oral care professional (aka doctor). This information may influence the arrangement of teeth in the final setup, bringing the positions and orientations of the teeth into conformance with a specification set by the doctor, within tolerances. In some implementations of the GDL Setup model, oral care parameters may be provided directly into the generator network as a separate input alongside the mesh data. In some implementations of GDL Setups, oral care parameters may be incorporated into the feature vector which is computed for each mesh element before the mesh elements are input to the generator for processing. Some implementations of a VAE Setup model may incorporate oral care parameters into the setups predictions. In some implementations, the procedure parameters K and/or the doctor preference information L may be concatenated with the latent space vector C. A doctor’s preferences (e.g., in an orthodontic context ) and/or doctor’s restoration preferences may be indicated in a treatment form, or they could be based upon characteristics in treatment plans such as final setup characteristics (e.g., amount of bite correction or midline correction in planned final setups), intermediate staging characteristics (e.g., treatment duration, tooth movement protocols, or overcorrection strategies), or outcomes (e.g., number of revisions/refinements).
[0048] Orthodontic procedure parameters may specify one or more of the following (with possible values shown in { }). Non-limiting categorical values for some example OPP are described below. In some implementations, a real value may be specified for one or more of these OPP. For example, the Overbite OPP may specify a quantity of overbite (e.g., in millimeters) which is desired in a setup, and may be received as input of a setups prediction model to provide that setups prediction model information about the amount of overbite which is desired in the setup. Some implementations may specify a numerical value for the Oveijet OPP, or other OPP. In some implementations, one or more OPP may be defined which correspond to one or more orthodontic metrics (OM). In some instances, a numerical value may be specified for such an OPP, for the purpose of controlling the output of a setups prediction model.
Teeth To Move: { AnteriorsOnly, AnteriorsAndBicuspids, FullArch}
Tooth Movement Restrictions: for each tooth, indicate if tooth is {DoNotMove, Missing, ToBeExtracted, Primary /Erupting, Clear}
Overbite: {ShowResultingOverbiteAfterAlignment, MaintainlnitialOverbite, CorrectOpenBite, CorrectDeepBite}
Oveijet: {ShowResultingOverjetAfterAlignment, MaintainfnitialOveijet, ImproveResultingOveijet} Anterior/Posterior (AP) Relationship
Maintain: {Right, Left, Both}
Improve canine relationship only: {Right, Left, Both}
Improve canine and/or molar relationship up to 4mm: {Right, Left, Both}
Correct to Class I (canine and molar): {Right, Left, Both}
Crossbite (if present)
Anterior: {DoNotCorrect, Correct, N/A}
Posterior: {DoNotCorrect, Correct, N/A}
Correction to Class I (canine and molar): {Right, Left, Both}
Correct with Posterior IPR: {yes, no}
Class II/III correction simulation (elastics required): {yes, no} Sequential Distalization (elastic recommended): {yes, no} Include cuts for elastics?: {yes, no} Preferred cuts for elastics: {UseButtonCutoutsOnMolarsAndHooksOnCanines, UseButtonCutoutsOnly, UseHooksOnly} Stage to start cuts for elastics: [integer]
LevelingOfUpperAnteriors: {Laterals0.5mmShorterThanCentral, LevellncisalEdges, LevelGingivalMargins, Aslndicated}
Spacing: {Close AllSpaces, LeaveSpecificSpaces}
Preferred Midline Position: {SetTheUpperMidlineToIdeal, MatchTheUpperAndLowerToEachOther} Resolve Upper Crowding by Expand: {Primarily, AsNeeded, None} Resolve Upper Crowding by Procline: {Primarily, AsNeeded, None}
Resolve Upper Crowding by IPR - Anterior: {Primarily, AsNeeded, None}
Resolve Upper Crowding by IPR - Posterior Right: {Primarily, AsNeeded, None}
Resolve Upper Crowding by IPR - Posterior Left: {Primarily, AsNeeded, None}
Resolve Lower Crowding by Expand: {Primarily, AsNeeded, None}
Resolve Lower Crowding by Procline: {Primarily, AsNeeded, None}
Resolve Lower Crowding by IPR - Anterior: {Primarily, AsNeeded, None}
Resolve Lower Crowding by IPR - Posterior Right: {Primarily, AsNeeded, None} Resolve Lower Crowding by IPR - Posterior Left: {Primarily, AsNeeded, None} Finishing Arch Form: {Patient’ sNatural, Aslndicated}
[doctor can specify an archform - selected from a set of options or custom-designed]
[0049] Other orthodontic procedure parameters may be defined, such as those which may be used to place standardized brackets at prescribed occlusal heights on the teeth. In some implementations, one or more orthodontic procedure parameters may be defined to specify at least one of the 2nd and 3 rd order rotation angles to be applied to a tooth (i.e., angulation and torque, respectively), which may enable a target setup arrangement where crown landmarks lie within a threshold distance of a common occlusal plane, for example. In some implementations, one or more orthodontic procedure parameters may be defined to specify the position in global coordinates where at least one landmark (e.g., a centroid) of a tooth crown (or root) is to be placed in a setup arrangement of teeth. Generally, an oral care parameter may be defined which corresponds to an oral care metric. For example, an orthodontic procedure parameter may be defined which corresponds to an orthodontic metric (e.g., to specify at the input of a setups prediction model an amount of a certain metric which is desired to appear in a predicted setup).
[0050] Doctor preferences may differ from orthodontic procedure parameters in that doctor preferences pertain to an oral care provider and may comprise of the means, modes, medians, minimums, or maximums (or some other statistic) of past settings associated with an oral care provider’s treatment decisions on past orthodontic cases. Procedure parameters, on the other hand, may pertain to a specific patient, and describe the needs of a particular patient’s treatment. Doctor preferences may pertain to a doctor and the doctor’s past treatment practices, whereas procedure parameters may pertain to the treatment of a particular patient. Doctor preferences (or “treatment preferences”) may specify one or more of the following (with some non-limiting possible values shown in { }). Other possible values are found elsewhere in this disclosure.
[0051] Doctor preferences may specify one or more of the following (with other possible values found elsewhere in this disclosure).
Deep Bite Cases (Amount of Bite Correction) - Final Overbite: [real value in millimeters, e.g., 0.5 mm] Option - Intrude Upper Anteriors: {yes, no} Option - Include lower canines in vertical overcorrection: {yes, no}
Midline Correction in Planned Final Setup: {MaintainlnitialMidline, ImproveMidlineWithfPR, Aslndicated}
Deep Bite Cases - Reverse Curve of Speed: {yes, no}
Anterior Open Bite Cases - Final Overbite: [real value in millimeters, e.g., 2 mm]
Is Arch Expansion a Priority for Your Cases?: {Yes, No}
If yes, specify acceptable expansion per quadrant in mm.
When expanding upper molars, apply buccal root torque: {yes, no}
Is IPR Acceptable of First Tx Design: {yes, no} Maximum IPR per contact:
Upper Anterior: [specify in mm] Lower Anterior: [specify in mm] Upper and Lower Anterior: [specify in mm] is Asymmetric iPR Acceptable?: {yes, no} Final Tooth Position (Overcorrection Strategy): {Ideal, Overcorrected} Root Movement: {MoveRootsAsNeededToAchieveTreatmentGoals, LimitPosteriorRootMovement, LimitAllRootMovement}
Final Occlusal Contacts: {AllContactsBalancedWhenPossible, NoOcclusalContactOnUpperlncisors, FinishWithHeavyPosteriorContacts, Other}
Is Asymmetric AP Shift Acceptable for Class Correction?: {yes, no, other} Treatment Duration: [count of stages]
Tooth Movement Protocol: {protocol A, protocol B, protocol C}
[0052] Existing techniques have attempted to move the teeth towards an archform V after a setup prediction has already been rendered by other method components, which introduces error to the resulting setup and diminishes performance in view of the purpose of the method components which made the setups prediction. The present disclosure provides several improvements over these existing techniques by enabling archform information to be introduced directly into the setups prediction neural network as an input to that neural network, with the technical improvement of providing setups predictions that more accurately meets the orthodontic treatment needs of the patient (thereby improving data precision). Archform information V may be provided as an input to any of the GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups and Diffusion Setups prediction neural networks. In some implementations, archform information V may be provided directly to one or more internal neural network layers in one or more of those setups applications.
[0053] The additional procedure parameters may include text descriptions of the patient’s medical condition and of the intended treatment. Such text descriptions may be analyzed via natural language processing operations, including tokenization, stop word removal, stemming, n-gram formation, text data vectorization, bag of words analysis, term frequency inverse document frequency (TF-IDF) analysis, sentiment analysis, naive Bayes classification, and/or logistic regression classification. The outputs of such analysis techniques may be used as input to one or more of the neural networks of this disclosure with the advantage of customizing and improving the predicted outputs (e.g., the predicted setups or predicted mesh geometries).
[0054] These additional orthodontic parameters and doctor preferences may also be incorporated into the neural networks of this disclosure, with the data precision and efficiency -related advantages of improving the predictive power of those neural networks and enabling those neural networks to predict output examples which better match the treatment needs of the individual patient.
[0055] In some implementations, a dataset used for training one or more of the neural network models of this disclosure may be filtered conditionally on one or more of the orthodontic procedure parameters described in this section. In some instances, patient cases which exhibit outlier values for one or more of these procedure parameters may be omitted from a dataset (alternatively used to form a dataset) for training one or more of the neural networks of this disclosure.
[0056] One or more procedure parameters and/or doctor preferences may be provided to a neural network during training. In this manner the neural network may be conditioned on the one or more procedure parameters and/or doctor preferences. Examples of such neural networks include a conditional generative adversarial network (cGAN) and/or a conditional variational autoencoder (cVAE), either of which may be used for the various neural network-based applications of this disclosure.
[0057] In some instances, tooth shape-based inputs may be provided to a neural network for setups predictions. In other instances, non-shape-based inputs can be used, such as a tooth name or designation, as it pertains to dental notation. In some implementations, a vector R of flags may be provided to the neural network, where a ‘ 1 ’ value indicates that the tooth is present and a ‘0’ value indicates that the tooth is absent from the patient case (though other values are possible). The vector R may comprise a 1- hot vector, where each element in the vector corresponds to a tooth type, name or designation.
Identifying information about a tooth (e.g., the tooth’s name) can be provided to the predictive neural networks of this disclosure, with the advantage of enabling the neural network to become trained to handle different teeth in tooth-specific ways. For example, the setups prediction model may learn to make setups transformations predictions for a specific tooth designation (e.g., upper right central incisor, or lower left cuspid, etc.). In the case of the mesh cleanup autoencoders (either for labelling mesh element or for in-filling missing mesh data), the autoencoder may be trained to provide specialized treatment to a tooth according to that tooth’s designation, in this manner. In the case of a setups classification neural network, a listing of tooth name(s) present in the patient’s arch may better enable the neural network to output an accurate determination of setup classification, because tooth designation is a valuable input to training such a neural network. Tooth designation/name may be defined, for example, according to the Universal Numbering System, Palmer System, or the FDI World Dental Federation notation (ISO 3950).
[0058] In one example, where all except the (up to four) wisdom teeth are present in the case, a vector R may be defined as an optional input to the setups prediction neural networks of this disclosure, where there is a 0 in the vector element corresponding to each of the wisdom teeth, and a 1 in the elements corresponding to the following teeth: UR7, UR6, UR5, UR4, UR3, UR2, UR1, ULI, UL2, UL3, UL4, UL5, UL6, UL7, LL7, LL6, LL5, LL4, LL3, LL2, LL1, LR1, LR2, LR3, LR4, LR5, LR6, LR7 [0059] In some instances, the position of the tooth tip may be provided to a neural network for setups predictions. In other instances, one or more vectors S of the orthodontic metrics described elsewhere in this disclosure may be provided to a neural network for setups predictions. The advantage is an improved capacity for the network to become trained to understand the state of a maloccluded setup and therefore be able to predict a more accurate final setup or intermediate stage.
[0060] In some implementations, the neural networks may take as input one or more indications of interproximal reduction (IPR) U, which may indicate the amount of enamel that is to be removed from a tooth during the course orthodontic treatment (either mesially or distally). In some implementations, IPR information (e.g., quantity of IPR that is to be performed on one or more teeth, as measured in millimeters, or one or more binary flags to indicate whether or not IPR is to be performed on each tooth identified by flagging) may be concatenated with a latent vector A which is produced by a VAE or a latent capsule T autoencoder. The vector(s) and/or capsule(s) resulting from such a concatenation may be provided to one or more of the neural networks of the present disclosure, with the technical improvement or added advantage of enabling that predictive neural network to account for IPR. IPR is especially relevant to setups prediction methods, which may determine the positions and poses of teeth at the end of treatment or during one or more stages during treatment. It is important to account for the amount of enamel that is to be removed ahead of predicted tooth movements.
[0061] In some implementations, one or more procedure parameters K and/or doctor preferences vectors L may be introduced to a setups prediction model. In some implementations, one or more optional vectors or values of tooth position N (e.g., XYZ coordinates, in either tooth local or global coordinates), tooth orientation O (e.g., pose, such as in transformation matrices or quaternions, Euler angles or other forms described herein), dimensions of teeth P (e.g., length, width, height, circumference, diameter, diagonal measure, volume - any of which dimensions may be normalized in comparison to another tooth or teeth), distance between adjacent teeth Q. These “dimensions of teeth P” may in some instances be used to describe the intended dimensions of a tooth for dental restoration design generation. [0062] In some implementations, tooth dimensions P such as length, width, height, or circumference may be measured inside a plane, such as the plane that intersects the centroid of the tooth, or the plane that intersects a center point that is located midway between the centroid and either the incisal-most extent or the gingival-most extent of the tooth. The tooth dimension of height may be measured as the distance from gums to incisal edge. The tooth dimension of width may be measured as the distance from the mesial extent to the distal extent of the tooth. In some implementations, the circularity or roundness of the tooth cross-section may be measured and included in the vector P. Circularity or roundness may be defined as the ratio of the radii of inscribed and circumscribed circles.
[0063] The distance Q between adjacent teeth can be implemented in different ways (and computed using different distance definitions, such as Euclidean or geodesic). In some implementations, a distance QI may be measured as an averaged distance between the mesh elements of two adjacent teeth. In some implementations, a distance Q2 may be measured as the distance between the centers or centroids of two adjacent teeth. In some implementations, a distance Q3 may be measured between the mesh elements of closest approach between two adjacent teeth. In some implementations, a distance Q4 may be measured between the cusp tips of two adjacent teeth. Teeth may, in some implementations, be considered adjacent within an arch. Teeth may, in some implementations, also be considered adjacent between opposing arches. In some implementations, any of QI, Q2, Q3 and Q4 may be divided by a term for the purpose of normalizing the resulting value of Q. In some implementations, the normalizing term may involve one or more of: the volume of a tooth, the count of mesh elements in a tooth, the surface area of a tooth, the cross-sectional area of a tooth (e.g., as projected into the XY plane), or some other term related to tooth size.
[0064] Other information about the patient’s dentition or treatment needs (or related parameters) may be concatenated with the other input vectors to one or more of MLP, GAN, generator, encoder structure, decoder structure, transformer, VAE, conditional VAE, regularized VAE, 3D U-Net, capsule autoencoder, diffusion model, and/or any of the neural networks models listed elsewhere in this disclosure.
[0065] The vector M may contain flags which apply to one or more teeth. In some implementations, M contains at least one flag for each tooth to indicate whether the tooth is pinned. In some implementations, M contains at least one flag for each tooth to indicate whether the tooth is fixed. In some implementations, M contains at least one flag for each tooth to indicate whether the tooth is pontic. Other and additional flags are possible for teeth, as are combinations of fixed, pinned and pontic flags. A flag that is set to a value that indicates that a tooth should be fixed is a signal to the network that the tooth should not move over the course of treatment. In some implementations, the neural network loss function may be designed to be penalized for any movement in the indicated teeth (and in some particular cases, may be heavily penalized). A flag to indicate that a tooth is pontic informs the network that the tooth gap is to be maintained, although that gap is allowed to move. In some cases, M may contain a flag indicating that a tooth is missing. In some implementations, the presence of one or more fixed teeth in an arch may aid in setups prediction, because the one or more fixed teeth may provide an anchor for the poses of the other teeth in the arch (i.e., may provide a fixed reference for the pose transformations of one or more of the other teeth in the arch). In some implementations, one or more teeth may be intentionally fixed, so as to provide an anchor against which the other teeth may be positioned. In some implementations, a 3D representation (such as a mesh) which corresponds to the gums may be introduced, to provide a reference point against which teeth can be moved.
[0066] Without the loss of generality, one or more of the optional input vectors K, L, M, N, O, P, Q, R, S, U and V described elsewhere in this disclosure may also be provided to the input or into an intermediate layer of one or more of the predictive models of this disclosure. In particular, these optional vectors may be provided to the MLP Setups, GDL Setups, RL Setups, VAE Setups, Capsule Setups and/or Diffusion Setups, with the advantage of enabling the respective model to generate setups which better meet the orthodontic treatment needs of the patient. In some implementations, such inputs may be provided, for example, by being concatenated with one or more latent vectors A which are also provided to one or more of the predictive models of this disclosure. In some implementations, such inputs may be introduced, for example, by being concatenated with one or more latent capsules T which are also provided to one or more of the predictive models of this disclosure.
[0067] In some implementations, one or more of K, L, M, N, O, P, Q, R, S, U and V may be introduced to the neural network (e.g., MLP or Transformer) directly in a hidden layer of the network. In some instances, one or more of K, L, M, N, O, P, Q, R, S, U and V may be introduced directly into the internal processing of an encoder structure.
[0068] In some implementations, a setups prediction model (such as GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, PT Setups, Similarity Setups and Diffusion Setups) may take as input one or more latent vectors A which correspond to one or more input oral care meshes (e.g., such as tooth meshes). In some implementations, a setups prediction model (such as GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups and Diffusion Setups) may take as input one or more latent capsules T which correspond to one or more input oral care meshes (e.g., such as tooth meshes). In some implementations, a setups prediction method may take as input both of A and T.
[0069] Some implementations of the setups prediction neural networks (e.g., GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, or FDG Setup, or other setups prediction network architectures) of this disclosure may take additional inputs to aid in setups prediction. Some of these inputs may reflect the geometrical attributes of one or more teeth or of a whole arch. In some implementations, an archform or arch curve may be provided to a setups prediction neural network, with the technical improvement of aiding that setups prediction neural network in finding a suitable set of final setups poses for the teeth in a patient case (with the technical improvements being directed to both resource footprint reduction by way of more efficient location capabilities and/or data precision in the form of locating a more pertinent final setup). The archform or arch curve may be encoded as a spline, a B-spline, NonUniform Rational B-Splines (NURBS), polynomial spline, non-polynomial spline, parabolic curve, hyperbolic curve or other parameterized curve. Such a curve may be computed as an average of multiple exemplars, such as exemplary final setups. Another non-limiting example of an archform is a Beta curve. In the case of the Setups VAE, the arch information may be provided to the encoder E2, as an additional input alongside E and D. In the case of the GDL Setups neural network, the arch information may be provided to the generator as an additional input to the mesh element lists and associated mesh element feature vectors. In some implementations, an archform may be described by one or more 3D representations, such as a 3D mesh, a set of 3D control points and/or as a 3D polyline. In some implementations, a Frenet frame may be overlaid onto an archform. The Frenet frame may locally describe the coordinate system corresponding to each point along the archform. Such a coordinate system may, in some implementations, be right-handed (or alternatively, in other implementations, left-handed). Such a coordinate system may, in some implementations, be determined, at least in part, by at least one of the tangent to the archform at the point and the archform’s curvature. In some implementations, a point may be described using an LDE coordinate frame relative to an archform, where L, D and E correspond to: 1) Length along the curve of the archform, 2) Distance away from the archform, and 3) Distance in the direction perpendicular to the L and D axes (which may be termed Eminence), respectively. Other geometrical inputs may also aid in the training of a setups prediction neural network.
[0070] Various loss calculation techniques are generally applicable to the techniques of this disclosure (e.g., GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, Setups Classification, Tooth Classification, VAE Mesh Element Labelling, MAE Mesh In-Filling and the imputation of procedure parameters).
[0071] These losses include LI loss, L2 loss, mean squared error (MSE) loss, cross entropy loss, among others. Losses may be computed and used in the training of neural networks, such as multi-layer perceptron’s (MLP), U-Net structures, generators and discriminators (e.g., for GANs), autoencoders, variational autoencoders, regularized autoencoders, masked autoencoders, transformer structures, or the like. Some implementations may use either triplet loss or contrastive loss, for example, in the learning of sequences.
[0072] Losses may also be used to train encoder structures and decoder structures. A KL- Divergence loss may be used, at least in part, to train one or more of the neural networks of the present disclosure, such as a mesh reconstruction autoencoder or the generator of GDL Setups, which the advantage of imparting Gaussian behavior to the optimization space. This Gaussian behavior may enable a reconstruction autoencoder to produce a better reconstruction (e.g., when a latent vector representation is modified and that modified latent vector is reconstructed using a decoder, the resulting reconstruction is more likely to be a valid instance of the inputted representation). There are other techniques for computing losses which may be described elsewhere in this disclosure. Such losses may be based on quantifying the difference between two or more 3D representations.
[0073] MSE loss calculation may involve the calculation of an average squared distance between two sets, vectors or datasets. MSE may be generally minimized. MSE may be applicable to a regression problem, where the prediction generated by the neural network or other machine learning model may be a real number. In some implementations, a neural network may be equipped with one or more linear activation units on the output to generate an MSE prediction. Mean absolute error (MAE) loss and mean absolute percentage error (MAPE) loss can also be used in accordance with the techniques of this disclosure. [0074] Cross entropy may, in some implementations, be used to quantify the difference between two or more distributions. Cross entropy loss may, in some implementations, be used to train the neural networks of the present disclosure. Cross entropy loss may, in some implementations, involve comparing a predicted probability to a ground truth probability. Other names of cross entropy loss include “logarithmic loss,” “logistic loss,” and “log loss”. A small cross entropy loss may indicate a better (e.g., more accurate) model. Cross entropy loss may be logarithmic. Cross entropy loss may, in some implementations, be applied to binary classification problems. In some implementations, a neural network may be equipped with a sigmoid activation unit at the output to generate a probability prediction. In the case of multi-class classifications, cross entropy may also be used. In such a case, a neural network trained to make multi-class predictions may, in some implementations, be equipped with one or more softmax activation functions at the output (e.g., where there is one output node for class that is to be predicted). Other loss calculation techniques which may be applied in the training of the neural networks of this disclosure include one or more of: Huber loss, Hinge loss, Categorical hinge loss, cosine similarity, Poisson loss, Logcosh loss, or mean squared logarithmic error loss (MSLE). Other loss calculation methods are described herein and may be applied to the training of any of the neural networks described in the present disclosure.
[0075] One or more of the neural networks of the present disclosure may, in some implementations, be trained, at least in part by a loss which is based on at least one of: a Point-wise Mesh Euclidean Distance (PMD) and an Earth Mover’s Distance (EMD). Some implementations may incorporate a Hausdorff Distance (HD) calculation into the loss calculation. Computing the Hausdorff distance between two or more 3D representations (such as 3D meshes) may provide one or more technical improvements, in that the HD not only accounts for the distances between two meshes, but also accounts for the way that those meshes are oriented, and the relationship between the mesh shapes in those orientations (or positions or poses). Hausdorff distance may improve the comparison of two or more tooth meshes, such as two or more instances of a tooth mesh which are in different poses (e.g., such as the comparison of predicted setup to ground truth setup which may be performed in the course of computing a loss value for training a setups prediction neural network).
[0076] Reconstruction loss may compare a predicted output to a ground truth (or reference) output. Systems of this disclosure may compute reconstruction loss as a combination of LI loss and MSE loss, as shown in the following line of pseudocode: reconstruction loss = 0.5*Ll(all_points_target,all_points_predicted) + 0.5*MSE(all_points_target,all_points_predicted). In the above example, all_points_target is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to ground tmth data (e.g., a ground truth tooth restoration design, or a ground tmth example of some other 3D oral care representation). In the above example, all_points_predicted is a 3D representation (e.g., a 3D mesh or point cloud) corresponding to generated or predicted data (e.g., a generated tooth restoration design, or a generated example of some other kind of 3D oral care representation). Other implementations of reconstruction loss may additionally (or alternatively) involve L2 loss, mean absolute error (MAE) loss or Huber loss terms. [0077] The entirety of the following paper is incorporated herein by reference in its entirety: "Attention Is All You Need"; Ashish Vaswani, Noam Shazeer, Niki Parmar, Niki Parmar, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin; NIPS 2017. The neural network-based models of this disclosure may provide additional advantages in implementations in which they are integrated with a neural network structure referred to as a “transformer.” FIG. 7 shows an example implementation of a transformer architecture.
[0078] Before recently developed models such as the transformer model, RNN-type models represented the state of the art for natural language processing (NLP). One example application of NLP is the generation of new text based upon prior words or text. Transformers have in turn provided significant improvements over GRU, LSTM and other such RNN-based NLP techniques due to an important attribute of the transformer model, which has the property of multi-headed attention. In some implementations, the NLP concept of multi-headed attention may describe the relationship between each word in a sentence (or paragraph or document or corpus of documents) and each other word in that sentence (or paragraph or document or corpus of documents). These relationships may be generated by a multiheaded attention module, and may be encoded in vector form. This vector may describe how each word in a sentence (or paragraph or document or corpus of documents) should attend to each other word in that sentence (or paragraph or document or corpus of documents). RNN, LSTM and GRU models process a sequence, such a sentence, one word at a time from the start to the end of the sequence. Furthermore, the model may only account for a given subset (called a window) of the sentence when making a prediction. However, transformer-based models may, in some instances, account for the entirety of the preceding text by processing the sequence in its entirety in a single step. Transformer, RNN, LSTM, and GRU models can all be adapted for use in predictive models in digital dentistry and digital orthodontics, particularly for the setup prediction task. In some implementations, an exemplary transformer model for use with 3D meshes and 3D transforms in setups prediction (or other oral care techniques) may be adapted from the Bidirectional Encoder Representation from Transformers (BERT) and/or Generative Pre-Training (GPT) models. For example, a GPT (or BERT) model may first be trained on other data, such as text or documents data, and then be used in transfer learning. Such a transfer learning process may receive a previously trained GPT or BERT model, and then do further training using data comprising 3D oral care representations. Such transfer learning may be performed to train oral care models such as: segmentation, mesh cleanup, coordinate system prediction, setups prediction, validation of 3D oral care representations, transform prediction for placement of oral care meshes (e.g., teeth, hardware, appliance components, fixture model components), tooth restoration design generation (or generation of other 3D oral care representations - such as appliance components, fixture models or archforms), classification of 3D oral care representations, imputation of missing oral care parameters, clustering of clinicians or clustering of clinician preferences, or the like.
[0079] Oral care data may comprise one or more of (or combinations of): 3D representations of tooth (e.g., meshes, point clouds or voxels), sections of tooth meshes (such as subsets of mesh elements), tooth transforms (such as in matrix, vector and/or quaternion form, or combinations thereof), transforms for appliance components, transforms for fixture model components, and mesh coordinate system definitions (such as represented by transforms, for example, transformation matrices) and/or other 3D oral care representations described herein.
[0080] Transformers may be trained for generating transforms to position teeth into setups poses (or to place appliance components for use in appliance generation or to place fixture model components for use in fixture model generation). Some implementations may operate in an offline prediction context, and some implementations operation in an online reinforcement learning (RL) context. In some implementations, a transformer may be initially trained in an offline context and then undergo further fine-tuning training in the online context. In the offline prediction context, the transformer may be trained from a dataset of cohort patient case data. In the online RL context, the transformer may be trained from either a physics model, or a CAD model, for example. The transformer may learn from static data, such as transformations (e.g., trajectory transformer). In some implementations, the transform may provide a mapping from malocclusion to setup (e.g., receiving transformation matrices as input and generating transformation matrices as ouput). Some implementations of transformers may be trained to process 3D representations, such as 3D meshes, 3D point clouds or voxels (e.g., using a decision transformer) takes as input geometry (e.g., mesh, point cloud, voxels etc.), outputs transformations. The decision transformer may be coupled with a representation generation module that encodes representation of the patient’s dentition (e.g., teeth), such as a VAE, a U-Net, an encoder, a transformer encoder, a pyramid encoder-decoder or a simple dense or fully connected network, or a combination thereof. In some implementations, the representation generation module (e.g., VAE, the U-Net, the encoder, the pyramid encoder-decoder or the dense network for generating the tooth representation) may be trained to generate the representation on one or more teeth. The representation generation module may be trained on all teeth in both arches, only the teeth within the same arch (either upper or lower), only anterior teeth, only posterior teeth, or some other subset of teeth. In some implementations, such a model may be trained on each individual tooth (e.g., an upper right cuspid), so that the model is trained or otherwise configured togenerate highly accurate representations for an individual tooth. In some implementations, an encoder structure may encode such a representation. In some implementations, a decision transformer may learn in an online context, in an offline context or both. An online decision transformer may be trained (e.g., using RL techniques) to output action, state, and/or reward. In some implementations, transformations may be discretized, to allow for piecewise or stepwise actions.
[0081] In some implementations, a transformer may be trained to process an embedding of the arch (i.e., to predict transforms for multiple teeth concurrently), to predict a setup. In some implementations, embeddings of individual teeth may be concatenated into a sequence, and then input into the transformer. A VAE may be trained to perform this embedding operation, a U-Net may be trained to perform such an embedding, or a simple dense or fully connected network may be trained, or a combination thereof. In some implementations, the transformer-based techniques of this disclosure may predict an action for an individual tooth, or may predict actions for multiple teeth (e.g., predict transformations for each of multiple teeth). [0082] A 3D mesh transformer may include a transformer encoder structure (which may encode oral care data), and may be followed by a transformer decoder structure. The 3D mesh transformer encoder may encode oral care data into a latent representation, which may be combined with attention information (e.g., to concatenate a vector of attention information to the latent representation). In some implementations, the attention information may help the decoder focus on the relevant oral care data during the decoding process (e.g., to focus on tooth order or mesh element connectivity), so that the transformer decoder can generate a useful output for the 3D mesh transformer (e.g., an output which may be used in the generation of an oral care appliance). Either or both of the transformer encoder or transformer decoder may generate a latent representation. The output of the transformer decoder (or transformer encoder) may be reconstructed using a decoder into, for example, one or more tooth transforms for a setup, one or more mesh element labels for segmentation, coordinate systems transforms for use in coordinate system generation, or one or more points of a point cloud or voxels or other mesh elements for another 3D representation). A transformer may include modules such as one or more of: multi-headed attention modules, feed forward modules, normalization modules, linear modules, and softmax modules, and convolution models for latent vector compression, and/or representation.
[0083] The encoder may be stacked one or more times, thereby further encoding the oral care data, and enabling different representations of the oral care data to be learned (e.g., different latent representations). These representations may be embedded with attention information (which may influence the decoder’s focus to the relevant portions of the latent representation of the oral care data) and may be provided to the decoder in continuous form (e.g., as a concatenation of latent representations - such as latent vectors). In some implementations, the encoded output of the encoder (e.g., latent representations) may be used by downstream processing steps in the generation of oral care appliances. For example, the generated latent representation may be reconstructed into transforms (e.g., for the placement of teeth in setups, or the placement of appliance components or fixture model components), or may be reconstructed into 3D representations (e.g., 3D point clouds, 3D meshes or others disclosed herein). Stated another way, the latent representation which is generated by the transformer (e.g., containing continuously encoded attention information) may be provided to a decoder which has been configured to reconstruct the latent representation into the specific data structure which is required by a particular domain area. Continuously encoded attention information may include attention information which has undergone processing by multiple multi-headed attention modules within the transformer encoder or transformer decoder, to name one example. Furthermore, a loss may be computed for a particular domain using data from that domain. The loss calculation may train the transformer decoder to accurately reconstruct the latent representation into the output data structure pertaining to a particular domain.
[0084] For example, when the decoder generates a transform for an orthodontic setup, the decoder may be configured with outputs that describe, for example, the 16 real values which comprise a 4x4 transformation matrix (other data structures for describing transforms are possible). Stated a different way, the latent output generated by the transformer encoder (or transformer decoder) may be used to predict setups tooth transforms for one or more teeth, to place those teeth in setup positions (e.g., either final setups or intermediate stages). Such a transformer encoder (or transformer decoder) may be trained, at least in part using a reconstruction loss (or a representation loss, among others described herein) function, which may compare predicted transforms to ground truth (or reference) transforms.
[0085] In a further example, when the decoder generates a transform for a tooth coordinate system, the decoder may be configured with outputs that describe, for example, the 16 real values which comprise a 4x4 transformation matrix (other data structures for describing transforms are possible). Stated a different way, the latent output generated by the transformer encoder (or transformer decoder) may be used to predict local coordinate systems for one or more teeth. Such a transformer encoder (or transformer decoder) may be trained, at least in part using a representation loss (or a reconstruction loss, among others described herein) function, which may compare predicted coordinate systems to ground truth (or reference) coordinate systems.
[0086] In a further example, when the decoder generates a 3D point cloud (or other 3D representation - such as 3D mesh, voxelized representation, or the like), the decoder may be configured with outputs that describe, for example, one or more 3D points (e.g., comprising XYZ coordinates). Stated a different way, the latent output generated by the transformer encoder (or transformer decoder) may be used to predict mesh elements for a generated (or modified) 3D representation. Such a transformer encoder (or transformer decoder) may be trained, at least in part using a reconstruction loss (or an LI, L2 or MSE loss, among others described herein) function, which may compare predicted 3D representations to ground truth (or reference) 3D representations.
[0087] In a further example, when the decoder generates mesh element labels for 3D representation segmentation or 3D representation cleanup, the decoder may be configured with outputs that describe, for example, labels for one or more mesh elements. Stated a different way, the latent output generated by the transformer encoder (or transformer decoder) may be used to predict mesh element labels for mesh segmentation or mesh cleanup. Such a transformer encoder (or transformer decoder) may be trained, at least in part using a cross entropy loss (or others described herein) function, which may compare predicted mesh element labels to ground truth (or reference) mesh element labels.
[0088] Multi-headed attention and transformers may be advantageously applied to the setups- generation problem. Multi-headed attention is a module in a 3D transformer encoder network which computes the attention weights for the provided oral care data and produces an output vector with encoded information on how each example of oral care data should attend to each other oral care data in an arch. An attention weight is a quantification of the relationship between pairs of oral care data.
[0089] A 3D representation of oral care data (e.g., comprising voxels, a point cloud, or a 3D mesh composed of vertices, faces or edges) may be provided to the transformer. The 3D representation may describe the patient's dentition, a fixture model (or components of a fixture model), an appliance (or components of an appliance), or the like. In some implementations, a transformer decoder (or a transformer encoder) may be equipped with multi-head attention. Multi -headed attention may enable the transformer decoder (or transformer encoder) to attend to different portions of the 3D representation of oral care data. For example, multi-headed attention may enable the transformer to attend to mesh elements within local neighborhoods (or cliques), or to attend to global dependencies between mesh elements (or cliques). For example, multi-headed attention may enable a transformer for setups prediction (e.g., a setups prediction model which is based on a transformer) to generate a transform for a tooth, and to substantially concurrently attend to each of the other teeth in the arch while that transform is generated. Stated another way, the transform for each tooth may be generated in light of the poses of one or more other teeth in the arch, leading to a more accurate transform (e.g., a transform which conforms more closely to the ground truth or reference transform). In the example of 3D representation generation (e.g., the generation of a 3D point cloud), a transformer model may be trained to generate a tooth restoration design. Multi-headed attention may enable the transformer to attend to multiple portions of the tooth (or to the surfaces of the adjacent teeth) while the tooth undergoes the generative process. For example, the transformer for restoration design generation may generate the mesh elements for the incisal edge of an incisor while, at least substantially concurrently, attending to the mesh elements of the mesial, distal, facial or lingual surfaces of the incisor. The result may be the generation of mesh elements to form an incisal edge for the tooth which merges seamlessly with the adjacent surfaces of the tooth. This use of multi-headed attentions results in more accurate modeling of the distribution of the training dataset, over techniques which do not apply multi-headed attention.
[0090] In some implementations of the present disclosure, one or more attention vectors may be generated which describe how aspects of the oral care data interacts with other aspects of the oral care data associated with the arch. In some implementations, the one or more attention vectors may be generated to describe how one or more portions of a tooth T1 interact with one or more portions of a tooth T2, a tooth T3, a tooth T4, and so one. A portion of a mesh may be described as a set of mesh elements, as defined herein. In some implementations, the interacting portions of tooth T1 and tooth T2 may be determined, in part, through the calculation of mesh correspondences, as described herein. Any of these models (RNN, GRU, LSTM and Transformer) may be advantageously applied to the task of setups transform prediction, such as in the models described herein. A transformer may be particularly advantageous in that a transformer may enable the transforms for multiple teeth, or even an entire arch to be generated at once, rather than individually, as may be the case with some other models, such as an encoder structure. In other implementations, attention-free transformers may be used to make predictions based on oral care data.
[0091] One implementation of the GDL Setups neural network model may include a representation generation module (e.g., containing a U-Net structure, an autoencoder encoder, a transformer encoder, another type of encoder-decoder structure, or an encoder, etc.) which may provide its output to a module which is trained to generate tooth transformers (e.g., a set of fully connected layers with optional skip connections, or an encoder structure) to generate the prediction of a transform for each individual tooth. Skip connections may, in some implementations, connect the outputs of a particular layer in a neural network to the inputs of another later in the neural network (e.g., a layer which is not immediately adjacent to the originating layer). The transform-generation module (e.g., an encoder) may handle the transform prediction one tooth at a time. Other implementations may replace this encoder structure with a transformer (e.g., transformer encoder or transformer decoder), which may handle all the predictions for all teeth substantially concurrently. Stated another way, a transformer may be configured to receive a large number of input values, larger than some other neural network models (e.g., than a typical MLP). This is because an increased number of inputs may be accommodated by the transformer, the predictions corresponding to those inputs may be generated substantially concurrently. The representation generation module (e.g., U-Net structure) may provide its output to the transformer, and the transformer may generate the setups transforms for all of the several teeth at once, with the technical advantage of improved accuracy (because the transforms for each tooth is generated in light of the transform for each of the adjacent or nearby teeth - leading to fewer collisions and better conformance with the goals of treatment). A transformer may be trained to output a transformation, such as a transform encoded by a 4x4 matrix (or some other size), a quaternion, a translation vector, Euler angles or some other form. The transformation may place a tooth into a setups pose, may place a fixture model component into a pose suitable for fixture model generation, or may place an appliance component into a pose suitable for appliance generation (e.g., dental restoration appliance, clear tray aligner, etc.). In some implementations, the transform may define a coordinate system for aspects of the patient’s dentition, such as a tooth mesh (e.g., a local coordinate system for a tooth). In some implementations, the inputs to the transformer may first be encoded using a neural network (e.g., a latent representation or embedding may be generated), such as one or more linear layers, and/or one or more convolutional layers. In some implementations, the transformer may first be trained on an offline dataset, and subsequently be trained using a secondary actor-critic network, which may enable online reinforcement learning.
[0092] Transformers may, in some implementations, enable large model capacity and/or enable an attention mechanism (e.g., the capability to pay attention and respond to certain inputs). The attention mechanisms (e.g., multi-headed attention) that are found within transformers may enable intra-sequence relationships to be encoded into neural network features. Intra-sequence relationships may be encoded, for example, by associating an order number (e.g., 1, 2, 3, etc.) with each tooth in an arch, or by associating an order number with each mesh element in a 3D representation (e.g., of a tooth). In implementations where latent vectors of teeth are provided to the transformer, intra-sequence relationships may be encoded, for example, by associating an order number (e.g., 1, 2, 3, etc.) with each element in the latent vector.
[0093] Transformers may be scaled by increasing the number of attention heads and/or by increasing the number of transformer layers. Stated differently, one or more aspects of a transformer may be independently trained to handle discrete tasks, and later combined to allow the resulting transformer to perform all of the tasks for which the individual components had been trained, without degrading the predictive accuracy of the neural network. Scaling a convolutional network may be more difficult, because the models may be less malleable or may be less interchangeable.
[0094] Convolution has an ability to be rotation and translation invariant, which leads to improved generalization, because a convolution model may not need to account for the manner in which the input data in rotated or translated. Transformers have an ability to be permutation invariant, since intrasequence relationships may be encoded into neural network features.
[0095] In some implementations for the generation or modification of 3D oral care representations, transformers may be combined with convolution-based neural networks, such as by vertically stacking convolution layers and attention layers. Stacking transformer blocks with convolutional blocks enables the resulting structure to have the translation invariance of convolution, and also the permutation invariance of a transformer. Such stacking may improve model capacity and/or model generalization. CoAtNet is an example of a network architecture which combines convolutional and attention-based elements and may be applied to the processing of oral care data. In some instances, a network for the modification or generation of 3D oral care representations may be trained, at least in part, from CoAtNet (or another model that combines convolution and self-attention/transformers) using transfer learning. [0096] The techniques of this disclosure may include operations such as 3D convolution, 3D pooling, 3D unconvolution and 3D unpooling. 3D convolution may aid segmentation processing, for example in down sampling a 3D mesh. 3D un-convolution undoes 3D convolution, for example, in a U- Net. 3D pooling may aid the segmentation processing, for example in summarized neural network feature maps. 3D un-pooling undoes 3D pooling, for example in a U-Net. These operations may be implemented by way of one or more layers in the predictive or generative neural networks described herein. These operations may be applied directly on mesh elements, such as mesh edges or mesh faces. These operations provide for technical improvements over other approaches because the operations are invariant to mesh rotation, scale, and translation changes. In general, these operations depend on edge (or face) connectivity, therefore these operations remain invariant to mesh changes in 3D space as long as edge (or face) connectivity is preserved. That is, the operations may be applied to an oral care mesh and produce the same output regardless of the orientation, position or scale of that oral care mesh, which may lead to data precision improvement. MeshCNN is a general-purpose deep neural network library for 3D triangular meshes, which can be used for tasks such as 3D shape classification or mesh element labelling (e.g., for segmentation or mesh cleanup). MeshCNN implements these operations on mesh edges. Other toolkits and implementations may operate on edges or faces.
[0097] In some implementations of the techniques of this disclosure, neural networks may be trained to operate on 2D representations (such as images). In some implementations of the techniques of this disclosure, neural networks may be trained to operate on 3D representations (such as meshes or point clouds). An intraoral scanner may capture 2D images of the patient's dentition from various views. An intraoral scanner may also (or alternatively) capture 3D mesh or 3D point cloud data which describes the patient's dentition. According to various techniques, autoencoders (or other neural networks described herein) may be trained to operate on either or both of 2D representations and 3D representations.
[0098] A 2D autoencoder (comprising a 2D encoder and a 2D decoder) may be trained on 2D image data to encode an input 2D image into a latent form (such as a latent vector or a latent capsule) using the 2D encoder, and then reconstruct a facsimile of the input 2D image using the 2D decoder. In the case of a handheld mobile app which has been developed for such analysis (e.g., for the analysis of dental anatomy), 2D images may be readily captured using one or more of the onboard cameras. In other examples, 2D images may be captured using an intraoral scanner which is configmed for such a function. Among the operations which may be used in the implementation a 2D autoencoder (or other 2D neural network) for 2D image analysis are 2D convolution, 2D pooling and 2D reconstruction error calculation. 2D convolution:
[0099] 2D image convolution may involve the "sliding" of a kernel across a 2D image and the calculation of elementwise multiplications and the summing of those elementwise multiplications into an output pixel. The output pixel that results from each new position of the kernel is saved into an output 2D feature matrix. In some implementations, neighboring elements (e.g., pixels) may be in well-defined locations (e.g., above, below, left and right) in a rectilinear grid.
2D pooling:
[00100] A 2D pooling layer may be used to down sample a feature map and summarize the presence of certain features in that feature map.
[00101] 2D reconstruction error may be computed between the pixels of the input and reconstmcted images. The mapping between pixels may be well understood (e.g., the upper pixel [23, 134] of the input image is directly compared to pixel [23,134] of the reconstructed image, assuming both images have the same dimensions). Among the advantages provided by the 2D autoencoder-based techniques of this disclosure is the ease of capturing 2D image data with a handheld device. In some instances, where outside data sources provide the data for analysis, there may be instances where only 2D image data are available. When only 2D image data are available, then analysis using a 2D autoencoder.
[00102] Modem mobile devices (such as commercially available smartphones) may also have the capability of generating 3D data (e.g., using multiple cameras and stereophotogrammetry, or one camera which is moved around the subject to capture multiple images from different views, or both), which in some implementations, may be arranged into 3D representations such as 3D meshes, 3D point clouds and/or 3D voxelized representations. The analysis of a 3D representation of the subject may in some instances provide technical improvements over 2D analysis of the same subject. For example, a 3D representation may describe the geometry and/or structure of the subject with less ambiguity than a 2D representation (which may contain shadows and other artifacts which complicate the depiction of depth from the subject and texture of the subject). In some implementations, 3D processing may enable technical improvements because of the inverse optics problem which may, in some instances, affect 2D representations. The inverse optics problem refers to the phenomenon where, in some instances, the size of a subject, the orientation of the subject and the distance between the subject and the imaging device may be conflated in a 2D image of that subject. Any given projection of the subject on the imaging sensor could map to an infinite count of {size, orientation, distance} pairings. 3D representations enable the technical improvement in that 3D representations remove the ambiguities introduced by the inverse optics problem.
[00103] A device that is configmed with the dedicated purpose of 3D scanning, such as a 3D intraoral scanner (or a CT scanner or MRI scanner), may generate 3D representations of the subject (e.g., the patient's dentition) which have significantly higher fidelity and precision than is possible with a handheld device. When such high-fidelity 3D data are available (e.g., in the application of oral care mesh classification or other 3D techniques described herein), the use of a 3D autoencoder is offers technical improvements (such as increased data precision), to extract the best possible signal out of those 3D data (i.e., to get the signal out of the 3D crown meshes used in tooth classification or setups classification). [00104] A 3D autoencoder (comprising a 3D encoder and a 3D decoder) may be trained on 3D data representations to encode an input 3D representation into a latent form (such as a latent vector or a latent capsule) using the 3D encoder, and then reconstruct a facsimile of the input 3D representation using the 3D decoder. Among the operations which may be used to implement a 3D autoencoder for the analysis of a 3D representation (e.g., 3D mesh or 3D point cloud) are 3D convolution, 3D pooling and 3D reconstruction error calculation.
[00105] For each mesh element, a 3D convolution may be performed to aggregate local features from nearby mesh elements. Processing may be performed above and beyond the techniques for 2D convolution, to account for the differing count and locations of neighboring mesh elements (relative to a particular mesh element). A particular 3D mesh element may have a variable count of neighbors and those neighbors may not be found in expected locations (as opposed to a pixel in 2D convolution which may have a fixed count of neighboring pixels which may be found in known or expected locations). In some instances, the order of neighboring mesh elements may be relevant to 3D convolution.
[00106] A 3D pooling operation may enable the combining of features from a 3D mesh (or other 3D representation) at multiple scales. 3D pooling may iteratively reduce a 3D mesh into mesh elements which are most highly relevant to a given application (e.g., for which a neural network has been trained). Similarly to 3D convolution, 3D pooling may benefit from processing beyond that entailed in 2D convolution, to account for the differing count and locations of neighboring mesh elements (relative to a particular mesh element). In some instances, the order of neighboring mesh elements may be less relevant to 3D pooling than to 3D convolution.
[00107] 3D reconstruction error may be computed using one or more of the techniques described herein, such as computing Euclidean distances between corresponding mesh elements, the two meshes. Other techniques are possible in accordance with aspects of this disclosure. 3D reconstruction error may generally be computed on 3D mesh elements, rather than the 2D pixels of 2D reconstruction error. 3D reconstruction error may enable technical improvements over 2D reconstruction error, because a 3D representation may, in some instances, have less ambiguity than a 2D representation (i.e., have less ambiguity in form, shape and/or structure). Additional processing may, in some implementations, be entailed for 3D reconstruction which is above and beyond that of 2D reconstruction, because of the complexity of mapping between the input and reconstructed mesh elements (i.e., the input and reconstructed meshes may have different mesh element counts, and there may be a less clear mapping between mesh elements than there is for the mapping between pixels in 2D reconstruction). The technical improvements of 3D reconstruction error calculation include data precision improvement. [00108] A 3D representation may be produced using a 3D scanner, such as an intraoral scanner, a computerized tomography (CT) scanner, ultrasound scanner, a magnetic resonance imaging (MRI) machine or a mobile device which is enabled to perform stereophotogrammetry. A 3D representation may describe the shape and/or structure of a subject. A 3D representation may include one or more 3D mesh, 3D point cloud, and/or a 3D voxelized representation, among others. A 3D mesh includes edges, vertices, or faces. Though interrelated in some instances, these three types of data are distinct. The vertices are the points in 3D space that define the boundaries of the mesh. These points would alternatively be described as a point cloud but for the additional information about how the points are connected to each other, as described by the edges. An edge is described by two points and can also be referred to as a line segment A face is described by a number of edges and vertices. For instance, in the case of a triangle mesh, a face comprises three vertices, where the vertices are interconnected to form three contiguous edges. Some meshes may contain degenerate elements, such as non-manifold mesh elements, which may be removed, to benefit Other mesh pre-processing operations are possible in accordance with aspects of this disclosure. 3D meshes are commonly formed using triangles, but may in other implementations be formed using quadrilaterals, pentagons, or some other n-sided polygon. In some implementations, a 3D mesh may be converted to one or more voxelized geometries (i.e., comprising voxels), such as in the case that sparse processing is performed. The techniques of this disclosure which operate on 3D meshes may receive as input one or more tooth meshes (e.g., arranged in one or more dental arches). Each of these meshes may undergo pre-processing before being input to the predictive architecture (e.g., including at least one of an encoder, decoder, pyramid encoder-decoder and U-Net). This pre-processing may include the conversion of the mesh into lists of mesh elements, such as vertices, edges, faces or in the case of sparse processing - voxels. For the chosen mesh element type or types (e.g., vertices), feature vectors may be generated. In some examples, one feature vector is generated per vertex of the mesh. Each feature vector may contain a combination of spatial and/or structural features, as specified in the following table:
Table 1
[00109] Table 1 discloses non-limiting examples of mesh element features. In some implementations, color (or other visual cues/identifiers) may be considered as a mesh element feature in addition to the spatial or structural mesh element features described in Table 1. As used herein (e.g., in Table 1), a point differs from a vertex in that a point is part of a 3D point cloud, whereas a vertex is part of a 3D mesh and may have incident faces or edges. A dihedral angle (which may be expressed in either radians or degrees) may be computed as the angle (e.g., a signed angle) between two connected faces (e.g., two faces which are connected along an edge). A sign on a dihedral angle may reveal information about the convexity or concavity of a mesh surface. For example, a positively signed angle may, in some implementations, indicate a convex surface. Furthermore, a negatively signed angle may, in some implementations, indicate a concave surface. To calculate the principal curvature of a mesh vertex, directional curvatures may first be calculated to each adjacent vertex around the vertex. These directional curvatures may be sorted in circular order (e.g., 0, 49, 127, 210, 305 degrees) in proximity to the vertex normal vector and may comprise a subsampled version of the complete curvature tensor. Circular order means: sorted in by angle around an axis. The sorted directional curvatures may contribute to a linear system of equations amenable to a closed form solution which may estimate the two principal curvatures and directions, which may characterize the complete curvature tensor. Consistent with Table 1, a voxel may also have features which are computed as the aggregates of the other mesh elements (e.g., vertices, edges and faces) which either intersect the voxel or, in some implementations, are predominantly or fully contained within the voxel. Rotating the mesh may not change structural features but may change spatial features. And, as described elsewhere in this disclosure, the term “mesh” should be considered in a nonlimiting sense to be inclusive of 3D mesh, 3D point cloud and 3D voxelized representation. In some implementations, apart from mesh element features, there are alternative methods of describing the geometry of a mesh, such as 3D keypoints and 3D descriptors. Examples of such 3D keypoints and 3D descriptors are found in “TONIONI A, et al. in ‘Learning to detect good 3D keypoints.’, Int J Comput. Vis. 2018 Vol .126, pages 1-20 3D keypoints and 3D descriptors may, in some implementations, describe extrema (either minima or maxima) of the surface of a 3D representation. In some implementations, one or more mesh element features may be computed, at least in part, via deep feature synthesis (DFS), e.g. as described in: J. M. Kanter and K. Veeramachaneni, "Deep feature synthesis: Towards automating data science endeavors," 2015 IEEE International Conference on Data Science and Advanced Analytics (DSAA), 2015, pp.
[00110] Predictive models which may operate on feature vectors of the aforementioned features include but are not limited to: GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, Tooth Classification, Setups Classification, Setups Comparison, VAE Mesh Element Labeling, MAE Mesh In-filling, Mesh Reconstruction Autoencoder, Validation Using Autoencoders, Mesh Segmentation, Coordinate System Prediction, Mesh Cleanup, Restoration Design Generation, Appliance Component Generation and/or Placement, and Archform Prediction, may input In some implementations, such feature vectors may be presented to one or more internal layers of a neural network which is part of one or more of those predictive models.
[00111] Representation generation neural networks based on autoencoders, U-Nets, transformers, other types of encoder-decoder structures, convolution and/or pooling layers, or other models may benefit from the use of mesh element features. Mesh element features may convey aspects of a 3D representation’s surface shape and/or structure to the neural network models of this disclosure. Each mesh element feature describes distinct information about the 3D representation that may not be redundantly present in other input data that are provided to the neural network. For example, a vertex curvature may quantify aspects of the concavity or convexity of the surface of a 3D representation which would not otherwise be understood by the network. Stated differently, mesh element features may provide a processed version of the structure and/or shape of the 3D representation; data that would not otherwise be available to the neural network. This processed information is often more accessible, or more amenable for encoding by the neural network. A system implementing the techniques disclosed herein has been utilized to mn a number of experiments on 3D representations of teeth. For example, mesh element features have been provided to a representation generation neural network which is based on a U-Net model, and also to a representation generation model based on a variational autoencoder with continuous normalizing flows. Based on experiments, it was found that systems using a full complement of mesh element features (e.g., “XYZ” coordinates tuple, “Normal vector”, “Vertex Curvature”, Points- Pivoted, and Normals-Pivoted) were at least 3% more accurate than systems that did not. Points-Pivoted describes “XYZ” coordinates tuples that have local coordinate systems (e.g., at the centroid of the respective tooth). Normals-Pivoted describes “Normal Vectors” which have local coordinate systems (e.g., at the centroid of the respective tooth). Furthermore, training converges more quickly when the full complement of mesh element features are used. Stated another way, the machine learning models trained using the full complement of mesh element features tended to be more accurate more quickly (at earlier epochs) than systems which did not. For an existing system observed to have a historical accuracy rate of 91%, an improvement in accuracy of 3% reduces the actual error rate by more than 30%.
[00112] As described herein, tooth movements specify one or more transformations that can be encoded in various ways to specify tooth positions and orientations within the setup and are applied to 3D representations of teeth. For instance, according to particular implementations, the tooth positions can be cartesian coordinates of a tooth's canonical origin location which is defined in some semantic context. Tooth orientations can be represented as rotation matrices, unit quaternions, or other 3D rotation representations such as Euler angles with respect to a frame of reference (either global or local). Dimensions are real valued 3D spatial extents and gaps can be binary presence indicators or real valued gap sizes between teeth especially in instances when certain teeth are missing. In some implementations, tooth rotations may be described by 3x3 matrices (or by matrices of other Tooth position and rotation information may, in some implementations, be combined into the same transform matrix, for example, as a 4x4 matrix, which may reflect homogenous coordinates, in some instances, affine spatial transformation matrices may be used to describe tooth transformations, for example, the transformations which describe the maloccluded pose of a tooth, an intermediate pose of a tooth and/or a final setup pose of a tooth. Some implementations may use relative coordinates, where setup transformations are predicted relative to malocclusion coordinate systems (e.g., a malocclusion-to-setup transformation is predicted instead of a setup coordinate system directly). Other implementations may use absolute coordinates, where setup coordinate systems are predicted directly for each tooth. In the relative mode, transforms can be computed with respect to the centroid of each tooth mesh (vs the global origin), which is termed “relative local.” Some of the advantages of using relative local coordinates include eliminating the need for malocclusion coordinate systems (landmarking data) which may not be available for all patient case datasets. Some of the advantages of using absolute coordinates include simplifying the data preprocessing as mesh data are originally represented as relative to the global origin. These details about tooth position encoding and tooth orientation encoding may, in some implementations, also apply one or more of the neural networks models of the present disclosure, including but not limited to: GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, FDG Setups, Setups Classification, Setups Comparison, VAE Mesh Element Labeling, MAE Mesh Infilling, Mesh Reconstruction VAE, and Validation Using Autoencoders.
[00113] According to some implementations of the techniques described herein, convolution layers in the various 3D neural networks described herein may use edge data to perform mesh convolution. The use of edge information reduces or potentially eliminates the model’s sensitivity to different input orders of 3D elements. In addition to or as an alternative to using edge data, the convolution layers may use vertex data to perform mesh convolution. The use of vertex information is advantageous in that there are typically fewer vertices than edges or faces, thereby enabling vertex-oriented processing to function at a lower processing overhead and lower computational resource cost. In addition to or separate from using edge data or vertex data, the convolution layers may use face data to perform mesh convolution. Furthermore, in addition to or separate from using edge data, vertex data, or face data, the convolution layers may use voxel data to perform mesh convolution. The use of voxel information provides a technical improvement in that, depending on the granularity chosen, there may be significantly fewer voxels to process compared to the vertices, edges or faces in the mesh. Sparse processing (with voxels) may lead to a lower processing overhead and lower computational cost (especially in terms of computer memory or RAM usage).
[00114] Examples are disclosed herein. The orthodontic metrics may be used to quantify the physical arrangement of an arch of teeth for the purpose of orthodontic treatment (as opposed to restoration design metrics - which pertain to dentistry and describe the shape and/or form of one or more pre-restoration teeth, for the purpose of supporting dental restoration). These orthodontic metrics can measure how badly maloccluded the arch is, or conversely the metrics can measure how correctly arranged the teeth are. In some implementations, the GDL Setups model RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity SetupsFDG Setups) may incorporate one or more of these orthodontic metrics, or other similar or related orthodontic metrics. In some implementations, such orthodontic metrics may be incorporated into the feature vector for a mesh element, where these perelement feature vectors are provided to the setups prediction network as inputs. In some implementations, such orthodontic metrics may be directly consumed by a generator, an MLP, a transformer, or other neural network as direct inputs (such as presented in one or more input vectors of real numbers S, such as described elsewhere in this disclosure The use of such orthodontic metrics in the training of the generator may improve the performance (e., correctness) of the , resulting in predicted transforms which place teeth more nearly in the correct final setups poses than would otherwise be possible. Such orthodontic metrics may be consumed by an encoder structure or by a U-Net structure (in the case of GDL Setups). Such orthodontic metrics may be consumed by an autoencoder, variational autoencoder, masked autoencoder or regularized autoencoder (in the case of the VAE Setups, VAE Mesh Element Labelling, MAE Mesh In-Filling). Such orthodontic metrics may be provided to a neural network which generates action predictions as a part of a reinforcement learning RL Setups model, orthodontic metrics a label to a setup arch (e.g., labels such as mal, staging or final setup). This description is non-limiting, as the orthodontic metrics may also be incorporated in other ways into the various techniques of this disclosure.
[00115] The various loss calculations of the present disclosure may, in some examples, incorporate one or more orthodontic metrics, with the of the correctness of the resulting neural An orthodontic metric may be used to directly compare a predicted example to the corresponding ground truth example (such as is with metrics In other examples, one or more orthodontic metrics may be incorporated into a loss computation. Such an orthodontic metric may be computed on the predicted example, and then the orthodontic metric would also be computed on the ground tmth example. These two orthodontic metrics results would then be consumed by the loss computation, with the advantage of improving the performance of the resulting neural network. In some implementations, one or more orthodontic metrics pertaining to the alignment of two or more adjacent teeth may be computed and incorporated into a loss function, for example, to train, at least in part, a setups prediction neural network In some implementations, such an orthodontic metric may influence the network to align the mesial surface of a tooth with the distal surface of an adjacent tooth. Backpropagation is an example algorithm by which a neural network may be trained using one or more loss values.
[00116] In some implementations, one or more orthodontic metrics may be used to evaluate the predicted output of a neural network, such as a setups prediction Such metric(s) may enable the training algorithm to determine how close the predicted output is to an acceptable output, for example, in a quantified sense. In some implementations, this use of an orthodontic metric may enable a loss value to be computed which does not depend entirely on a comparison to a ground truth In some implementations, such a use of an orthodontic metric may enable loss calculation and network training to proceed without the need for a comparison against a ground truth example. The advantage of such an approach is that loss may be computed based on a general principle or specification for the predicted output (such as a setup) rather than tying loss calculation to a specific ground tmth example (which may have been defined by a particular doctor, clinician, or technician, whose treatment may differ from that of other technicians or doctors). In some implementations, such an orthodontic metric may be defined based on a Frechet Inception Distance) score.
[00117] The following is a description of some of the orthodontic metrics which are used to quantify the state of a set of teeth in an arch for the purpose of orthodontic treatment. These orthodontic metrics indicate the degree of malocclusion that the teeth are in at a given stage of clear tray aligner treatment. [00118] An orthodontic metric that can be computed using tensors may be especially advantageous when training one of the neural networks of the present disclosure, because tensor operations may promote efficient computations. The more efficient (and faster) the computation, the faster the rate at which training can proceed.
[00119] In some examples, an error pattern may be identified in one or more predicted outputs of an ML model (e.g., a transformation matrix for a predicted tooth setup, a labelling of mesh elements for mesh cleanup, an addition of mesh elements to a mesh for the purpose of mesh in-filling, a classification label for a setup, a classification label for a tooth mesh, etc.). One or more orthodontic metrics may be selected to become an input to the next round of ML model training, to address any pattern of errors or deficiencies which may be identified in the one or more predicted outputs.
[00120] may be defined relative to an archfrom coordinate frame, the LDE coordinate system. In some implementations, a point may be described using an LDE coordinate frame relative to an archform, where L, D and E correspond to: 1) Length along the curve of the archform, 2) Distance away from the archform, and 3) distance in the direction perpendicular to the L and D axes (which may be termed Eminence), respectively. [00121] Various of the OM and other techniques of the present disclosure may compute collisions between 3D representations (e.g., of oral care objects, such as teeth). Such collisions may be computed as at least one of: 1) penetration distance between 3D tooth representations, 2) count of overlapping mesh elements between 3D tooth representations, and 3) volume of overlap between 3D tooth representations. In some implementations, an OM may be defined to quantify the collision of two or more 3D representations of oral care structures, such as teeth. Some optimization algorithms, such as setups prediction techniques, may seek to minimize collisions between oral care structures (such as teeth). Between-arch orthodontic metrics are described as follows.
[00122] Six (6) metrics for the comparison of two or more arches are listed below. Other suitable comparison orthodontic metrics are found elsewhere in this disclosure, such as in the section for the Setups Comparison technique.
1. Rotation geodesic distance (rotation between predicted example and ground truth setup example)
2. Translation distance (gap between predicted example and ground truth setup example)
3. Normalized translation distance
4. 3D alignment error that measures the distance between predicted mesh elements and ground truth mesh elements, in units of mm.
5. Normalized 3D alignment
6. Percent overlap (% overlap) by volume (alternatively % overlap by mesh elements) of predicted example and corresponding ground truth example
[00123] Within-arch orthodontic metrics are described as follows.
Alignment - A 3D tooth orientation vector may be calculated using the tooth's mesial-distal axis. A 3D vector, which may be tangent vector to the archform at the position of the tooth may also be calculated. The XY components (i.e., which may be 2D vectors) may then be used to compare the orientation of the archform at the tooth's location to the tooth's orientation in XY space. Cosine similarity may be used to calculate the 2D orientation difference (angle) between the archform tangent and the tooth's mesial-distal axis.
Arch Symmetry - For each left-right pair of teeth (e.g., lower left lateral incisor and/or lower right lateral incisor) the absolute difference may be calculated between each tooth’s X-coordinate and the global coordinate reference frame’s X-axis, delta may indicate the arch asymmetry for a given tooth pair. The result of such a calculation may be the mean X-axis delta of one or more tooth-pairs from the arch. This calculation may, in some implementations, be performed relative to the Y-axis withy-coordinates (and/or relative to the Z axis with Z-coordinates).
Archform D-axis Differences - dimension difference (e., the positional difference in the facial -lingual direction) between two arch states, for one or more teeth, some implementations, a dictionary of the D- direction tooth movement for each tooth, with tooth UNS number as the key. May use the LDE coordinate system relative to an archform. Archform (Lower) Length Ratio - the ratio between the current lower arch length and the arch length as it was in the original maloccluded lower arch.
Archform (Upper) Length Ratio - the ratio between the current upper arch length and the arch length as it was in the original maloccluded upper arch.
Archform Parallelism (Full arch) - For at least one local tooth coordinate system origin in the upper arch, the one or more nearest origins (e.g., tooth local coordinate system origins) in the lower arch. In some implementations, the two nearest origins may be used, the straight line distance from the upper arch point to the line formed between the origins of the two teeth in the opposing (lower) arch. May return the standard deviation of the set of “point-to-line" distances mentioned above, where the set may be composed of the point-to-line distances for each tooth in the arch.
Archform Parallelism (Individual tooth) - This metric may share some computational with the archform_parallelism_global orthodontic metric, except that this metric may input the mean distance from a tooth origin to the line formed by the neighboring teeth in opposing arches (e.g., a tooth in the upper arch and the corresponding tooth in the lower arch). The mean distance may be computed for one or more such pairs of teeth. In some implementations, this may be computed for all pairs of teeth. Then the mean distance may be subtracted from the distance that is computed for each tooth pair. This OM may yield the deviation of a tooth from a “typical” tooth parallelism in the arch.
Buccolingual Inclination - For at least one molar or premolar, the corresponding tooth on the opposite side of the same arch (i.e., for a tooth on the left side of the arch, find the same type of tooth on the right side and vice versa), an n-element list for each tooth (e.g. n may equal 2). This list may contain at least the tooth IDs of the teeth in each pair of teeth (e.g., LeftLowerFirstMolar and RightLowerFirstMolar in a list = [left tooth idx l, right_tooth_idx_2]). Such an n-element vector may be computed for each molar and each premolar in the upper and lower arches. The buccal cusps identified on the molars and premolars on each of the left and right sides of the arch. Draw a line between the buccal cusps of the left tooth and the buccal cusps on the right tooth. Make a plane using this line and the z-axis of the arch. The lingual cusps may be projected onto the plane (i.e., at this point the angle of inclination may be determined). By performing an additional projection, the approximate vertical distance between the lingual cusps and the buccal cusps may be computed. This distance may be used as the buccolingual inclination OM.
Canine Overbite - The upper and lower canines may be identified. The first premolar for the given side of the mouth may be identified. On a given side of the arch, a distance may be computed between the upper canine and the lower canine, and also between the upper pre-molar and the lower pre-molar. The average ( median, or mode or some other statistic) may be computed for the measured distances. The z- component of this result indicates the degree of overbite. Overbite may be computed between any tooth in one arch and the corresponding tooth in the other arch.
Canine Overjet Contact - May calculate the collisions (e.g., collision distances) between pairs of canines on opposing arches.
Canine Overjet Contact KDE - May take an orthodontic metric score for the current patient case as input, and may convert that score into to a log-likelihood using a previously trained kernel density estimation (KDE) model or distribution. This operation may yield information about where in the distribution of "typical" values this patient case lies.
Canine Overjet - This OM may share some computational steps with the canine overbite OM. In some implementations, average distances may be computed. In some implementations, the distance calculation may compute the Euclidean distance of the XY components of a tooth in the upper arch and a tooth in the lower arch, to yield oveget (i.e., as opposed to computing the difference in Z-components, as may be performed for canine overbite). Oveget may be computed between any tooth in one arch and the corresponding tooth in the other arch.
Canine Class Relationship (also applies to first, second and third molars) - This OM may, in some implementations comprise two functions (e.g., written in Python). get_canine_landmarks(): Get landmarks for each tooth which may be used to compute the class relationship, and then, in some implementations, map those landmarks onto the global coordinate space so that measurements may be made between teeth. class_relationship_score_by_side(): May compute the average position of at least one landmark on at least one tooth in the lower arch, and may compute the same for the upper arch. Then may compute the vector from the upper arch landmark position to the lower arch landmark position, and finally projects this vector onto the lower arch to yield a quantification (e.g., as a scalar) of the amount of delta in “arch 1-axis" position there is. This OM may compute how far forward or behind the tooth is positioned on the 1-axis relative to the tooth or teeth of interest in the opposing arch.
Crossbite - Fossa in at least one upper molar may be located by finding the halfway point between distal and mesial marginal ridge saddles of the tooth. A lower molar cusp may lie between the marginal ridges of the corresponding upper molar. This OM may compute a vector from the upper molar fossa midpoint to the lower molar cusp. This vector may be projected onto the d-axis of the archform, yielding a lateral measure of distance from the cusp to the fossa. This distance may define the crossbite magnitude.
Edge Alignment - This OM may identify the leftmost and rightmost edges of a tooth, and may identify the for that tooth’s neighbor.
The OM may then draw a vector from the leftmost edge of the tooth to the leftmost edge of the tooth’s neighbor.
The OM may then draw a vector from the rightmost edge of the tooth to the rightmost edge of the tooth’s neighbor.
The OM may then calculates the linear fit error between the two vectors.
Such a calculation may involve making two vectors:
Vec tooth = right tooths leftside to left tooths leftside
Vec neighbor = right tooths rightside to left tooths leftside
And then may involve computing the dot-product of these two vectors and subtracting the result from 1. (i.e., EdgeAlignment score = 1 - abs(dot(Vec_tooth, Vec neighbor)) ).
A score of 0 may indicate perfect alignment. A score of 1 may mean perpendicular alignment. Incisor Interarch Contact KDE - May identify the deviation of the IncisorlnterarchContact from the mean of a modeled distribution of such statistics across a dataset of one or more other patient cases.
Leveling - May compute a measure of leveling between a tooth and its neighbor. This OM may calculate the difference in height between two or more neighboring teeth. For molars, this OM may use the midpoint between the mesial and distal saddle ridges as the height of the molar. For non-molar teeth, this OM may use the length of the crown from gums to tip. In some implementations, the tip may be the origin of the local coordinate space of the tooth. Other implementations may place the origin in other locations. A simple subtraction between the heights of neighboring teeth may yield the leveling delta between the teeth (e.g., by comparing Z components).
Midline - May compute the position of the midline for the upper incisors and/or the lower incisors, and then may compute the distance between them.
Molar Interarch Contact KDE - May compute a molar interarch contact score (i.e., a collision depth or other type of collision), and then may identify where that score lies in a pre-defined KDE (distribution) built from representative cases.
Occlusal Contacts - For a particular tooth from the arch, this OM may identify one or more landmarks (e.g., mesial cusp, or central cusp, etc.). Get the tooth transform for that tooth. For each cusp on the current tooth, the cusp may be scored according to how well the cusp contacts the neighboring (corresponding) tooth in the opposite arch. A vector may be found from the cusp of the tooth in question to the vertical intersection point in the corresponding tooth of the opposing arch. The distance and/or direction (i.e., up or down) to the opposing arch may be computed. A list may be returned that contains the resulting signed distances, one for each cusp on the tooth in question.
Overbite - The upper and lower central incisors may be compared along the z-axis. The difference along the z-axis may be used as the overbite score.
Overjet - The upper and lower central incisors may be compared along the y-axis. The difference along the y-axis may be used as the oveijet score.
Molar Interarch Contact - May calculate the contact score between molars, and may use collision measurement(s) (such as collision depth).
Root Movement d - The tooth transforms for an initial state and a next state may be recieved. The archform axes at a point L along the archform may be computed. This OM may return a distance moved along the d-axis. This may be accomplished by projecting the root pivot point onto the d-axis.
Root Movement 1 - The tooth transforms for an initial state and a next state may be received. The archform axes at a point L along the archform may be computed. This OM may return a distance moved along the 1-axis. This may be accomplished by projecting the root pivot point onto the 1-axis.
Spacing - May compute the spacing between each tooth and its neighbor. The transforms and meshes for the arch may be received. The left and right edges of each tooth mesh may be computed. One or more points of interest may be transformed from local coordinates into the global arch coordinate frame. The spacing may be computed in a plane (e.g., the XY plane) between each tooth and its neighbor to the "left". May return an array of one or more Euclidean distances (e.g., such as inthe XY plane) which may represent the spacing between each tooth and its neighbor to the left.
Torque - May compute torque (i.e., rotation around and axis, such as the x-axis). For one or more teeth, one or more rotations may be converted from Euler angles into one or more rotation matrices. A component (such as a x-component) of the rotations may be extracted and converted back into Euler angles. This x- component may be interpreted as the torque for a tooth. A list maybe returned which contains the torque for one or more teeth, and may be indexed by the UNS number of the tooth.
[00124] The neural networks of this disclosure may exploit one or more benefits of the operation of parameter tuning, whereby the inputs and parameters of a neural network are optimized to produce more data-precise results. One parameter which may be tuned is neural network learning rate (e.g., which may have values such as 0.1, 0.01, 0.001, etc.). Data augmentation schemes may also be tuned or optimized, such as schemes where “shiver” is added to the tooth meshes before being input to the neural network (i.e., small random rotations, translations and/or scaling may be applied to vary the dataset and make the neural network robust to variations in data).
A subset of the neural network model parameters available for tuning are as follows: o Learning rate (LR) decay rate (e.g., how much the LR decays during a training run) o Learning rate (LR). The floating-point value (e.g., 0.001) that is used by the optimizer. o LR schedule (e.g., cosine annealing, step, exponential) o Voxel size (for cases with sparse mesh processing operations) o Dropout % (e.g., dropout which may be performed in a linear encoder) o LR decay step size (e.g., decay evety 10 or 20 or 30 epochs) o Model scaling, which may increase or decrease the count of layers and/or the count of parameters per layer.
[00125] Parameter tuning may be advantageously applied to the training of a neural network for the prediction of final setups or intermediate staging to provide data precision-oriented technical improvements. Parameter tuning may also be advantageously applied to the training of a neural network for mesh element labeling or a neural network for mesh in-filling. In some examples, parameter tuning may be advantageously applied to the training of a neural network for tooth reconstruction. In terms of classifier models of this disclosure, parameter tuning may be advantageously applied to a neural network for the classification of one or more setups (i.e., classification of one or more arrangements of teeth). The advantage of parameter tuning is to improve the data precision of the output of a predictive model or a classification model. Parameter tuning may, in some instances, provide the advantage of obtaining the last remaining few percentage points of validation accuracy out of a predictive or classification model. [00126] Some techniques of the present disclosure, for example the setups comparison technique, the setups prediction techniques (e.g., such as GDL Setups, MLP Setups, VAE Setups and the like), may benefit from a processing step which may align (or register) arches of teeth (e.g., where a tooth may be represented by a 3D point cloud, or some other type of 3D representation described herein). Such a processing setup may, for example, be used to register a ground truth setup arch from a patient case with the maloccluded arch from that same case, before these mat and ground truth setup arches are used to train a setups prediction neural network model. Such a step may aid in loss calculation, because the predicted arch (e.g., an arch outputted by a generator) may be in better alignment with the ground truth setup arch, a condition which may facilitate the calculation of reconstruction loss, representation loss, LI loss, L2 loss, MSE loss and/or other kinds of losses described herein. In some implementations, an iterative closest point (ICP) technique may be used for such registration. ICP may minimize the squared errors between corresponding entities, such as 3D representations. In some implementations, linear least squares calculations may be performed. In some implementations, non-linear least squares calculations may be performed. Various registration models may incorporate portions of the following algorithms, in whole or in part: Levenberg-Marquardt ICP, Least Square Rigid transformation, Robust Rigid transformation, random sample consensus (RANSAC) ICP, K-means based RANSAC ICP and Generalized ICP (GICP). Registration may, in some instances, decrease the subjectivity and/or randomness that may, in some instances, occur in reference ground truth setup designs which have been designed by technicians (i.e., two technicians may produce different but valid final setups outputs for the same case) or by other optimization techniques.
[00127] Various neural network models of this disclosure may draw benefits from data augmentation. Examples include models of this which are trained on 3D meshes, such as GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, FDG Setups, Setups Classification, Setups Comparison, VAE Mesh Element Labeling, MAE Mesh In-filling, Mesh Reconstruction VAE, and Validation Using Autoencoders. Data augmentation, such as by way of the method shown in FIG. 1, may increase the size of the training dataset of dental arches. Data augmentation can provide additional training examples by adding random rotations, translations, and/or rescaling to copies of existing dental arches. In some implementations of the techniques of this disclosure, data augmentation may be carried out by perturbing or jittering the vertices of the mesh, in a manner similar to that described in (“Equidistant and Uniform Data Augmentation for 3D Objects”, IEEE Access, Digital Object Identifier 10.1109/ACCESS.2021.3138162). The position of a vertex may be perturbed through the addition of Gaussian noise, for example with zero mean, and 0.1 standard deviation. Other mean and standard deviation values are possible in accordance with the techniques of this disclosure.
[00128] FIG. 1 shows a data augmentation method that systems of this disclosure may apply to 3D oral care representations. A non-limiting example of a 3D oral care representation is a tooth mesh or a set of tooth meshes. Tooth data 100 (e.g., 3D meshes) are received at the input. The systems of this disclosure may generate copies of the tooth data 100 (102). In the example of FIG. 1, the systems of this disclosure may apply one or more stochastic rotations to the tooth data 100 (104). In the example of FIG. 1, the systems of this disclosure may apply stochastic translations to the tooth data 100 (106). The systems of this disclosure may apply stochastic scaling operations to the tooth data 100 (108). The systems of this disclosure may apply stochastic perturbations to one or more mesh elements of the tooth data 100 (110). The systems of this disclosure may output augmented tooth data 112 that are formed by way of the method of FIG. 1.
[00129] Because generator networks of this disclosure can be implemented as one or more neural networks, the generator may contain an activation function. When executed, an activation lunction outputs a determination of whether or not a neuron in a neural network will fire (e.g., send output to the next layer). Some activation functions may include: binary step functions, or linear activation functions. Other activation functions impart non-linear behavior to the network, including: sigmoid/logistic activation functions, Tanh (hyperbolic tangent) functions, rectified linear units (ReLU), leaky ReLU functions, parametric ReLU functions, exponential linear units (ELU), softmax function, swish function, Gaussian error linear unit (GELU), or scaled exponential linear unit (SELU). A linear activation function may be well suited to some regression applications (among other applications), in an output layer. A sigmoid/logistic activation function may be well suited to some binary classification applications (among other applications), in an output layer. A softmax activation function may be well suited to some multiclass classification applications (among other applications), in an output layer. A sigmoid activation function may be well suited to some multilabel classification applications (among other applications), in an output layer. A ReLU activation function may be well suited in some convolutional neural network (CNN) applications (among other applications), in a hidden layer. A Tanh and/or sigmoid activation function may be well suited in some recurrent neural network (RNN) applications (among other applications), for example, in a hidden layer. There are multiple optimization algorithms which can be used in the training of the neural networks of this disclosure (such as in updating the neural network weights), including gradient descent (which determines a training gradient using first-order derivatives and is commonly used in the training of neural networks), Newton's method (which may make use of second derivatives in loss calculation to find better training directions than gradient descent, but may require calculations involving Hessian matrices), and conjugate gradient methods (which may yield faster convergence than gradient descent, but do not require the Hessian matrix calculations which may be required by Newton's method). In some implementations, additional methods may be employed to update weights, in addition to or in place of the techniques described above. These additional methods include the Levenberg-Marquardt method and/or simulated annealing. The backpropagation algorithm is used to assign the results of loss calculation back into the network so that network weights can be adjusted, and learning can progress.
[00130] Neural networks contribute to the lunctioning of many of the applications of the present disclosure, including but not limited to: GDL Setups, RL Setups, VAE Setups, Capsule Setups, MLP Setups, Diffusion Setups, PT Setups, Similarity Setups, Tooth Classification, Setups Classification, Setups Comparison, VAE Mesh Element Labeling, MAE Mesh In-filling, Mesh Reconstruction Autoencoder, Validation Using Autoencoders, imputation of oral care parameters, 3D mesh segmentation (3D representation segmentation), Coordinate System Prediction, Mesh Cleanup, Restoration Design Generation, Appliance Component Generation and/or Placement, or Archform Prediction. The neural networks of the present disclosure may embody part or all of a variety of different neural network models. Examples include the U-Net architecture, multi-later perceptron (MLP), transformer, pyramid architecture, recurrent neural network (RNN), autoencoder, variational autoencoder, regularized autoencoder, conditional autoencoder, capsule network, capsule autoencoder, stacked capsule autoencoder, denoising autoencoder, sparse autoencoder, conditional autoencoder, long/short term memory (LSTM), gated recurrent unit (GRU), deep belief network (DBN), deep convolutional network (DCN), deep convolutional inverse graphics network (DCIGN), liquid state machine (LSM), extreme learning machine (ELM), echo state network (ESN), deep residual network (DRN), Kohonen network (KN), neural Turing machine (NTM), or generative adversarial network (GAN). In some implementations, an encoder structure or a decoder structure may be used. Each of these models provides one or more of its own particular advantages. For example, a particular neural networks architecture may be especially well suited to a particular ML technique. For example, autoencoders are particularly suited to the classification of 3D oral care representations, due to the ability to encode the 3D oral care representation into a form which is more easily classifiable.
[00131] In some implementations, the neural networks of this disclosure can be adapted to operate on 3D point cloud data (alternatively on 3D meshes or 3D voxelized representation). Numerous neural network implementations may be applied to the processing of 3D representations and may be applied to training predictive and/or generative models for oral care applications, including: PointNet, PointNet++, SO-Net, spherical convolutions, Monte Carlo convolutions and dynamic graph networks, PointCNN, ResNet, MeshNet, DGCNN, VoxNet, 3D-ShapeNets, Kd-Net, Point GCN, Grid-GCN, KCNet, PD-Flow, PU-Flow, MeshCNN and DSG-Net. Oral care applications include, but are not limited to: setups prediction (e.g., using VAE, RL, MLP, GDL, Capsule, Diffusion, etc. which have been trained for setups prediction), 3D representation segmentation, 3D representation coordinate system prediction, element labeling for 3D representation clean-up (VAE for Mesh Element labeling), in-filling of missing elements in 3D representation (MAE for Mesh In-Filling), dental restoration design generation, setups classification, appliance component generation and/or placement, archform prediction, imputation of oral care parameters, setups validation, or other validation applications and tooth 3D representation classification.
[00132] Some implementations of the techniques of this disclosure incorporate the use of an autoencoder. Autoencoders that can be used in accordance with aspects of this disclosure include but are not limited to: AtlasNet, FoldingNet and 3D-PointCapsNet. Some autoencoders may be implemented based on PointNet.
[00133] Representation learning may be applied to setups prediction techniques of this disclosure by training a neural network to learn a representation of the teeth, and then using another neural network to generate transforms for the teeth. Some implementations may use a VAE or a Capsule Autoencoder to generate a representation of the reconstruction characteristics of the one or more meshes related to the oral care domain (including, in some instances, information about the structures of the tooth meshes). Then that representation (either a latent vector or a latent capsule) may be used as input to a module which generates the one or more transforms for the one or more teeth. These transforms may in some implementations place the teeth into final setups poses. These transforms may in some implementations place the teeth into intermediate staging poses. In some implementations, a transform may be described by a 9x1 transformation vector (e.g., that specifies a translation vector and a quaternion). In other implementations, a transform may be described by a transformation matrix (e.g., a 4x4 affine transformation matrix).
[00134] In some implementations, systems of this disclosure may implement a principal components analysis (PCA) on an oral care mesh, and use the resulting principal components as at least a portion of the representation of the oral care mesh in subsequent machine learning and/or other predictive or generative processing.
[00135] Systems of this disclosure may implement end-to-end training. Some of the end-to-end training-based techniques of this disclosure may involve two or more neural networks, where the two or more neural networks are trained together (i.e., the weights are updated concurrently during the processing of each batch of input oral care data). End-to-end training may, in some implementations, be applied to setups prediction by concurrently training a neural network which leams a representation of the teeth, along with a neural network which generates the tooth transforms.
[00136] According to some of the transfer learning-based implementations of this disclosure, a neural network (e.g., a U-Net) may be trained on a first task (e.g., such as coordinate system prediction). The neural network trained on the first task may be executed to provide one or more of the starting neural network weights for the training of another neural network that is trained to perform a second task (e.g., setups prediction). The first network may learn the low-level neural network features of oral care meshes and be shown to work well at the first task. The second network may exhibit faster training and/or improved performance by using the first network as a starting point in training. Certain layers may be trained to encode neural network features for the oral care meshes that were in the training dataset. These layers may thereafter be fixed (or be subjected to minor changes over the course of training) and be combined with other neural network components, such as additional layers, which are trained for one or more oral care tasks (such as setups prediction). In this manner, a portion of a neural network for one or more of the techniques of the present disclosure (e.g., setups prediction) may receive initial training on another task, which may yield important learning in the trained network layers. This encoded learning may then be built upon with further task-specific training of another network.
[00137] In accordance with this disclosure, transfer learning may be used for setups prediction, as well as for other oral care applications, such as mesh classification (e.g., tooth or setups classification), mesh element labeling, mesh element in-filling, procedure parameter imputation, mesh segmentation, coordinate system prediction, restoration design generation, mesh validation (for any of the applications disclosed herein). In some implementations, a neural network trained to output predictions based on oral care meshes may first be partially trained on one of the following publicly available datasets, before being further trained on oral care data: Google PartNet dataset, ShapeNet dataset, ShapeNetCore dataset, Princeton Shape Benchmark dataset, ModelNet dataset, ObjectNet3D dataset, ThingilOK dataset (which is especially relevant to 3D printed parts validation), ABC: A Big CAD Model Dataset For Geometric Deep Learning, ScanObjectNN, VOCASET, 3D-FUTURE, MCB: Mechanical Components Benchmark, PoseNet dataset, PointCNN dataset, MeshNet dataset, MeshCNN dataset, PointNet++ dataset, PointNet dataset, or PointCNN dataset.
[00138] In some implementations, a neural network which was previously trained on a first dataset (either oral care data or other data) may subsequently receive further training on oral care data and be applied to oral care applications (such as setups prediction). Transfer learning maybe employed to further train any of the following networks: GCN (Graph Convolutional Networks), PointNet, ResNet or any of the other neural networks from the published literature which are listed above.
[00139] In some implementations, a first neural network may be trained to predict coordinate systems for teeth (such as by using the techniques described in WO2022123402A1 or US Provisional Application No. US63/366492). A second neural network may be trained for setups prediction, according to any of the setups prediction techniques of the present disclosure (or a combination of any two or more of the techniques described herein). Transfer learning may assign at least a portion of the knowledge or capability of the first neural network to the second neural network. As such, transfer learning may provide the second neural network an accelerated training phase to reach convergence. In some implementations, the training of the second network may, after being augmented with the transferred learning, then be completed using one or more of the techniques of this disclosure.
[00140] Systems of this disclosure may train ML models with representation learning. The advantages of representation learning include that the generative network (e.g., neural network that predicts a transform for use in setups prediction) can be configured to receive input with a known size and/or standard format, as opposed to receiving input with a variable size or structure. Representation learning may produce improved performance over other techniques, because noise in the input data may be reduced (e.g., because the representation generation model extracts hierarchical neural network features and/or reconstruction characteristics of an inputted representation (e.g., a mesh or point cloud) through loss calculations or network architectures chosen for that purpose).
[00141] Reconstruction characteristics may comprise values in of a latent representation (e.g., a latent vector) that describe aspects of the shape and/or structure of the 3D representation that was provided to the representation generation module that generated the latent representation. The weights of the encoder module of a reconstruction autoencoder, for example, may be trained to encode a 3D representation (e.g., a 3D mesh, or others described herein) into a latent vector representation (e.g., a latent vector). Stated another way, the capability to encode a large set (e.g., hundreds, thousands or millions) of mesh elements into a latent vector (e.g., of hundreds or a thousand real values - e.g., 512, 1024, etc.) may be learned by the weights of the encoder. Each dimension of that latent vector may contain a real number which describes some aspect of the shape and/or structure of the original 3D representation. The weights of the decoder module of the reconstruction autoencoder may be trained to reconstruct the latent vector into a close facsimile of the original 3D representation. Stated another way, the capability to interpret the dimensions of the latent vector, and to decode the values within those dimensions, may be learned by the decoder. In summary, the encoder and decoder neural network modules are trained to perform the mapping of a 3D representation into a latent vector, which may then be mapped back (or otherwise reconstructed) into a 3D representation that is substantially similar to an original 3D representation for which the latent vector was generated.
[00142] Returning to loss calculation, examples of loss calculation may include KL-divergence loss, reconstruction loss or other losses disclosed herein. Representation learning may reduce the size of the dataset required for training a model, because the representation model learns the representation, enabling the generative network to focus on learning the generative task. The result may be improved model generalization because meaningful neural network features of the input data (e.g., local and/or global features) are made available to the generative network. Stated another way, a first network may learn the representation, and a second network may make the predictive decision. By training two networks to perform their own separate tasks, each of the networks may generate more accurate results for their respective tasks than with a single network which is trained to both learn a representation and make a decision. In some instances, transfer learning may first train a representation generation model. That representation generation model (in whole or in part) may then be used to pre-train a subsequent model, such as a generative model (e.g., that generates transform predictions). A representation generation model may benefit from taking mesh element features as input, to improve the capability of a second ML module to encode the structure and/or shape of the inputted 3D oral care representations in the training dataset.
[00143] One or more of the neural networks models of this disclosure may have attention gates integrated within. Attention gate integration provides the advantage of enabling the associated neural network architecture to focus resources on one or more input values. In some implementations, an attention gate may be integrated with a U-Net architecture, with the advantage of enabling the U-Net to focus on certain inputs, such as input flags which correspond to teeth which are meant to be fixed (e.g.,. prevented from moving) timing orthodontic treatment (or which require other special handling). An attention gate may also be integrated with an encoder or with an autoencoder (such as VAE or capsule autoencoder) to improve predictive accuracy, in accordance with aspects of this disclosure. For example, attention gates can be used to configure a machine learning model to give higher weight to aspects of the data which are more likely to be relevant to correctly generated outputs. As such, and because a machine learning model configured with these attention gates (or mechanisms) utilizes aspects of the data that are more likely to be relevant to correctly generated outputs, the ultimate predictive accuracy of those machine learning models is improved.
[00144] The quality and makeup of the training dataset for a neural network can impact the performance of the neural network in its execution phase. Dataset filtering and outlier removal can be advantageously applied to the training of the neural networks for the various techniques of the present disclosure (e.g., for the prediction of final setups or intermediate staging, for mesh element labeling or a neural network for mesh in-filling, for tooth reconstruction, for 3D mesh classification, etc.), because dataset filtering and outlier removal may remove noise from the dataset. And while the mechanism for realizing an improvement is different than using attention gates, that ultimate outcome is that this approach allows for the machine learning model to focus on relevant aspects of the dataset, and may lead to improvements in accuracy similar to improvements in accuracy realized vis-a-vis attention gates. [00145] In the case of a neural network configured to predict a final setup, a patient case may contain at least one of a set of segmented tooth meshes for that patient, a mal transform for each tooth, and/or a ground tmth setup transform for each tooth. In the case of a neural network to predict a set of intermediate stage setups, a patient case may contain at least one of a set of segmented tooth meshes for that patient, a mal transform for each tooth, and/or a set of ground truth intermediate stage transforms for each tooth. In some implementations, a training dataset may exclude patient cases which contact passive stages (i.e., stages where the teeth of an arch do not move). In some implementations, the dataset may exclude cases where passive stages exist at the end of treatment. In some implementations, a dataset may exclude cases where overcrowding is present at the end of treatment (i.e., where the oral care provider, such as an orthodontist or dentist) has chosen a final setup where the tooth meshes overlap to some degree. In some implementations, the dataset may exclude cases of a certain level (or levels) of difficulty (e.g., easy, medium and hard).
[00146] In some implementations, the dataset may include cases with zero pinned teeth (or may include cases where at least one tooth is pinned). A pinned tooth may be designated by a technician as they design the treatment to stop the various tools from moving that particular tooth. In some implementations, a dataset may exclude cases without any fixed teeth (conversely, where at least one tooth is fixed). A fixed tooth may be defined as a tooth that shall not move in the course of treatment. In some implementations, a dataset may exclude cases without any pontic teeth (conversely, cases in which at least one tooth is pontic). A pontic tooth may be described as a “ghost” tooth that is represented in the digital model of the arch but is either not actually present in the patient’ s dentition or where there may be a small or partial tooth that may benefit from future work (such as the addition of composite material through a dental restoration appliance). The advantage of including a pontic tooth in a patient case is to leave space in the arch as a part of a plan for the movements of other teeth, in the course of orthodontic treatment. In some instances, a pontic tooth may save space in the patient’s dentition for future dental or orthodontic work, such as the installation of an implant or crown, or the application of a dental restoration appliance, such as to add composite material to an existing tooth that is too small or has an undesired shape.
[00147] In some implementations, the dataset may exclude cases where the patient does not meet an age requirement (e.g., younger than 12). In some implementations, the dataset may exclude cases with interproximal reduction (IPR) beyond a certain threshold amount (e.g., more than 1.0 mm). The dataset to train a neural network to predict setups for clear tray aligners (CTA) may exclude patient cases which are not related to CTA treatment. The dataset to train a neural network to predict setups for an indirect bonding tray product may exclude cases which are not related to indirect bonding tray treatment. In some implementations, the dataset may exclude cases where only certain teeth are treated. In such implementations, a dataset may comprise of only cases where at least one of the following are treated: anterior teeth, posterior teeth, bicuspids, molars, incisors, and/or cuspids. [00148] The mesh comparison module may compare two or more meshes, for example for the computation of a loss function or for the computation of a reconstruction error. Some implementations may involve a comparison of the volume and/or area of the two meshes. Some implementations may involve the computation of a minimum distance between corresponding vertices/faces/edges/voxels of two meshes. For a point in one mesh (vertex point, mid-point on edge, or triangle center, for example) compute the minimum distance between that point and the corresponding point in the other mesh. In the case that the other mesh has a different number of elements or there is otherwise no clear mapping between corresponding points for the two meshes, different approaches can be considered. For example, the open-source software packages CloudCompare and MeshLab each have mesh comparison tools which may play a role in the mesh comparison module for the present disclosure. In some implementations, a Hausdorff Distance may be computed to quantify the difference in shape between two meshes. The open-source software tool Metro, developed by the Visual Computing Lab, can also play a role in quantifying the difference between two meshes. The following paper describes the approach taken by Metro, which may be adapted by the neural networks applications of the present disclosure for use in mesh comparison and difference quantification: "Metro: measuring error on simplified surfaces" by P. Cignoni, C. Rocchini and R. Scopigno, Computer Graphics Forum, Blackwell Publishers, vol. 17(2), June 1998, pp 167-174.
[00149] Some techniques of this disclosure may incorporate the operation of, for one or more points on the first mesh, projecting a ray normal to the mesh surface and calculating the distance before that ray is incident upon the second mesh. The lengths of the resulting line segments may be used to quantify the distance between the meshes. According to some techniques of this disclosure, the distance may be assigned a color based on the magnitude of that distance and that color may be applied to the first mesh, by way of visualization.
[00150] The setups prediction techniques described herein may generate a transform to place a tooth in a setup pose. Such a predicted transform may entail both the position and the orientation of the tooth, which is a significant improvement over existing techniques which use one neural network to generate a position prediction and another neural network to generate a pose prediction. In setups prediction, the predicted position and the predicted orientation affect each other. Generating the predicted position and the predicted orientation substantially concurrently offers improvements in predictive accuracy relative to generating predicted position and predicted orientation separately (e.g., predicting one without the benefit of the other).
[00151] The MLP Setups, VAE Setups, and Capsule Setups models of the present disclosure improve upon existing techniques with the addition of (among other things) a latent space input: either the latent space vector A of an oral care mesh or the latent capsule T of an oral care mesh. Prior setups prediction techniques did not train a reconstruction autoencoder to generate representations of teeth, and therefore could not verify the correctness of their outputs. The advantage of using a reconstruction autoencoder to generate tooth representations is that the latent representation (e.g., A or T) may be reconstructed by the reconstruction autoencoder. Reconstruction error (as described herein) may be computed, to demonstrate the correctness of the latent encoding (e.g., to demonstration that the latent representation correctly describes the shape and/or structure of the tooth). Results with a high reconstruction error may be excluded from downstream (e.g., further or additional) processing, which leads to a more accurate system as a whole. Either or both of A and T may be reconstructed (via a decoder) into a facsimile of an inputted oral care 3D representation (e.g., an inputted tooth mesh). One or more latent space vectors A (or latent capsules T) may be provided to the MLP Setups model. One or more latent space vectors A (or latent capsules T) may also be provided to the VAE Setups model. One or more latent capsules T (or latent vectors A) may also be provided to the Capsule Autoencoder Setups model.
[00152] The latent space vector A (or latent capsule T) for a tooth mesh (which may comprise thousands of interconnected mesh elements) describes the reconstruction characteristics of that tooth mesh in a compact form, for example a vector of length N (e.g., where in one example N = 128). This latent space vector A (or latent capsule T) may be reconstmcted into a close facsimile of the input tooth mesh through the operation of a decoder that has been trained for that task. The latent space vector A (or latent capsule T) is powerful because, although A (or T) is relatively extremely compact, A (or T) describes sufficient characteristics of the inputted oral care mesh (e.g., tooth mesh) to enable such a reconstruction of that oral care mesh (e.g., tooth mesh). In some implementations, the latent space vector A (or latent capsule T) can be used as an additional input to predictive or generative models of this disclosure. The latent space vector A (or latent capsule T) can be used as an additional input to at least one of an MLP, an encoder, a transformer, a regularized autoencoder, or a VAE of this disclosure. The latent space vector A (or latent capsule T) can be used as an input to the GDL Setups model described in the present disclosure. Furthermore, the latent space vector A (or latent capsule T) can be used as an input to the RL Setups model described in the present disclosure. The advantage of training a setups prediction neural network to take a latent space vector A (or latent capsule T) as an input is to provide information about the reconstruction characteristics of the tooth mesh to the network. Reconstruction characteristics may contain information about local and/or global attributes of the mesh. Reconstruction characteristics may include information about mesh structure. Information about shape may, in some instances, be included. An awareness of these reconstruction characteristics may better enable the trained setups prediction model to predict a final setup or intermediate staging, thereby providing the technical improvement of improved data precision. A further advantage of using the latent space vector A (or latent capsule T) is the vector’s size. A neural network may encode an understanding of the input mesh and pose data more resource-efficiently if those data are presented in a compact form (such as a vector of 128 real values), as opposed to inputting the full mesh (which may contain thousands of mesh elements). The latent representation of a mesh (or multiple meshes) may provide a more favorable signal-to-noise ratio than the original form of that mesh or those meshes, thereby improving the capability of a subsequent ML model (such as a neural network or SVM) to form predictions, draw inferences, and/or otherwise generate outputs (such as transforms or meshes) based on the input mesh(es). [00153] FIG. 2 shows how some of various setups prediction models can take as input either 1) tooth meshes or 2) latent space vectors (or latent capsules) which represent tooth meshes in reduced- dimensionality form.
[00154] This portion of the disclosure is directed to a reinforcement learning machine learning model to place an oral care mesh (e.g., such as placing one or more teeth to predict a final setup or to predict a series of intermediate stages that implement orthodontic treatment). The soft actor critic (SAC) model may be applied to the creation of final setups.
[00155] The machine learning technique of reinforcement learning may train a decision-making agent through a system of rewarding desired behavior and/or punishing undesirable behavior. The agent may be able to generate actions and learn through a process of interacting with its environment.
[00156] The RL Setups model is a novel adaptation of the Soft Actor Critic (SAC) reinforcement learning algorithm, which is trained to generate final setups and/or intermediate stages for orthodontic treatment. The RL Setups model has at least 3 main components: the RL environment, the RL neural networks (the decision-making entities, of which there are 5 or more) and the replay buffer (which stores the accumulated results of past executions of the algorithm). FIG. 6 describes five of the neural networks which may contribute to the training method of FIG. 3. Other implementations of SAC may be trained on other kinds of 3D oral care representations (e.g., hardware or appliance components), to place those oral care representations relative to one or more other oral care representations. In the example of setups prediction, a transform may be predicted for a tooth to place that tooth in a pose relative to at least one other tooth. In the example of hardware placement, a transform may be predicted for a hardware element to place that hardware element relative to a tooth. In other examples, a transform may be predicted to place an appliance or an appliance component (such as a library component) relative to one or more teeth or one or more other appliance components.
[00157] The Soft Actor Critic (SAC) network architecture may contain 5 (or more) neural networks, a replay buffer function 300 and an environment 336. The remainder of the specification in this section (“SAC Neural Network Architecture”) applies to some implementations. SAC is a form of Q-Leaming. Actor Network: a policy network which may predict an action based on a state.
Critic Networks: two q-value networks - a minimum value predicted by these network may be used to update, at least in part, the policy and value functions.
Value Networks : two or more networks which may predict how valuable is a state - the target value network may be softly updated based on the main value network.
Optimization Algorithm: ADAM (alternatives are described in “Neural network architecture”) Feedforward: ReLu nonlinear function (alternatives are described in “Neural network architecture”) Activation Function: Tanh (alternatives are described in “Neural network architecture”) Replay Buffer: may save tuples containing at least one of state, action, reward, and next state. Maximum Entropy: may encourage exploration by randomly selecting actions.
Reparameterization: may be used to update the Policy Actor network and may reduce problems with backpropagating errors. Loss: simple mean square error (MSELoss).
Weights and Bias: Network weights may be modified in small steps (e.g., ± 3e-3).
Policy Network: When training the policy actor network, a standard deviation with minimum of -20 and maximum of 2 is defined. Other implementations may follow other minimums and maximums, for example to avoid negative values and/or reduce the range of values (e.g., [-5,2] and [0.000001,1] respectively).
Environment: The environment may be considered to be the universe in which the agent operates (e.g., an orthodontic setup). The environment may contain a module that applies movement (rotation, translation), may compute one or more rewards, may provide action samples, may define data dimensions and, may flag if/when the setup/goal may be achieved during training.
Neural Network Architecture: Each of the 5 neural networks may, in some implementations, be implemented using one or more MLPs consisting of linear layers and ReLu activation functions. Other possible neural networks and/or neural network components are disclosed elsewhere in this disclosure. Agent: The agent may be interpreted to be who/what performs/applies actions to a state and then who may receive the response (e.g., transformations to be applied to teeth, such as translations and/or rotations, after one or more actions are applied and one or more rewards are issued) from the environment. The actor’s experience may be recoded in a loop of steps taken, where each step may correspond to an action which may be applied to one or more teeth.
Rewards: A reward may be assigned after each action is applied.
Action Sample: An action sample may, in some implementations, be a 232-dimensional vector (other dimensions are possible). An example action sample is: [upper arch form, upper arch teeth translation and rotation, lower arch-form, lower arch teeth translation and rotation]
State: A state may in some implementations be a 232-dimensional vector (or a vector of another size), and may have some of the same characteristics as an action, [upper arch form, upper arch teeth translation and rotation, lower arch-form, lower arch teeth translation and rotation]
Action applied to State: Applying an action to a state may be interpreted to mean that each tooth is transformed, based, at least in part, on the current state of the tooth.
[00158] In the context of setups generation, using the SAC training paradigm may provide several advantages. SAC may seek to maximize a policy’s entropy (e.g., which may involve a stochastic element to the moving of teeth in order to achieve a setup). SAC may, in some implementations, be an off-policy algorithm. SAC may in some implementations be used for intermediate staging, in which case motion limits may be advantageous (i.e., limits on rotation magnitude and/or translation magnitude).
[00159] FIG. 3 shows an example training method for a reinforcement learning model to predict transforms for 3D oral care representations (e.g., such as placing teeth for setups prediction). At the start of a round of training for the RL Setups algorithm, a tuple 302 may be received by the agent 304. A tuple 302 may comprise {a current state, an action, a next state, a reward, and a ‘done’ flag}, though other tuples are possible in other implementations. The state may have one or more vectors which represent each tooth in each of the two arches. Within one of such vectors, a vector element may contain real numbers which may describe the translation and rotation transformations for one or more teeth (e.g., as translation vector and/or quaternion, or as transformation matrices). The state may also, in some implementations, have an archform (e.g., as encoded by a set of control points or otherwise described elsewhere in this disclosure).
[00160] The Value Network 328 and the Target Value Network 308 comprise two of the five (or more) neural networks. The 2 value networks may compute values for the current state 310 and next state (aka new state) 312 that may be used in updating at least one of the critic loss 316 or the value loss 320. State data (e.g., current state 310 and next state 312) from the tuples may be provided to these two neural networks, and the networks may provide the corresponding outputs (e.g., the state value and the target value tuple 318) to the calculation of losses: the value loss (used in updating the Value Network 328 and Target Value Network 308) and the critic loss (used in updating the 2 Critic Networks 334 and 306).
[00161] The purpose of the two critic networks may be to compute predicted Q-values 314 which may be used in updating the policy loss and the critic loss. A tuple 338 comprising the current state and the predicted action (provided by the Actor Policy Network) may be provided to each of the two critic networks, and each critic network may generate a predicted Q-value 314. The minimum of these two predicted Q-values may be provided to the calculation of the critic loss 316. In some implementations, a Q-value may map state and action information to one or more values which represent an expected longterm reward.
[00162] The RL techniques described herein may, in some implementations, learn an optimal policy which optimizes (e.g., maximizes) long-term cumulative rewards. The purpose of the Actor Policy Network 330 may be to generate a predicted action 324. The Actor Policy Network (APN) 330 may also generate a log probability 322 (i.e., the log probability of a state vector, after the state vector may have undergone a ReLU operation), which may be used in the calculation of the policy loss 326, which may, in turn, be used in training the Actor Policy Network 330. At the completion of a round of training, after each of the five (or more) neural networks may have been updated by its respective loss function, a new tuple may be outputted to the environment 336. This new tuple 332 may encapsulate the learning which took place during that round of training and may include any reward which accrued to the agent from this latest round of execution. The new tuple may be received by the replay buffer 300, which may store one or more tuples and which may provide samples of tuples to the neural networks during future rounds of training. The Actor Policy Network 330 comprises the neural network that may ultimately be trained by the training process and may get deployed for use in an operational RL Setups prediction model.
[00163] The APN may take the current state as input and may output an action (i.e., vector of transformations for the teeth to place the teeth into their setup poses). This action may be provided to the Environment, which may then proceed to generate a new state (i.e., which apply the transformations to the one or more teeth of the one or more aches). This process may iterate until a 'done' criterion is reached. In some implementations, a 'done' criterion may be reached when one or more metrics (as described elsewhere in this disclosure) show that the one or more arches reflect tooth positions which are acceptable for use in a setup (such as a final setup or intermediate stage). In some implementations, one or more of the iterations of the training procedure described in FIG. 3 (and further illustrated in FIG. 5) may generate actions which are used to form an incremental set of movements for one or more teeth of the one or more arches, where such incremental movements define a stage. In some implementations, motion limits may be imposed, so that a given tooth does not move too far (in terms of translation and/or rotation) for a given stage.
[00164] Each patient case may be represented by a vector (the size of which may vary depending on the contents of the vector). The vector therefore may contain one or more arch forms, one or more teeth (for example 28 teeth - in the case that upper and lower 8s, namely, wisdom teeth, are excluded), translation information and/or rotation information. One example of such a vector may be of 232 dimensions: [upper arch form, upper arch teeth translation and rotation, lower arch-form, lower arch teeth translation and rotation]
[00165] The environment (shown in FIG. 4) may represent the ‘world’ of the task of transform generation (e.g., for orthodontic setups transform generation, or appliance component transform generation). In the aligners treatment solution, the environment may represent the arches of teeth in orthodontic treatment, where every predicted action may be applied to one or more tooth, which may result in the returning of one or more new states, rewards and/or done flags (i.e., to signify that the algorithm has reached completion).
[00166] The training method may involve a number of episodes that run multiple steps. Every step may consist of applying an action (either predicted by the policy network or a sample action) to a state and which may yield a new state and/or a reward from the environment. These tuples (of state, action, reward, new state, done flag) may then be pushed to the replay buffer that further is used to update the critic, value and policy networks.
[00167] The goal is to maximize rewards and reduce network losses. Specifically, a reward may influence the critic network and, by consequence, at least one of the value and policy networks. FIG. 6 provides an illustration of various neural networks trained for use in the reinforcement learning methods of this disclosure (e.g., RL Setups). Some implementations may incorporate other neural networks in place of or in addition to the neural networks shown in FIG. 6, such as one or more of the neural networks listed elsewhere in this disclosure.
[00168] Actions: The Environment may be interpreted as the place where actions are applied to the arch state (i.e., each tooth has a state, and translation and/or rotation may be applied to a tooth state). An RL arch vector may comprise one or more (e.g., 28) tooth transforms, and a list of archform control points (or some other representation of the archform). There may be lower and upper arches. Actions may be applied to one or both arches. An action may be a tooth movement (e.g., involving at least one of a translation and a rotation). As an action is applied, the environment may return at least three data points: an arch state, a reward, and a flag indicating whether the final state has been reached.
[00169] In some implementations, the dataset may be mesh-free, and may operate entirely on tooth transforms (e.g., transformation matrices). In other implementations, the dataset may involve tooth meshes. In some implementations, the input vectors may be composed of tooth translation and/or rotation information (e.g., in the form of translation vectors and quaternions, or in the form of transformation matrices which entail at least one of translations and rotations), and an archform (i.e., a set of points that defines the form of the arch - for example a spline with 6 control points in some examples). In some implementations, the starting state of a patient case may comprise a translation vector and a quaternion for each tooth, such as to define the maloccluded poses of the teeth.
[00170] An action may comprise of a tooth transformation, which may have translation and/or rotation components. Action samples may be used to train the neural network models. An action sample may, for example, be a set of tooth movements between the states of a given patient case (e.g., embodied as a vector that represents translation and rotation for each tooth in each arch - using quaternions and translation vectors). Multiple teeth may move in correspondence to each action sample. Action samples may be provided to the training environment to drive learning.
Each of the five neural networks is described below.
[00171] An Actor Policy Network (APN) may predict an action based on the arch state. Input may be a state, and output may be an action. Two critic networks (with the same goal) may each predict a Q- value. Input may be an arch state and/or an action. Each may generate one or more real valued Q-value outputs. The minimum of these two values may ultimately be outputted. Two value networks (with the same architecture) may each implement a goal to compute the value of an arch state. Input to one network may be the current arch state. Input to the other network may be the new or predicted (next) arch state. Each network may generate one (or more) real-value outputs. Both of these values are used in computing the value loss. One or both of these values may also used in computing the critic loss. The current state output value may be used to update the target value network. The reward, in the form of a “done” flag, and the discount factor (gamma) may be used to transform the target value into the final target value. The final target value may be used to compute one or both of the critic loss and the value loss. The purpose of the policy loss may be to update the actor policy network. The purpose of the actor policy network may be to predict an action.
[00172] Additional details on RL model training are described below. One or more vectors that defines the current case state may be provided, and one or more action samples may be provided (e.g., to train the five or more neural networks on how to receive input in the form of a case state and move the teeth into setup poses). The goal state may be defined, at least in part, by an example of a ground truth setup. In some implementations, loss may be computed as a difference between a predicted action and a ground tmth action (e.g., a quantification of the difference in tooth transformations). Such a loss may, in some implementations, be used to train one or more neural networks in the RL Setups model.
[00173] Episodes and steps may, in some implementations, describe portions of the training process. An episode may be a group of actions which may be implemented over N steps (e.g., ten steps). Execution may go through a predetermined count of action (e.g., ten actions) to attempt to reach a desired setup configuration (e.g., either final setup or intermediate stage). A step may be viewed as the application of an action. If there are ten steps, then there may be ten corresponding actions. In one example, there may be five episodes. For three of the episodes, the agent provides the actions from ground tmth action samples. For the remaining two episodes, the model generates the actions. The agent may be the logic which is applied to train the five (or more) neural network models.
[00174] The five (or more) neural network models may be trained via backpropagation. Loss may, in some implementations, be computed as a function of rewards received. Loss may be a real number. The reward may be computed by the environment as a mean squared error (MSE) between the generated state and the goal state (i.e., by way of a vector comparison). MSE may be computed on two vectors. A state may contain the translation and/or rotation information for one or more teeth for one or both arches, alongside the archform information. The rewards (and associated neural network loss values) may be used to train, at least in part, the five (or more) neural network models via backpropagation.
[00175] The replay buffer may store tuples having a {state, action, reward, flag to indicate whether goal was achieved, new predicted state} structure. Such tuples may comprise samples of what has been accomplished via past iterations of learning (training passes). Every time an action is applied, a tuple may be created and stored in the replay buffer. When the replay buffer is full (e.g., at 200 samples), a portion of the samples may be used to update one or more of the neural networks. The replay buffer may be used indirectly in loss calculation. The replay buffer may provide samples that may be used as inputs to the neural networks. The frequency of updates to the neural networks may be dependent on the size of the replay buffer. From time during training, the agent may request tuples from the replay buffer to drive learning.
[00176] Techniques of this disclosure may provide additional inputs to RL setups as well. In some implementations of RL Setups (such as in implementations that use the SAC algorithm), one or more of the input values from one or more of the following list of input vector may be introduced to the RL Setups model: B, K, L, M, N, O, R, S, P, Q, U, V. In some implementations, one or more of the inputs B, K, L, M, N, O, R, S, P, Q, U, V may be provided to one or more of the neural networks of RL Setups: Actor Policy Network, Value Network, Target Value Network and both Critic Networks (i.e., in addition to the state information and/or predicted action information which various of these networks may also receive as input, according to various implementations).
[00177] For example, in some implementations, one or more of the inputs B, K, L, M, N, O, R, S, P, Q, U, V may be received by the Actor Policy Network (APN), along with the input of the current state, with the advantage of better enabling the APN to generate actions (e.g., incremental tooth transforms) which produce effective setups predictions which meet the needs of patients (e.g., to generate an effective final setup or intermediate stage).
[00178] It will be appreciated that various modifications to the algorithm(s) can be used in accordance with aspects of this disclosure as well. Some implementations may adapt the RL Setups technique described above to perform operations on 3D oral care representations, such as tooth meshes, gums, hardware, appliances and appliance components (e.g., prefab library components). Such appliances and appliance components may pertain to dental restoration, and may be used to shape dental composites to produce veneers. [00179] In such implementations, representation learning may be used in conjunction with reinforcement learning to generate transforms. Such transform may, in some implementations, be applied to tooth for setups prediction. In such examples, the APN may contain at least a first machine learning model and a second machine learning model (e.g., potentially in addition to other machine learning models). The current state and/or the next state may maintain information about the tooth meshes. A first ML model (e.g., a U-Net, autoencoder, a pyramid encoder-decoder, a 3D encoder, or a series of convolution and/or pooling layers) may be used to generate representations of the teeth. A second ML model (e.g., an encoder, a transformer, an autoencoder, an MLP or a network comprising one or more global average pooling layers) may be trained to map those representations of the teeth (e.g., which may contain a representation vector corresponding to each mesh element) to one or more vectors of transformation values for one or more teeth (e.g., the APN may map the tooth representations into transforms which may place those teeth into setups poses).
[00180] Such implementations may generate tooth transforms which may be applied to corresponding tooth meshes. In the course of execution of this modified version of the above RL Setups technique, collision detection may be performed on the tooth meshes through the course of training and deployment. These techniques of the present disclosure may incorporate such collision detection between meshes into the calculation of loss values, which may be used to train, at least in part, the neural networks of the RL Setups technique, such as the APN. The resulting APN may generate tooth transforms which seek to mitigate or reduce or potentially even avoid collisions between transformed tooth meshes. Such an RL model may incorporate mesh collision mitigate or collision avoidance measures as described in US patent application with Publication No. US20220054234A1.
[00181] In terms of policy, various RL-based techniques of this disclosure may implement one or more of a state-action-reward-state-action (S ARSA), which is a policy-based training method. The S ARSA policy gives information about the probability that a certain tooth transformation or pose may be advantageous. Examples of policy -free training methods include Q-learning (in which the agent for adjusting the poses of teeth explores the environment in a self-directed method) and/or Deep Q-learning (which is an implementation of Q-Leaming using neural networks).
[00182] Other RL-based approaches that are in accordance with aspects of this disclosure are described below. In some implementations, the following RL algorithms may be trained to be applied in whole, in part or in modified form to the setups generation task. Some implementations of these RL models for use in RL Setups generation may incorporate one or more of the neural networks described elsewhere in this disclosure.
1. Multiobjective RL: for taking into account the refinement penalty, extended treatment duration penalty, and biological constraints
2. Autoregressive generative model: for modeling the temporal transition of stages, and propose a diverse set of plans
3. POMDP: modeling the unobserved staging results
4. Offline contextual bandit: train the autoregressive model with offline expert data
5. Imitation learning to learn from experts 6. Safe RL: utilizing collisions as a risk metric and using one of the methods that balances risk and reward to identify high reward staging that simultaneously limits collisions, thus increasing speed of staging.
7. A2C / A3C (Asynchronous Advantage Actor-Critic)
8. PPO (Proximal Policy Optimization)
9. TRPO (Trust Region Policy Optimization)
10. DDPG (Deep Deterministic Policy Gradient)
11. ID 3 (Twin Delayed DDPG)
12. DQN (Deep Q-Networks)
13. C51 (Categorical 51 -Atom DQN)
14. QR-DQN (Quantile Regression DQN)
15. HER (Hindsight Experience Replay)
16. World Models
17. 12 A (Imagination- Augmented Agents)
18. MBMF (Model-Based RL with Model-Free Fine-Tuning)
19. MB VE (Model-Based Value Expansion)
20. AlphaZero
[00183] The RL techniques described herein may, in some implementations, use imitation learning. In such implementations, an expert may provide input to help train the model in the form of sparse rewards (e.g., a manually specified reward function). An agent may, in some implementations, attempt to learn an optimal policy by following the decisions of the expert. An example of Imitation Learning that techniques of this disclosure may incorporate is described as follows. Generative approaches using Imitation learning may be used in the generation of intermediate staging. For instance, the Generative Adversarial Imitation Learning (GAIL) algorithm is an effective method to improve the quality of setups (final and/or staging), reduce the expected number of required refinement stages (i.e., stages requiring rework after aligners have already been made); and decrease the count of stages required for treatment (leading to fewer aligners). GAIL is an algorithm which may learn to “mimic” an expert’s behavior using one or more exemplary ground truth demonstrations (e.g., from an expert). Conceptually, GAIL may harness the so-called “generative adversarial training” concept to fit a distribution over states and actions visited by the expert. Ground truth setups (whether for final setups or for intermediate staging) and/or the generated (predicted) setups (final or intermediate staging) from a policy network (the generator network) may be provided to a discriminator network. The discriminator may be implemented as a binary -classifier which classifies each input as 0 or 1 (i.e., whether the input is from a ground truth source or from an artificial intelligence-based algorithm). Upon successfully training both the generator and discriminator networks, the generator (after achieving convergence) will eventually be able how to generate setups (final or intermediate staging) that are indistinguishable from those provided as ground truth, even by a highly trained discriminator network.
[00184] Techniques described herein may be trained to generate transforms which may place the patient’s teeth into poses suitable for use in orthodontic setups (e.g., intermediate stages or final setups), according to the specification of the oral care arguments which may, in some implementations, be provided to the generative model. Oral care arguments may include oral care parameters as disclosed herein, or other real-value, text-based or categorical inputs which specify intended aspects of the one or more 3D oral care representations which are to be generated. In some instances, oral care arguments may include oral care metrics, which may describe intended aspects of the one or more 3D oral care representations which are to be generated. Oral care arguments are specifically adapted to the implementations described herein. For example, the oral care arguments may specify the intended designs (e.g., including shape and/or structure) of 3D oral care representations which may be generated (or modified) according to techniques described herein. In short, implementations using the specific oral care arguments disclosed herein generate more accurate 3D oral care representations than implementations that do not use the specific oral care arguments.
[00185] In some instances, a text encoder may encode a set of natural language instructions from the clinician (e.g., generate a text embedding). A text string may comprise tokens. An encoder for generating text embeddings may, in some implementations, apply either mean-pooling or max-pooling between the token vectors. In some instances, a transformer (e.g., BERT or Siamese BERT) may be trained to extract embeddings of text for use in digital oral care (e.g., by training the transformer on examples of clinical text, such as those given below). In some instances, such a model for generating text embeddings may be trained using transfer learning (e.g., initially trained on another corpus of text, and then receive further training on text related to digital oral care). Some text embeddings may encode text at the word level. Some text embeddings may encode text at the token level. A transformer for generating a text embedding may, in some implementations, be trained, at least in part, with a loss calculation which compares predicted outputs to ground truth outputs (e.g., softmax loss, multiple negatives ranking loss, MSE margin loss, cross-entropy loss or the like). In some instances, the non-text arguments, such as real values or categorical values, may be converted to text, and subsequently embedded using the techniques described herein. The following are examples of natural language instructions that may be issued by a clinician to the generative models described herein: “Generate a setup to set to Class I molar and canine, 2 mm overbite and add 2mm of expansion 5-5Z5-5”, “Generate a setup to align with proclination and expansion and finish with .5 mm spaces U2-2 for future restorative”, or “Adjust the setup with no second molar movement, rotate upper first molars mesial out for Class I, level lower to a reverse curve of Spee 2 mm, advance the mandible with elastics to Class I canine.” [00186] In some instances, a local coordinate system for a 3D oral care representation, such as a tooth, may be described by one or more transforms (e.g., an affine transformation matrix, translation vector or quaternion). Systems of this disclosure may be trained for coordinate system prediction using past cohort patient case data. The past patient data may include at least: one or more tooth meshes or one or more ground truth tooth coordinate systems. Machine learning models such as: U-Nets, encoders, autoencoders, pyramid encoder-decoders, transformers, or convolution and/or pooling layers, may be trained for coordinate system prediction. Representation learning may determine a representation of a tooth (e.g., encoding a mesh or point cloud into a latent representation, for example, using a U-Net, encoder, transformer, convolution and/or pooling layers or the like), and then predict a transform for that representation (e.g., using a trained multilayer perceptron, transformer, encoder, transformer, or the like) that defines a local coordinate system for that representation (e.g., comprising one or more coordinate axes). In the instance where the coordinate system is predicted for a tooth mesh, the mesh convolutional techniques described herein can leverage invariance to rotations, translations, and/or scaling of that tooth mesh to generate predications that techniques that are not invariant to the rotations, translations, and/or scaling of that tooth mesh cannot generate. Pose transfer techniques may be trained for coordinate system prediction, in the form of predicting a transform for a tooth. Reinforcement learning techniques may be trained for coordinate system prediction, in the form of predicting a transform for a tooth. [00187] Machine learning models such as: U-Nets, encoders, autoencoders, pyramid encoderdecoders, transformers, or convolution and/or pooling layers, may be trained as a part of a method for hardware (or appliance component) placement. Representation learning may train a first module to determine an embedded representation of a 3D oral care representation (e.g., encoding a mesh or point cloud into a latent form using an autoencoder, or using a U-Net, encoder, transformer, block of convolution and/or pooling layers or the like). That representation may comprise a reduced dimensionality form and/or information-rich version of the inputted 3D oral care representation. In some implementations, the generation of a representation may be aided by the calculation of a mesh element feature vector for one or more mesh elements (e.g., each mesh element). In some implementations, a representation may be computed for a hardware element (or appliance component). Such representations are suitable to be provided to a second module, which may perform a generative task, such as transform prediction (e.g., a transform to place a 3D oral care representation relative to another 3D oral care representation, such as to place a hardware element or appliance component relative to one or more teeth) or 3D point cloud generation. Such a transform may comprise an affine transformation matrix, translation vector or quaternion or the like.
[00188] Machine learning models which may be trained to predict a transform to place a hardware element (or appliance component) relative to elements of patient dentition include: MLP, transformer, encoder, or the like. Systems of this disclosure may be trained for 3D oral care appliance placement using past cohort patient case data. The past patient data may include at least: one or more ground truth transforms and one or more 3D oral care representations (such as tooth meshes, or other elements of patient dentition). In the instance where a U-Net (among other neural networks) is trained to generate the representations of tooth meshes, the mesh convolution and/or mesh pooling techniques described herein leverage invariance to rotations, translations, and/or scaling of that tooth mesh to generate predications that techniques that are not invariant to the rotations, translations, and/or scaling of that tooth mesh cannot generate. Reinforcement learning techniques may be trained for hardware or appliance component placement.
[00189] Techniques of this disclosure may, in some implementations, use PointNet, PointNet++, or derivative neural networks (e.g., networks trained via transfer learning using either PointNet or PointNet++ as a basis for training) to extract local or global neural network features from a 3D point cloud or other 3D representation (e.g., a 3D point cloud describing aspects of the patient’s dentition - such as teeth or gums). Techniques of this disclosure may, in some implementations, use U-Nets to extract local or global neural network features from a 3D point cloud or other 3D representation. [00190] 3D oral care representations are described herein as such because 3-dimensional representations are currently state of the art. Nevertheless, 3D oral care representations are intended to be used in a non-limiting fashion to encompass any representations of 3 -dimensions or higher orders of dimensionality (e.g., 4D, 5D, etc.), and it should be appreciated that machine learning models can be trained using the techniques disclosed herein to operate on representations of higher orders of dimensionality.
[00191] In some instances, input data may comprise 3D mesh data, 3D point cloud data, 3D surface data, 3D polyline data, 3D voxel data, or data pertaining to a spline (e.g., control points). An encoderdecoder structure may comprise one or more encoders, or one or more decoders. In some implementations, the encoder may take as input mesh element feature vectors for one or more of the inputted mesh elements. By processing mesh element feature vectors, the encoder is trained in a manner to generate more accurate representations of the input data. For example, the mesh element feature vectors may provide the encoder with more information about the shape and/or structure of the mesh, and therefore the additional information provided allows the encoder to make better-informed decisions and/or generate more-accurate latent representations of the mesh. Examples of encoder-decoder structures include U-Nets, autoencoders or transformers (among others). A representation generation module may comprise one or more encoder-decoder structures (or portions of encoders-decoder structures - such as individual encoders or individual decoders). A representation generation module may generate an information-rich (optionally reduced-dimensionality) representation of the input data, which may be more easily consumed by other generative or discriminative machine learning models.
[00192] A U-Net may comprise an encoder, followed by a decoder. The architecture of a U-Net may resemble a U shape. The encoder may extract one or more global neural network features from the input 3D representation, zero or more intermediate-level neural network features, or one or more local neural network features (at the most local level as contrasted with the most global level). The output from each level of the encoder may be passed along to the input of corresponding levels of a decoder (e.g., by way of skip connections). Like the encoder, the decoder may operate on multiple levels of global-to-local neural network features. For instance, the decoder may output a representation of the input data which may contain global, intermediate or local information about the input data. The U-Net may, in some implementations, generate an information-rich (optionally reduced-dimensionality) representation of the input data, which may be more easily consumed by other generative or discriminative machine learning models.
[00193] An autoencoder may be configured to encode the input data into a latent form. An autoencoder may train an encoder to reformat the input data into a reduced-dimensionality latent form in between the encoder and the decoder, and then train a decoder to reconstruct the input data from that latent form of the data. A reconstruction error may be computed to quantify the extent to which the reconstructed form of the data differs from the input data. The latent form may, in some implementations, be used as an information-rich reduced-dimensionality representation of the input data which may be more easily consumed by other generative or discriminative machine learning models. In most scenarios, an autoencoder may be trained to input a 3D representation, encode that 3D representation into a latent form (e.g., a latent embedding), and then reconstruct a close facsimile of that input 3D representation at the output.
[00194] A transformer may be trained to use self-attention to generate, at least in part, representations of its input. A transformer may encode long-range dependencies (e.g., encode relationships between a large number of inputs). A transformer may comprise an encoder or a decoder. Such an encoder may, in some implementations, operate in a bi-directional fashion or may operate a self-attention mechanism. Such a decoder may: operate a masked self-attention mechanism; operate a cross-attention mechanism; or operate in an auto-regressive manner, according to particular implementations. The self-attention operations of the transformers described herein may, in some implementations, relate different positions or aspects of an individual 3D oral care representation in order to compute a reduced-dimensionality representation of that 3D oral care representation. The cross-attention operations of the transformers described herein may, in some implementations, mix or combine aspects of two (or more) different 3D oral care representations. The auto-regressive operations of the transformers described herein may, in some implementations, consume previously generated aspects of 3D oral care representations (e.g., previously generated points, point clouds, transforms, etc.) as additional input when generating a new or modified 3D oral care representation. The transformer may, in some implementations, generate a latent form of the input data, which may be used as an information-rich reduced-dimensionality representation of the input data, which may be more easily consumed by other generative or discriminative machine learning models.
[00195] In some implementations, an encoder-decoder structure may first be trained as an autoencoder. In deployment, one or more modifications may be made to the latent form of the input data. This modified latent form may then proceed to be reconstructed by the decoder, yielding a reconstructed form of the input data which differs from the input data in one or more intended aspects. Oral care arguments, such as oral care parameters or oral care metrics may be provided to the encoder, the decoder, or may be used in the modification of the latent form, to guide the encoder-decoder structure in generating a reconstructed form that has desired characteristics (e.g., characteristics which may differ from that of the input data).
[00196] Techniques of this disclosure may, in some instances, be trained using federated learning. Federated learning may enable multiple remote clinicians to iteratively improve a machine learning model (e.g., validation of 3D oral care representations, mesh segmentation, mesh cleanup, other techniques which involve labeling mesh elements, coordinate system prediction, non-organic object placement on teeth, appliance component generation, tooth restoration design generation, techniques for placing 3D oral care representations, setups prediction, generation or modification of 3D oral care representations using autoencoders, generation or modification of 3D oral care representations using transformers, generation or modification of 3D oral care representations using diffusion models, 3D oral care representation classification, imputation of missing values), while protecting data privacy (e.g., the clinical data may not need to be sent “over the wire” to a third party). Data privacy is particularly important to clinical data, which is protected by applicable laws. A clinician may receive a copy of a machine learning model, use a local machine learning program to further train that ML model using locally available data from the local clinic, and then send the updated ML model back to the central hub or third party. The central hub or third party may integrate the updated ML models from multiple clinicians into a single updated ML model which benefits from the learnings of recently collected patient data at the various clinical sites. In this way, a new ML model may be trained which benefits from additional and updated patient data (possibly from multiple clinical sites), while those patient data are never actually sent to the 3rd party. Training on a local in-clinic device may, in some instances, be performed when the device is idle or otherwise be performed during off-hours (e.g., when patients are not being treated in the clinic). Devices in the clinical environment for the collection of data and/or the training of ML models for techniques described herein may include intra-oral scanners, CT scanners, X- ray machines, laptop computers, servers, desktop computers or handheld devices (such as smart phones with image collection capability). In addition to federated learning techniques, in some implementations, contrastive learning may be used to train, at least in part, the ML models described herein. Contrastive learning may, in some instances, augment samples in a training dataset to accentuate the differences in samples from difference classes and/or increase the similarity of samples of the same class.
Reinforcement Learning by Human Feedback
[00197] According to some of the RL-based techniques of this disclosure, a reward model (RM) for setups prediction, also referred to as a preference model, may be generated which is calibrated or improved using human preferences. The preference model may be trained to output the term r omega, which may be a scalar value which may describe at least one aspect of preferability of a setup. A policy may be a setups prediction model (or alternatively, a model that predicts a transform for another type of oral care representation or a model which generates an oral care representation) which receives input pertaining to the ranking or quality of one or more setups from a clinician or technician (e.g., received by way of a prompt). The policy may output one or more transforms for one or more teeth, or in some implementations, may output transforms for other types of oral care representation or may output 3D oral care representations themselves. The action space of this policy may pertain to the space of possible transforms and the observation space may pertain to the distribution of possible transforms, which may be large. The reward function may pertain to a combination of the preference model and a constraint on policy shift.
[00198] FIG. 8 shows an example training method. Patient case data 802, including 3D representations of the patient’s teeth, malocclusion tooth transforms, and/or ground truth setup tooth transforms, may be provided to the model Fl 804 and model F2 806. A setups prediction model Fl 804 may be trained on a dataset of cohort patient cases in a supervised/semi-supervised fashion (e.g., GDL Setups, Transformer Setups, Autoencoder Setups, Pose Transfer Setups and the like). Such models may include ML components such as: U-Nets, encoders, decoders, pyramid encoder-decoders, transformers, autoencoders, multi-layer perceptrons (MLP), or the like. For example, a U-Net may be used to extract hierarchical neural network features from the tooth meshes of the patient’s arch and generate latent representations as embedding vectors. The embedding vectors may be provided to an MLP (e.g., four fully connected layers with optional skip connections), which may generate the setups transforms for the teeth. The performance of setup prediction model Fl 804 may be improved upon using reinforcement learning from human feedback (RLHF). RLHF may entail the providing of feedback from a human to train a reward model. Stated another way, RLHF may incorporate feedback from a human user into the data that are used to train a reward model. The trained reward model may be used as a reward function to optimize the policy of the agent, using reinforcement learning. Setup prediction model Fl 804 may be trained for tooth placement in the setups application or may also be trained more broadly to place 3D oral care representations other than teeth. For example, model Fl may place a 3D oral care representation relative to at least one other 3D oral care representation or relative to at least one global coordinate axis (e.g., to place a fixture model component or an appliance component relative to one or more teeth). [00199] A copy may be made of model Fl, referred to as setups prediction model F2 806. The training of Fl may be frozen to prevent further learning by setups prediction model Fl 804, whereas the further learning may be applied to setups prediction model F2 806 through RL optimizations (e.g., referred to herein as 'fine-tuning'). A software application, such as implemented using a web interface on a PC or mobile device, may enable feedback to be collected from a clinician or technician. The following process may undergo multiple iterations in the course of training model F2 through RLHF. A maloccluded arch of teeth 802 from the training dataset may be provided to model Fl 804, yielding a predicted setup Y1 812 as output. The same maloccluded arch of teeth may be provided to model F2 806, yielding a predicted setup Y2 814 as output. Model Fl 804 represents the original setups model, while model F2 806 may undergo “fine-tuning” or improvement with each iteration.
[00200] A third model, Fpreferencemodei 816 is defined which may enable grading/ranking/scoring of the predicted setup Y2. Fpreferencemodei 816 may in some instances have already undergone training. 816 may in some instances have undergone human-in-the-loop training or otherwise been trained (at least in part) on at least one aspect of a dataset of human-in-the-loop responses. Predicted setup Y2 may be provided to Fpreferencemodei 816, which may then output a notion of “preferability” (e.g., as a scalar or categorical value), called r omega. r omega may, in some instances, describe the fitness, suitability, or quality of the predicted setup predicted setup Y2, as assigned by a clinician or technician (e.g., using the software interface described above, or via other channels).
[00201] Predicted setup Y1 and predicted setup Y2 may then be compared using a technique such a scaled KL-Divergence loss or another of the loss calculation methods described herein. The output of this loss calculation is denoted by r_KL 818. Expressed as a formula, the output may be computed (818) according to the following: r_KL = KL(Y1,Y2). The KL-divergence loss may penalize the RL policy from moving substantially away from the initial pretrained model with each iteration of the RLHF training. This penalty may benefit training, by biasing the model to output setups which are coherent, valid setups, at least within an acceptable margin of error. Without the penalty imposed by the KL- Divergence term, the RLHF optimization algorithm may start to generate poor-quality setups which nonetheless manage to compute high rewards values (e.g., artificially high reward values which “fool” the RLHF algorithm to invoke a false positive). In some instances, a reward may be derived, at least in part, from a ranking supplied by a clinician. In some instances, a reward may be derived, at least in part, from a model performance score, such as an ADD 10 or ADD05 score. Other scores may be used, as long as the scores enable setups to be ranked. The KL-divergence term may be approximated via sampling from both the Y 1 and Y2 distributions. A reward for the RL optimization function (PPO), which may be used as a part of an update rule in training model F2, may be computed (820) as follows: r = r omega - scaler term * r_KL. The summed reward 820 may be used to update the policy 810.
[00202] Upon calculation of a reward, the F2 setups prediction model may undergo updating. One or more aspects of model F2 (e.g., neural network weights) may be updated (808), based at least in part on the reward. In some implementations, a policy -gradient RL algorithm may be used to train model F2, such as Proximal Policy Optimization (PPO). PPO may maximize the reward metric for a current batch of training data. PPO may, in some instances, impose constraints on the gradient calculation during the model weights updating step, to prevent destabilization of the learning process.
[00203] This updating operation completes an iteration of training by RLHF. Model F2 may, in some instances, continue to be trained until the performance of model F2 reaches a threshold, the performance of model F2 supersedes model Fl in some metric and/or until no further learning can otherwise be achieved.
[00204] Imitation Learning (which is one example of Reinforcement Learning) may be used to train an ML model to place a 3D oral care representation (e.g., a tooth for a clear tray aligner, an appliance component for a dental restoration appliance or a component for a fixture model). For example, a neural network may be trained using Imitation Learning to generate transforms to place 3D oral care representations (e.g., teeth for restoration, pontics, appliance components or fixture model components) into poses which are suitable for oral care appliance generation. The training may be performed using cohort patient case data. In some implementations, an RL model may use a reward function to learn, at least in part, a policy. The trained policy may be used by a deployed RL model to make decisions. In some implementations, anRL model (e.g., an Imitation Learning model) may use reference or ground truth data (e.g., a reference or ground truth transformation matrix for each tooth) to learn, at least in part, a policy. A policy may represent actions for the agent to taken in a given state. An example of an “action” is a transform that may be applied to a tooth mesh, an appliance, an appliance component, or other 3D representation. As a data-driven learning process, some implementations may be trained on reference (or ground truth) patient data which includes state-action pairs, such as [current state tooth transforms, action transforms]. For example, such a pair may take the form of: [[4x4 tooth transform], [4x4 tooth transform]]). An Imitation Learning model may use supervised learning to learn a policy, which may minimize loss. For example, the supervised learning model may comprise one or more neural networks (e.g., an MLP comprising an input layer, one or more hidden layers and an output layer). In some implementations, the training data may include tooth movement data (e.g., quaternions, or 4x4 transformation matrices). In some implementations, the training data may include latent representations of the patient’s dentition, such as latent representations of the patient’s teeth generated using an autoencoder, a transformer or the other representation generation methods described herein. For example, transforms or latent representations may be provided to the training module (e.g., a neural network) which generates tooth transforms for setups predictions. Some implementations may train a module which combine reference behaviors (actions, such as reference tooth movements) with a reinforcement learning reward function to train an algorithm that learns to replicate the optimal rewards of reference behaviors (state-actions pairs).
[00205] Some Reinforcement Learning models may train a Multi-armed Bandit model to generate transforms to place 3D representation of teeth, appliance components, fixture model components (among other items) into poses which are suitable for appliance generation. Multiple candidate actions (e.g., data pertaining to tooth movements - such as tooth transforms for use in setups) may be provided as inputs to a Multi-armed Bandits model. The neural network module of the Multi-armed Bandit model may be trained to ‘select’ one of the provided candidate actions, an action which may optimize a reward. The reward may quantify the results of an action taken by the agent. A higher reward may correspond to an action that more-highly benefits the agent. A Multi-armed Bandit model may comprise at least: one or more reinforcement learning agents, one or more environments, or one or more reward functions. A Multi-armed Bandit model may then compute a reward (e.g., mean square error between the expected tooth state and the actual tooth state) based on the chosen action (e.g., a predicted tooth transform) and train (e.g., via penalizing or rewarding) the neural network module through a loss function (e.g., vertex to vertex Euclidean distance between expect position and actual tooth position after applied the action). Some implementations of a Multi-armed Bandit model may target automated setups generation (e.g., through the selection of training data). Such implementations may be trained on both arches and all teeth in a patient’s mouth, or the implementation may be trained on the social six (e.g., top front six teeth), or be trained on some other specific set of teeth or malocclusion. Examples of malocclusion include open bite or deep bite cases. The training dataset may contain state-action pair vectors (e.g., [[4x4 tooth transform], [4x4 tooth transform]]), each of which may contain information about the pose of the 3D representation which is to be placed (e.g., such as tooth position, rotation or translation). In some implementations, the rewards may be computed only based on the actions, and independently of the predicted next state of the teeth. In other implementations, the rewards may take into consideration information about neighboring teeth, collisions between teeth or the distance between a tooth’s current position and the final setup position of that tooth. For example, such rewards may be computed position adjacent teeth in a manner so as to minimize collisions. In some implementations, rewards may be subject to constraints. In some implementations, constraints may include information pertaining to tooth movement limits (e.g., max clinical rotation allowed, or max translation allowed), space between teeth (e.g., max or min), collisions between adjacent teeth (e.g., maximum allowed collisions between the mesh elements of adjacent teeth), or root movement (e.g., max or min).
[00206] In some implementations, reinforcement learning may be used to train an ML model to label mesh elements (e.g., points, edges/faces/vertices, voxels), for example, for mesh segmentation or mesh cleanup. In some implementations, a segmentation parameters module may be trained using RL techniques to generate segmentation parameters. Such parameters may identify portions of a mesh which may be subject to more fine-grained segmentation than other portions of the mesh. For example, a portion of the mesh with a higher density (e.g., point cloud density, number of points, point distance, etc.) than other portions of the mesh may receive a finer-grained segmentation than those other portions of the mesh. A 3D representation (e.g., 3D point cloud or 3D mesh, etc. of the patient’s dentition) may be provided to the segmentation parameters module. The generated parameters may be provided to a subsequent ML model for segmentation (e.g., a U-Net, pyramid encoder-decoder, 3D SWIN, MASK- CNN or CNN) that then generates mesh element labels for the 3D representation based at least in part on the segmentation parameters. One or more segmented meshes may be generated (e.g., meshes for individual teeth of the patient’s dentition) through the application of the mesh element labels (e.g., labels on faces, vertices, edges, voxels or points). The newly segmented meshes may be used to compute an RL reward. Examples of the reward include Euclidean distance between the labelled mesh elements of a newly segmented mesh and a ground truth mesh (e.g., pairwise Euclidean distances between the mesh of the same mesh element label between the two meshes). The newly segmented meshes (e.g., the mesh element labels) may be provided to the ML model for segmentation as a new state. In some implementations, a representation generation ML model (e.g., an encoder from an autoencoder, an encoder from a transformer, or a stand-alone encoder) may generate one or more representations of the 3D representation. The latent representation may comprise a reduced-dimensionality form of the 3D representation. In some implementations, the latent representation may describe aspects of dental anatomy (e.g., fossae, tips, ridges, and the like). Denoising, downsampling or filtering operations may be applied to improve the fidelity of the 3D representation (e.g., prior to segmentation). Data augmentation of the training dataset may enable better performance and reduce overfitting of the segmentation ML model. In some implementations, the mesh element labeling techniques of this disclosure may be performed for the purpose of mesh cleanup (e.g., to label mesh elements corresponding to hardware that is to be removed from a tooth crown or extraneous material that is to be removed from a scanned dental arch).
[00207] In some implementations, an automated setups prediction model may be trained to generate a setup with a customized curve-of-spee (e.g., a curve-of-spee which conforms to the intended outcome of the treatment of the patient). Such a model may be trained on cohort patient case data. One or more oral care metrics may be computed on each case to quantify or measure aspects of that case's curve-of- spee. At training time, one or more of such metrics may be provided to the setups prediction model, for example, to influence the model regarding the geometry and/or structure of each case's curve-of- spee. Upon deployment of the setups prediction model, that same input pathway to the trained neural network may be configured with one or more values as instructions to the model about an intended curve- of-spee. Such values may automatically generate a setup with a curve-of-spee which meets the aesthetic and/or medical treatment needs of the particular patient case.
[00208] In some implementations, a curve-of-spee metric may measure the curvature of the occlusal or incisal surfaces of the teeth on either the left or right sides of the arch, with respect to the occlusal plane. The occlusal plane may, in some instances, be computed as a surface which averages the incisal or occlusal surfaces of the teeth (for one or both arches). In some implementations, a curvature metric may be computed along a normal vector, such as a vector which is normal to the occlusal plane. In other implementations, a curvature metric may be computed along the normal vector of another plane. In some implementations, an XY plane may be defined to correspond to the occlusal plane. An orthogonal plane may be defined as the plane that is orthogonal to the occlusal plane, which also passes through a curve- of-spee line segment, where the curve-of-spee line segment is defined by a first endpoint which is a landmarking point on a first tooth (e.g., canine) and a second endpoint which is a landmarking point on the most-posterior tooth of the same side of the arch. A landmarking point can in some implementations be located along the incisal edge of a tooth or on the cusp of a tooth. In some instances, the landmarking points for the intermediate teeth (e.g., teeth which are located between the first tooth and the most posterior tooth) on either the left or right sides of the arch may form a curved path, such as may be described by a polyline. The following is a non-limiting list of curve-of-spee oral care metrics.
[00209] 1) Measure the vertical height between a line segment and a point. Stated another way, measure a distance between a line segment and a point along the z-axis. The line segment is defined by joining the highest cusp of the most-posterior tooth (in the lower arch) and the cusp of the first tooth on that side (in the lower arch). Given the subset of teeth between the first tooth and the most-posterior tooth, the point is defined by the highest cusp of the lowest tooth of this subset. Stated another way, a curve-of-spee metric may be computed using the following 4 steps, i) Line: Form a line between the highest cusp on the most posterior tooth and the cusp of the first tooth, ii) Curve Point A: Given the set of teeth between the most posterior tooth and the first tooth, find the highest point of the lowest tooth, iii) Curve Point B: Project Curve Point A onto the Line to find a point (Curve Point B) along the line that is closest to Curve Point A. iv) Curve-Of-Spee: Find the height difference between Curve Point B and Curve Point A.
[00210] 2) Project one or more intermediate landmark points (e.g., points on the teeth which lie between the first tooth and the most-posterior tooth on that side of the arch) and the curve-of-spee line segment onto the orthogonal plane. Compute the curve-of-spee metric by measuring the distance between the farthest of the projected intermediate points to the projected curve-of-spee line segment. This yields a measure for the curvature of the arch relative to the orthogonal plane.
[00211] 3) Project one or more intermediate landmark points and the curve-of-spee line segment onto the occlusal plane. Compute Curve of Spee in this plane by measuring the distance between the farthest of the projected intermediate points to the projected curve-of-spee line segment. This yields a measure for the curvature of the arch relative to the occlusal plane.
[00212] 4) Skip the projection and compute the distances and curvatures in the 3D space. Compute
Curve of Spee by measuring the distance between the farthest of the intermediate points to the curve-of- spee line segment. This yields a measure for the curvature of the arch in 3D space.
[00213] 5) Compute the slope of the projected curve-of-spee line segment on the occlusal plane.
[00214] 6) Compute the slope of the projected curve-of-spee line segment in the orthogonal plane. [00215] Curve-of-spee metrics 5 and 6 may help the network to reduce some more degrees of freedom in defining how the patient’s arch is curved in the posterior of the mouth.
[00216] In some implementations, the RL techniques of this disclosure may generate transforms to place appliance components (e.g., as described herein) or fixture model components (e.g., as described herein) into poses which are suitable for oral care appliance generation (e.g., dental restoration appliances or orthodontic appliances - such as CTA).
[00217] In some implementations, the RL models for transform generation of this disclosure may be trained to generate transforms for one or more oral care appliance components (e.g., to generate transforms to place the one or more appliance components relative to one or more teeth of the patient). Losses (e.g., LI, L2, or reconstruction loss, among others described herein) may be computed between the predicted transforms (e.g., transforms predicted for the appliance components, or other 3D representation of oral care data described herein) and corresponding ground truth transforms that are provided with the training data. Such losses may be used to train, at least in part, the second ML module. Examples of pre-defined (or library) appliance components which may be placed using techniques of this disclosure include: vents, rear snap clamps, door hinges, door snaps, an incisal registration feature, center clips, custom labels, a manufacturing case frame, a diastema matrix handle, among others.
[00218] A digital fixture model may comprise 3D representations of the patient’s dentition, with optional fixture model components attached to that dentition or placed in relation to the dentition. The RL models for transform generation described herein may be trained to generate transforms to place digital fixture model components into poses that are suitable for oral care appliance generation. Fixture model components may include 3D representations (e.g., 3D point clouds, 3D meshes, or voxelized representations) of one or more of the following non-limiting items: 1) interproximal webbing - which may fill-in space or smooth-out the gaps between teeth to ensure aligner removability. 2) blockout - which may be added to the fixture model to remove overhangs that might interfere with plastic tray thermoforming or to ensure aligner removability. 3) bite blocks - occlusal features on the molars or premolars intended to prop the bite open. 4) bite ramps - lingual features on incisors and cuspids intended to prop the bite open. 5) interproximal reinforcement - a structure on the exterior of an oral care appliance (e.g., an aligner tray), which may extend from a first gingival edge of the appliance body on a labial side of the appliance body along an interproximal region between the first tooth and the second tooth to a second gingival edge of the appliance body on a lingual side of the appliance body. The effect of the interproximal reinforcement on the appliance body at the interproximal region may be stiffer than a labial face and a lingual face of the first shell. This may allow the aligner to grasp the teeth on either side of the reinforcement more firmly. 6) gingival ridge structure - a structure which may extend along the gingival edge of a tooth in the mesial-distal direction for the purpose of enhancing engagement between the aligner and a given tooth. 7) torque points - structures which may enhance force delivered to a given tooth at specified locations. 8) power ridges - structures which may enhance force delivered to a given tooth at a specified location. 9) dimples - structures which may enhance force delivered to a given tooth at specified locations. 10) digital pontic tooth - structure which may hold space open or reserve space in an arch for a tooth which is partially erupted, or the like. In aligners, a physical pontic is a tooth pocket that does not cover a tooth when the aligner is installed on the teeth. The tooth pocket may be filled with tooth-colored wax, silicone, or composite to provide a more aesthetic appearance. 11) power bars - blockout added in an edentulous space to provide strength and support to the tray. A power bar may fill- in voids. Abutments or healing caps may be blocked-out with a power bar. 12) trim line - digital path along the digital fixture model, which may approximately follow the contours of the gingiva (e.g., may be biased 1 or 2 mm in the gingival direction). The trimline may define the path along which a clear aligner may be cut or separated from a physical fixture model, after 3D printing. 13) undercut fill - material which is added to the fixture model to avoid the formation of cavities between the fixture model’s height of contour and another boundary (e.g., the gingiva or the plane that the plane that undergirds the physical fixture model after 3D printing).
[00219] A digital pontic tooth may be placed using the RL techniques of this disclosure. A pontic tooth is a digital 3D representation of a tooth which may act as a placeholder in an arch. In an aligner, a pontic may function as a tooth pocket that may be filled with a tooth-colored material (e.g., wax) to improve aesthetics. A pontic tooth may hold a space in the arch open during orthodontic setups generation. When automated setups prediction is performed, transforms for intermediate stages may be generated. One or more pontic teeth may be defined to hold open space for missing teeth or act as a placeholder as space closes or opens in the arch during setups predicted staging or an unerupted tooth may erupt into a space where a pontic is present. Creating a pocket for the tooth to erupt into.
[00220] Digital pontic teeth may be placed in (or generated within) an arch (e.g., during setups generation or fixture model generation) to reserve space in an arch for missing or extracted teeth (e.g., so that adjacent teeth do not encroach upon that space over the course of successive intermediate stages of orthodontic treatment). In some implementations, digital pontic tooth may be used when the space (e.g., the space that is to be held open) is at least a threshold dimension (e.g., a width of 4mm, among others). In some instances, digital pontic teeth for UL4-UR4 or LL4-LR4 may be placed (or generated or modified) when space is available or when there is a partially erupted tooth within the space. When there is a partially erupted tooth in the space, a digital pontic tooth may be placed over the empting tooth to maintain a space for the erupting tooth to erupt into. Pontic teeth may be placed to be inside the gingiva, to minimize (or avoid) heavy occlusal contacts (e.g., contacts between the chewing surfaces of the upper or lower arches), or to cover an erupting tooth (when present), among other conditions.
[00221] An RL-based ML model may be trained to generate (or modify) oral care data describing a current state of a treatment. Examples of such data include 3D representations of oral care data (e.g., teeth, appliance components, fixture model components, etc.), or transforms which may be applied to those 3D representations of oral care data. For example, a pre-restoration tooth mesh (or point cloud) may be provided to an RL-based ML model for 3D representation modification. The RL-based ML model may modify the pre-restoration tooth mesh, according to one or more oral care arguments (e.g., restoration design metrics) which may influence the RL-base ML model to modify the shape and/or structure of the pre-restoration tooth mesh in a manner so that the generated post-restoration tooth mesh is suitable for use in treatment of the patient. Other 3D representations of oral care data described herein may also be modified using RL: appliance components (e.g., parting surfaces, etc.), fixture model components (e.g., digital pontic teeth, or interproximal webbing, etc.), archforms, among other examples described herein.
[00222] In some implementations, RL techniques may be used to train ML models described herein (e.g., techniques based on encoder-decoder structures, such as autoencoders or transformers) for the generation or modification of 3D representations (e.g., 3D point clouds, voxels, etc.). An RL method of training an ML model to generate a 3D representation may involve elements such as: state, actions, rewards, and policy.
[00223] In some implementations, the state of a 3D representation may be described by the structures, positions, and/or mesh element feature vectors of a set of mesh elements. A mesh element feature vector may be computed, and be provided to the ML models described herein. For example, the state of a 3D representation (e.g., point cloud) may comprise the set of XYZ coordinates for the points of the point cloud. In a further example, the state of a 3D mesh may comprise the XYZ coordinates of the vertices, and/or the edge or face lists which describe the structure of the connections between the vertices.
[00224] Furthermore, positional and/or structural state data may be augmented by one or more mesh element features (e.g., vertex normals, face normals, edge dihedral angles, vertex curvatures, edge curvatures, vertex concavity /convexity measurements, distance to nearest opposing surface), as described herein. In some implementations, the state of a 3D representation may comprise, at least in part, the set of mesh element feature vectors associated with the 3D representation’s mesh elements. Each mesh element feature vector may be modified by actions, explicitly or implicitly, and/or may be used to inform actions and/or rewards.
[00225] An agent (i.e., a trained MLP or other of the neural networks described herein) may generate one or more actions which modify the state of the 3D representation which is being generated. In some implementations, the actions may be operable to modify the configuration of mesh elements (e.g., points of a point cloud, etc.) into a shape and/or structure which is suitable for use in generating an oral care appliance. One or more oral care arguments (e.g., oral care procedure parameters or oral care metrics) may be provided to the agent, to influence the agent to generate actions which result in the 3D representation being modified into a shape and/or structure which is suitable for use in generating an oral care appliance. Actions may include modifying the positions and/or orientations of mesh elements. Nonlimiting examples of actions include:
1) scaling mesh elements
2) adjusting the positions of mesh elements along one or more axes (e.g., global or local axes)
3) adjusting a surface to become more concave or convex
4) adjusting a surface to become less concave or convex
5) taking a step in the interpolation between a source 3D representation and a target 3D representation incrementally advancing a morph from a source 3D representation to a target 3D representation (e.g., a predicted target 3D representation which is predicted using techniques of this disclosure) [00226] A morphing action, as used herein, may comprise modifying the shape and/or structure of an initial 3D representation to more nearly reflect the shape and/or structure of a target 3D representation, for example, over a series of evolving steps.
[00227] Actions samples may be provided as training data. Actions could also be computed based on the mesh-morphing technique. The same technique can be used to compute rewards when using 3D representations.
[00228] The RL environment may comprise one or more constraints (e.g., on motion, rotation, translation, width, length, etc.), a function that applies actions or a reward function. The reward may quantify the magnitude of success of a new state (e.g., after application of the predicted action), for example, a modified shape and/or structure of the 3D representation (e.g., in shaping one or more teeth for use in restoring the patients’ dentition, in shaping a fixture model component, etc.).
[00229] In some implementations, the reward may impact loss computation (e.g., comparison between predicted action and the ground truth action, such using reconstruction loss or mean square error of a list of KPIs (e.g., mesh elements, positions, etc.) or another of the losses described herein). The reward may, in some implementations, involve the calculation of oral care metrics (e.g., “Bilateral Symmetry & Ratios,” “Proximal Contacts,” or “Tooth Morphology” when a point cloud for a tooth restoration design is generated). Such metrics may contribute to assessing the progress of the 3D representation towards taking-on a shape and/or structure that is indicated by the one or more oral care arguments. For example, the metrics may measure how well the point cloud describes a tooth restoration design, a fixture model component, a generated appliance component, or another 3D oral care representation that is suitable for use in appliance generation. The generated (or modified) point cloud may then be used as part of the generation of an oral care appliance. The dataset may comprise 3D oral care representations that are divided into train, test and validation sets (e.g., 75%, 15%, 10%).
[00230] An RL model may be trained to generate (or modify) 3D representations, in some implementations, the RL model may function alongside another ML model (e.g., a transformer or an autoencoder) to generate (or modify) a 3D representation. For example, an RL model (e.g., an MLP with optional skip connections) may be trained to generate a tooth restoration design (or to modify an existing tooth design) in conjunction with an encoder-decoder structure of this disclosure (e.g., an encoderdecoder which is trained as a reconstruction autoencoder). The encoder-decoder structure may generate a 3D representation of the patient’s dentition (e.g., one or more tooth crowns), which may be provided to the RL model. The RL model may generate one or more actions to be taken to improve the provided 3D representation of patient's dentition. The agent may comprise one or more neural networks (e.g., an encoder-decoder structure such as an autoencoder or transformer, or an MLP) which are trained with multiple sample actions (e.g., from past data from cohort patient cases) so the agent leams how to generate actions. The actions generated by the RL model may be applied to a partially generated or formed 3D representation (e.g., a 3D representation of the patient’s dentition that is under modification - such as to incorporate fixture model components into the dentition as a part of fixture model generation), resulting in a modified 3D representation of the patient’s dentition with evolving fixture model components attached. Ground truth (or reference) cohort patient case data may be provided to the training workflow. One or more loss values may be computed by comparing the generated (or modified) 3D representation to the corresponding ground truth data 3D representation. The one or more loss values (e.g., reconstruction loss, chamfer loss or other losses described herein) may be used to train, at least in part, the one or more neural networks which comprise the agent of the RL model.
[00231] A reward may, in some implementations, be computed by comparing a modified 3D representation (e.g., new state or next state) to a corresponding ground truth 3D representation. The reward function may influence the agent to execute actions which change the state of the 3D representation, and to change the state of the 3D representation in a manner which causes the 3D representation to assume a shape and/or structure which is more nearly suitable for use in generating an oral care appliance. This computed reward may be used to optimize a agent or policy neural network, enabling that the agent or policy neural network to predict or generate better actions.
[00232] In reinforcement learning (RL), loss values may be computed based on a comparison between predicted actions and corresponding ground truth actions. Rewards may be computed based upon a comparison between predicted next states and corresponding ground truth next states. Both losses and rewards may, in some implementations, be computed upon vectors of real values. In loss calculation, those vectors describe actions. In reward calculation, those vectors describe states. While losses and rewards both reflect differences, losses reflect differences in actions, and rewards reflect differences in states.
[00233] When using an RL model to generate intermediate setups for orthodontic treatment, an action may, in some implementations, comprise a transform that describes a transition of a tooth from a current pose to a next pose. A current state may, in some implementations, comprise a transform that describes a current pose of a tooth, and a next state may comprise a different transform that describes a next pose of a tooth.
[00234] When using an RL model to generate (or modify) a 3D representation (e.g., a point cloud, a mesh, etc.), an action may, in some implementations, comprise an adjustment to the concavity of a surface (among others described herein). A current state of a 3D representation may, in some implementations, comprise the set of mesh elements (and optionally the associated mesh element feature vectors). A next current state of a 3D representation may, in some implementations, comprise the set of mesh elements (and optionally the associated mesh element feature vectors) after the application of at least one action. That is, the next state may comprise the set of mesh elements, where at least one mesh element (e.g., the position of a vertex, among others) has been modified by the application of an action. [00235] The loss calculation methods described herein may compare two or more vectors of real values, such as MSE, LI, L2, reconstruction loss, among others. Rewards and losses may, in some examples, be computed based upon these method of comparing vectors of real values. For example, in some implementations, loss may be computed (at least in part) using reconstruction loss, and the reward may be computed using MSE. [00236] In some implementations, a loss may be computed, at least in part, based upon a reward. That is, a reward value may contribute to the calculation of a loss value, in some implementations. [00237] In reinforcement learning (RL), an agent neural network goal is to learn a policy. A policy maps a state to an action. In the continuous space there are infinite numbers of possible states and therefore infinite number of possible actions. For this reason, a machine learning model is trained to learn actions from states.
[00238] The reinforcement environment represents the common processes of applying actions to states. The RL environment may also apply constraints to the actions or to the new states. Constraints may be represented by motions limits (e.g. : translation, rotation) or shape limits (e.g. : magnitude of shape change or deformation).

Claims

CLAIMS WHAT IS CLAIMED IS:
1. A method of generating a transformation for a three-dimensional (3D) representation of oral care data, the method comprising: receiving, by processing circuitry of a computing device, oral care data describing a current state of a treatment; providing, by the processing circuitry, the oral care data describing a current state of a treatment to a reinforcement learning (RL) model as an input; executing, by the processing circuitry and using the oral care data describing a current state of a treatment, one or more trained neural networks to generate a representation that specifies a predicted action associated with the oral care data describing a current state of a treatment; and generating, by the processing circuitry and using the current state and the action, a predicted next state of the treatment.
2. The method of claim 1, wherein the oral care data describing a current state of a treatment comprises at least one of: one or more 3D representations of the oral care data, and one or more transforms that when applied to a 3D representation of oral care data modifies one or more poses of the 3D representation of the oral care data.
3. The method of claim 1 , wherein one or more rewards are computed to quantify a difference between the predicted next state and a corresponding ground truth next state.
4. The method of claim 3, wherein one or more loss values are computed which quantify the difference between the predicted action and a corresponding ground truth action.
5. The method of claim 4, wherein the one or more losses are computed, at least in part, based upon the one or more rewards.
6. The method of claim 2, wherein the 3D representation of oral care data is at least one of a 3D mesh, a 3D point cloud, a 3D voxelized representation, or a 3D surface.
7. The method of claim 2, wherein the one or more 3D representations of the oral care data represents one or more of: one or more teeth of a dental arch, one or more appliance components, or one or more fixture model components.
8. The method of claim 2, wherein the action moves at least one aspect of the 3D representation of the oral care data into a pose which is suitable for use in digital oral care.
9. The method of claim 2, further comprising: receiving, by processing circuitry of a computing device, the 3D representation of the oral care data; and wherein generating, by trained neural network, the representation uses at least one of the one or more transforms and the 3D representation of the oral care data.
10. The method of claim 1, wherein executing the trained neural network comprises executing the RL model according to a policy -based methodology.
11. The method of claim 1, wherein executing the trained neural network comprises executing the RL model according to a policy -free methodology.
12. The method of claim 1, wherein the neural network is trained according to one or more of a Q- leaming paradigm, a soft actor critic (SAC) learning paradigm, generative adversarial imitation learning (GAIL) paradigm.
13. The method of claim 1, wherein the computing device is deployed at a clinical context, and wherein the method is performed in near real-time during an encounter with a patient.
14. The method of claim 1, wherein the RL model is trained, at least in part, through reinforcement learning by human feedback (RLHF).
15. The method of claim 7, wherein tire RL model is used to generate a design for at least one of an orthodontic appliance and a dental restoration appliance.
16. The method of claim 2, wherein the one or more transforms are used to modify the pose of one or more teeth of the patient, to generate an orthodontic setup.
17. The method of claim 16, wherein the orthodontic setup is a final setup which reflects the poses of the one or more teeth after completion of an orthodontic treatment of a patient.
18. The method of claim 16, wherein the orthodontic setup is an intermediate setup which reflects the poses of the one or more teeth in a non-final stage during a course of an orthodontic treatment of a patient.
19. The method of claim 1, further comprising, providing, by the processing circuitry, as additional input data to the RL model, at least one of: (i) one or more 3D geometries describing one or more teeth, (ii) one or more vectors P containing at least one value pertaining to at least one method of computing a dimension of at least one tooth, (iii) one or more vectors Q containing at least one value pertaining to at least one method of computing a distance between adjacent teeth, (iv) one or more vectors B containing latent vector information about one or more teeth, (v) one or more vectors N containing at least one value pertaining to the position of at least one tooth, (vi) one or more vectors O containing at least one value pertaining to the orientation of at least one tooth, (vii) one or more vectors R at least one of tooth name, designation, tooth type and tooth classification.
20. A computing device for generating a transformation for a three-dimensional (3D) representation of oral care data for use in orthodontic alignment treatment, the computing device comprising: interface hardware configured to receive the 3D representation of oral care data; processing circuitry configured to: receive, by the processing circuitry of a computing device, oral care data describing a current state of a treatment; provide, by the processing circuitry, the oral care data describing a current state of a treatment to a reinforcement learning (RL) model as an input; execute, by the processing circuitry and using the oral care data describing a current state of a treatment, one or more trained neural networks to generate a representation that specifies a predicted action associated with the oral care data describing a current state of a treatment; and generate, by the processing circuitry and using the current state and the action, a predicted next state of the treatment.
EP23828816.1A 2022-12-14 2023-12-14 Reinforcement learning for final setups and intermediate staging in clear tray aligners Pending EP4634936A1 (en)

Applications Claiming Priority (3)

Application Number Priority Date Filing Date Title
US202263432627P 2022-12-14 2022-12-14
US202363460262P 2023-04-18 2023-04-18
PCT/IB2023/062694 WO2024127303A1 (en) 2022-12-14 2023-12-14 Reinforcement learning for final setups and intermediate staging in clear tray aligners

Publications (1)

Publication Number Publication Date
EP4634936A1 true EP4634936A1 (en) 2025-10-22

Family

ID=89378644

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23828816.1A Pending EP4634936A1 (en) 2022-12-14 2023-12-14 Reinforcement learning for final setups and intermediate staging in clear tray aligners

Country Status (3)

Country Link
EP (1) EP4634936A1 (en)
CN (1) CN120303746A (en)
WO (1) WO2024127303A1 (en)

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118798139B (en) * 2024-07-11 2025-11-14 华中科技大学 A Controllable Text Generation Method Based on Fine-Grained Feedback Reinforcement Learning
CN118885814B (en) * 2024-09-19 2024-12-20 四川长园工程勘察设计有限公司 Battery charging and discharging optimization method, system and medium based on deep reinforcement learning
CN119723102B (en) * 2024-11-19 2026-02-17 南京一目智能科技有限公司 A method, apparatus, electronic device and storage medium for detecting clothing
CN120755872B (en) * 2025-07-17 2026-01-30 北京炎凌嘉业智能科技股份有限公司 Training method, spraying method and system for intelligent spraying model of curved surface

Family Cites Families (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2020026117A1 (en) 2018-07-31 2020-02-06 3M Innovative Properties Company Method for automated generation of orthodontic treatment final setups
EP3902499A4 (en) 2018-12-26 2022-09-07 3M Innovative Properties Company Methods to automatically remove collisions between digital mesh objects and smoothly move mesh objects between spatial arrangements
US11735306B2 (en) * 2019-11-25 2023-08-22 Dentsply Sirona Inc. Method, system and computer readable storage media for creating three-dimensional dental restorations from two dimensional sketches
EP4161435A4 (en) 2020-06-03 2024-10-02 Solventum Intellectual Properties Company SYSTEM FOR PRODUCING A GRADUATED TREATMENT OF AN ORTHODONTIC ALIGNER
JP2023552589A (en) 2020-12-11 2023-12-18 スリーエム イノベイティブ プロパティズ カンパニー Automatic processing of dental scans using geometric deep learning
WO2022135722A1 (en) * 2020-12-23 2022-06-30 Hirsch Dynamics Holding Ag Automatic creation of a virtual model of at least a part of an orthodontic appliance
KR102469288B1 (en) * 2022-05-13 2022-11-22 주식회사 쓰리디오엔에스 Methods, devices and computer programs for automated orthodontic planning

Also Published As

Publication number Publication date
WO2024127303A1 (en) 2024-06-20
CN120303746A (en) 2025-07-11

Similar Documents

Publication Publication Date Title
EP4633528A1 (en) Denoising diffusion models for digital oral care
EP4634798A1 (en) Neural network techniques for appliance creation in digital oral care
EP4634936A1 (en) Reinforcement learning for final setups and intermediate staging in clear tray aligners
WO2024127309A1 (en) Autoencoders for final setups and intermediate staging in clear tray aligners
WO2024127316A1 (en) Autoencoders for the processing of 3d representations in digital oral care
EP4634797A1 (en) Machine learning models for dental restoration design generation
US20250364117A1 (en) Mesh Segmentation and Mesh Segmentation Validation In Digital Dentistry
US20250366959A1 (en) Geometry Generation for Dental Restoration Appliances, and the Validation of That Geometry
WO2024127302A1 (en) Geometric deep learning for final setups and intermediate staging in clear tray aligners
EP4634931A1 (en) Metrics calculation and visualization in digital oral care
WO2023242761A1 (en) Validation for the placement and generation of components for dental restoration appliances
EP4633526A1 (en) Transformers for final setups and intermediate staging in clear tray aligners
EP4633527A1 (en) Pose transfer techniques for 3d oral care representations
US20260011442A1 (en) Validation of Tooth Setups for Aligners in Digital Orthodontics
WO2025074322A1 (en) Combined orthodontic and dental restorative treatments
US20250366958A1 (en) Validation for Rapid Prototyping Parts in Dentistry
US20260020937A1 (en) Bracket and Attachment Placement in Digital Orthodontics, and the Validation of Those Placements
US20250359964A1 (en) Coordinate System Prediction in Digital Dentistry and Digital Orthodontics, and the Validation of that Prediction
WO2024127314A1 (en) Imputation of parameter values or metric values in digital oral care
WO2024127308A1 (en) Classification of 3d oral care representations
WO2025126117A1 (en) Machine learning models for the prediction of data structures pertaining to interproximal reduction
WO2024127310A1 (en) Autoencoders for the validation of 3d oral care representations
EP4634930A1 (en) Force directed graphs for final setups and intermediate staging in clear tray aligners

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250624

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

DAV Request for validation of the european patent (deleted)
DAX Request for extension of the european patent (deleted)