WO2025199235A1 - Compressed meta-optical encoder for image classification - Google Patents

Compressed meta-optical encoder for image classification

Info

Publication number
WO2025199235A1
WO2025199235A1 PCT/US2025/020565 US2025020565W WO2025199235A1 WO 2025199235 A1 WO2025199235 A1 WO 2025199235A1 US 2025020565 W US2025020565 W US 2025020565W WO 2025199235 A1 WO2025199235 A1 WO 2025199235A1
Authority
WO
WIPO (PCT)
Prior art keywords
meta
optics
images
electronic
kernels
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/US2025/020565
Other languages
French (fr)
Inventor
Anna-Wirth Singh
Jinlin XIANG
Minho Choi
Johannes Emanuel FRÖCH
Luocheng HUANG
Shane COLBURN
Eli Shlizerman
Arka MAJUMDAR
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
University of Washington
Original Assignee
University of Washington
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by University of Washington filed Critical University of Washington
Publication of WO2025199235A1 publication Critical patent/WO2025199235A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G02OPTICS
    • G02BOPTICAL ELEMENTS, SYSTEMS OR APPARATUS
    • G02B1/00Optical elements characterised by the material of which they are made; Optical coatings for optical elements
    • G02B1/002Optical elements characterised by the material of which they are made; Optical coatings for optical elements made of materials engineered to provide properties not available in nature, e.g. metamaterials
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/06Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons
    • G06N3/067Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons using optical means
    • G06N3/0675Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons using optical means using electro-optical, acousto-optical or opto-electronic means
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/084Backpropagation, e.g. using gradient descent
    • BPERFORMING OPERATIONS; TRANSPORTING
    • B82NANOTECHNOLOGY
    • B82YSPECIFIC USES OR APPLICATIONS OF NANOSTRUCTURES; MEASUREMENT OR ANALYSIS OF NANOSTRUCTURES; MANUFACTURE OR TREATMENT OF NANOSTRUCTURES
    • B82Y20/00Nanooptics, e.g. quantum optics or photonic crystals

Definitions

  • CNNs are composed of several electronics convolutional layers that adaptively learn spatial representations from input images. While powerful, the electronics convolution operation is computationally expensive, leading to high latency and power consumption. In fact, it has been estimated that about 80% of the total runtime of CNNs is used performing convolution operations. Reducing this latency and power consumption may include free- space optical systems as a solution. Beyond power consumption and latency reduction, optical information processing may be characterized by high bandwidth, spatial parallelism, and low-loss transmission. However, a challenge to all CNN approaches is that non- linear layers are interspersed with the linear layers. For example, AlexNet, a known baseline structure in computer vision challenges, consists of five convolutional layers followed by three fully connected layers (electrical layers), employing ReLU activations.
  • the classification accuracy of the CNN may drop by about 20%.
  • the nonlinear layers cannot be implemented using simple lens like optics; to implement them optically, some physical nonlinearity must be introduced, for instance by using an atomic vapor cell or image intensifier.
  • hybrid approaches involving repeated transduction of the signal (e.g., optical images) to perform linear operations in optics and non-linear operations in electronics provide little benefit due to large latency and power consumption in signal transduction. Implementing only one of many required convolution operations does not provide significant benefit in terms of speed and latency.
  • Knowledge distillation may circumvent the need for nonlinearity without a significant reduction in the performance by transferring knowledge from a larger, pre-trained network (the ‘teacher’ network) to a more compact network (the ‘student’ network).
  • AlexNet is the teacher network and a single convolutional layer coupled with a single fully connected layer is the student network.
  • inventive technology uses a hybrid meta-optical platform, where an optical frontend is based on a single meta-optic performing the convolution operation, followed by an electronic backend which contains a linear calibration layer and a fully connected layer.
  • the optical frontend of the network may be realized using inverse-designed meta- optics.
  • the meta-optic may include arrays of sub-wavelength scatterers which act as phase masks, imparting spatially-coded phase shifts to incident light.
  • the meta-optics is designed to realize a phase mask which performs the desired convolutional steps of the CNN by engineering the PSF.
  • inventive technology achieves a classification accuracy of 93%, which is comparable to the 96% accuracy achieved using an all-electronic network compressed using knowledge distillation.
  • MAC multiply-accumulation
  • a polychromatic optical encoder with PSF-engineered meta- optic is used to classify the Canadian Institute For Advanced Research in 10 classes (CIFAR-10) dataset. This makes the inventive system compact and fully compatible with the conventional optical imaging systems.
  • a convolutional neural network (CNN) system for image classification includes a meta-optics having a plurality of convolutional kernels, each convolutional kernel having a plurality of nanostructures.
  • the plurality of kernels are configured for convolving source images by applying their respective point spread functions (PSFs) to the source images.
  • An image capturing camera is configured for capturing convolved images.
  • convolving images includes multi-wavelength kernel convolutions that are executed simultaneously.
  • the multi-wavelength kernel convolutions include convolutions of a red (R) wavelength, a green (G) wavelength, and a blue (B) wavelength of light.
  • the plurality of kernels are distributed over a monolithic meta- optics.
  • the PSFs of the plurality of kernels are at least in part obtained from a knowledge distillation that includes compressing operations of multiple electronic convolutional layers into an operation of one electronic convolutional layer.
  • the PSFs are determined by simulation prior to manufacturing of the meta-optics.
  • the electronic backend software includes a calibration layer configured for compensating fabrication and misalignment errors of the meta-optics.
  • the calibration layer is configured for assigning weights to the convolved images, where the weights are determined by comparing the convolved images obtained by the meta-optics with the convolved images obtained by electronic convolution.
  • the nanostructures are made of silicon nitride on a quartz substrate.
  • the electronic backend software is a trained and reconfigurable artificial intelligence (AI).
  • the convolved images are equivalent to an output of 2- dimensional convolutional layers.
  • the system also includes a source of light configured to emit in a visible spectrum.
  • a method for convolutional neural network (CNN) image classification includes illuminating a meta-optics having a plurality of convolutional kernels with source images.
  • Each convolutional kernel includes a plurality of nanostructures.
  • the method also includes convolving the source images by applying point spread functions (PSFs) of the convolutional kernels to the source images; and capturing convolved images by an image capturing camera. Based on the convolved images, the method classifies the source images with an electronic backend software.
  • PSFs point spread functions
  • FIGURE 1 is a scanning electron micrograph image of a meta-optics in accordance with an embodiment of the present technology
  • FIGURES 2A-2C illustrate several views of meta-optic’s nanoposts in accordance with embodiments of the present technology
  • FIGURE 3 illustrates operation of a meta-optic according to embodiments of the present technology
  • FIGURE 4A is a schematic view of a convolutional neural network (CNN) for image classification based on all-electronic multi-layered CNN in accordance with embodiments of the present technology
  • FIGURE 4B is a schematic view of a CNN for image classification based on all- electronic compressed CNN in accordance with embodiments of the present technology
  • FIGURE 4C is a schematic view of a CNN for image classification based on a hybrid CNN which combines optical meta-optic front end
  • the designation value may vary by plus or minus twelve percent, or eleven percent, or ten percent, or nine percent, or eight percent, or seven percent, or six percent, or five percent, or four percent, or three percent, or two percent, or one percent.
  • the use of the term “at least one” will be understood to include one as well as any quantity more than one, including but not limited to, 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 100, etc.
  • the term “at least one” may extend up to 100 or 1000 or more, depending on the term to which it is attached; in addition, the quantities of 100/1000 are not to be considered limiting, as lower or higher limits may also produce satisfactory results.
  • the words “comprising” (and any form of comprising, such as “comprise” and “comprises”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “includes” and “include”) or “containing” (and any form of containing, such as “contains” and “contain”) are inclusive or open-ended and do not exclude additional, unrecited elements or method steps.
  • the term “or combinations thereof” as used herein refers to all permutations and combinations of the listed items preceding the term.
  • A, B, C, or combinations thereof is intended to include at least one of: A, B, C, AB, AC, BC, or ABC and, if order is important in a particular context, also BA, CA, CB, CBA, BCA, ACB, BAC, or CAB.
  • expressly included are combinations that contain repeats of one or more item or term, such as BB, AAA, AB, BBC, AAABCCCC, CBBAAA, CABABB, and so forth.
  • BB BB
  • AAA AAA
  • AB BBC
  • AAABCCCCCC CBBAAA
  • CABABB CABABB
  • FIG.1 is an optical image of a meta-optics (also referred as a metalens or meta- optic encoder) in accordance with an embodiment of the present technology.
  • Illustrated meta-optics 100 includes a number of nanostructures (also referred to as nanoposts or scatterers) 110 that are carried by a substrate (also referred to as a carrier) 115.
  • the nanostructures 110 may be nanoscale structures that are generally cylindrical or rectangular, and are characterized by one or more characteristic scales (e.g., cylinder diameter d, width w, height t, etc.).
  • the nanostructures 110 may have different sizes, as illustrated in FIG.1.
  • the meta-optics 100 may be manufactured by the process described below.
  • a 600 nm layer of silicon nitride is first deposited via plasma-enhanced chemical vapor deposition (PECVD) on a quartz substrate, followed by spin-coating with a high-performance positive electron beam resist (e.g., ZEP-520A).
  • PECVD plasma-enhanced chemical vapor deposition
  • ZEP-520A high-performance positive electron beam resist
  • An 8 nm Au/Pd charge dissipation layer is then sputtered followed by subsequent exposure to an electron-beam lithography system (e.g., JEOL JBX6300FS).
  • the Au/Pd layer may then be removed with a thin film etchant (e.g., type TFA gold etchant), and the samples may be developed in amyl acetate.
  • a thin film etchant e.g., type TFA gold etchant
  • 50 nm of aluminum is evaporated and lifted off via sonication in methylene chloride, acetone, and isopropyl alcohol.
  • the samples are then dry etched using a CHF3 and SF6 chemistry and the aluminum is removed by immersion in AD-10 photoresist developer.
  • FIGS.2A-2C illustrate several views of nanoposts in accordance with embodiments of the present technology.
  • FIG. 2A is an isometric view of a nanopost 110 that is carried by a substrate 115.
  • the illustrated nanopost 110 is cylindrical, but in other embodiments the nanopost 110 may have other shapes, for example, an elliptical cross-section, a square cross-section, a rectangular cross-section or other cross-sectional shape that maintain center-to-center spacing at a sub-wavelength value.
  • FIG.2B is a top view of two adjacent nanoposts that are separated by a distance “p” (pitch). Only two nanoposts are illustrated in FIG.2B for simplicity. However, for a practical meta-optics 100, many more nanoposts are distributed over the substrate 115.
  • FIG. 2C is a side view of a nanopost 110 that is carried by a substrate 115.
  • the nanoposts (scatterers, nanostructures) 110 are made of silicon nitride due to its broad transparency window and CMOS compatibility.
  • the illustrated nanoposts 110 are characterized by a height “t” and diameter “d”.
  • the values of “d” may range from about 100 nm to about 300 nm.
  • the value of “t” (height) is constant (within the limits of manufacturing tolerance) for all diameters “d” for a given metalens.
  • the values of “t” may range from about 500 nm to about 800 nm.
  • the nanoposts may be polarization-insensitive cylindrical nanoposts 110 arranged in a square lattice on a quartz substrate 115.
  • the phase shift mechanism of these nanoposts arises from an ensemble of oscillating modes within the nanoposts that couple amongst themselves at the top and bottom interfaces of the post.
  • the modal composition varies, modifying the transmission coefficient through the nanoposts.
  • FIG.3 illustrates operation of meta-optics 100. In operation, an image is projected onto the meta-optics 100 as an optical input.
  • the meta-optics 100 convolves the input image, and produces an output image that is projected-on and acquired-by an optical sensor 180 (e.g., a digital camera).
  • an optical sensor 180 e.g., a digital camera
  • FIG.4A is a schematic view of a convolutional neural network (CNN) 1001 for image classification based on all-electronic multi-layered CNN in accordance with embodiments of the present technology.
  • FIG.4B is a schematic view of a CNN for image classification based on compressed all-electronic CNN 1002 in accordance with embodiments of the present technology.
  • FIG.4C is a schematic view of a CNN 1003 for image classification based on a hybrid CNN which combines an optical meta-optic front end with electronic backend in accordance with embodiments of the present technology.
  • the CNNs 1001 and 1002 take input 11 that is a digitized image (e.g., an image acquired by a digital camera, coupled with a capability of assigning numerical values to individual pixels of the digital camera like a digital image of a scene where the values represented at each pixel of the image correspond to intensity values of an imaged scene or an image wherein the values represented at each pixel of the image correspond to intensity values of an imaged scene).
  • a digitized image e.g., an image acquired by a digital camera, coupled with a capability of assigning numerical values to individual pixels of the digital camera like a digital image of a scene where the values represented at each pixel of the image correspond to intensity values of an imaged scene or an image wherein the values represented at each pixel of the image correspond to intensity values of an imaged scene.
  • Such input is then electronically processed by multiple convolutional layers 300 and multiple fully connected layers 400 of the teacher network 1001, or by a single convolutional layer 300 and a single fully connected layer 400 of the student network 1002.
  • FIG. 4A teacher network 1001 into a single convolution layer 300 of FIG.4B.
  • the convolution layers 300 are electronic convolution layers, thus the teacher network 1001 and student network 1003 are both fully electronic networks.
  • a modified AlexNet, denoted AlexNet-Mod can be used as the teacher network of FIG. 4A, whereas a single convolutional layer 300 coupled with a single fully connected layer400 can be used as the student network of FIG.4B.
  • 4C illustrates a hybrid meta-optics platform 1003, where an optical frontend based on a single meta-optics 100 performs the linear convolution operation of the light signal coming from the light source 10 (e.g., display showing a source image to be classified, a laser, a light emitting diode, a digital display with a representation of a digital image, a reflective or transmissive target with a representation of an object or scene, an ambient scene illuminated by sunlight or man-made sources containing objects to be detected or classified, etc.), followed by an electronic backend which contains a linear calibration layer 130 and a fully connected layer 400.
  • the light source 10 e.g., display showing a source image to be classified, a laser, a light emitting diode, a digital display with a representation of a digital image, a reflective or transmissive target with a representation of an object or scene, an ambient scene illuminated by sunlight or man-made sources containing objects to be detected or classified, etc.
  • the electronic backend which contains
  • the meta-optics 100 includes arrays of sub- wavelength scatterers which act as phase masks, imparting spatially-coded phase shifts to incident light. In operation, these spatially-coded phase shifts to incident light impart a point spread function (PSF) onto the input image as the input image is convolved by the meta-optics 100.
  • PSF point spread function
  • the meta-optics 100 is designed to apply a phase mask which performs the desired convolutional steps of the CNN by engineering the required PSF.
  • such meta-optics 100 is fabricated and experimentally validated using incoherent green light illumination, e.g., light illumination from a light emitting diode, centered at 525 nm.
  • incoherent green light illumination e.g., light illumination from a light emitting diode, centered at 525 nm.
  • other sources of light are also possible, for example, lasers or light emitting diodes operating at different wavelengths.
  • FIG.5A is a point spread function (PSF) measurement setup using a monochromatic point light source (left) and optical convolution using a micro-LED display (right) in accordance with embodiments of the present technology. As explained above with respect to FIGs.
  • PSF point spread function
  • compressed network 1002 with a single convolutional layer can be obtained by distilling knowledge from a teacher network 1001 (e.g., AlexNet) to a student network of the desired structure.
  • the knowledge distillation approach assumes that the teacher network is already trained and performs the desired task with high accuracy.
  • AlexNet as the teacher network 1001 may achieve over 98% classification accuracy for both training and testing MNIST datasets over repeated trials.
  • a sample AlexNet consists of five electronical convolutional layers 300 and three fully connected electronic layers 400 (total of 8 layers) which can be compressed to the student system 1002 having one electronical convolutional layer 300 and one fully connected layer 400 (two layers total).
  • AlexNet teacher network 1001 may employ nonlinear activation functions (ReLU) to optimize performance.
  • ReLU nonlinear activation functions
  • the meta-optics 100 (i.e., the compressed convolutional layer) is divided into positive and negative kernels, for example 8 positive and 8 negative kernels that are characterized by their individual PSFs, as further explained below.
  • each kernel may be 6 ⁇ 6 pixels in size (as captured on the pixels of the camera 200), but other pixel numbers allocated to each positive/negative kernel as well as other numbers of kernels are also possible in different embodiments.
  • the illustrated meta-optic 100 contains 16 different sub-optics (or kernels), which can be understood as segments of one monolithic meta-optic 100, the kernels being spatially distributed within a single convolution layer to operate in parallel for the classification tasks.
  • the images acquired by the camera 200 may be forwarded to the fully connected (electronic backend) layer 400 for classification.
  • the left-hand side illustration of the CNN in FIG. 5A includes a source of the light 10 (e.g., a fiber coupled LED or laser) that provides PSF-determining images to the camera 200 and further to the fully connected (electronic back-end) layer 400.
  • the right-hand side illustration of the CNN in FIG. 5A includes a real image (in the illustrated sample case, an image of a digit ‘7’) that projects on the meta-optics 100 and further onto the camera 200.
  • the left-hand side setup of the CNN 1003 in FIG.5A may be used for determining or testing the real PSFs (i.e., 8 PSFs for positive kernels and 8 PSFs for negative kernels) of a given meta-optics 100 using, for example, a monochromatic source of light 10.
  • the PSF of each sub-optic kernel can be measured using a single-mode fiber 10 as a point light source.
  • the PSFs from all 16 sub-optics can be simultaneously captured by the two-dimensional CMOS camera 200.
  • the 5A produces convolved images captured by the camera 200, where the captured images represent f(x, y)*PSF(x,y) for the 16 kernels (8 positive and 8 negative), which a person of ordinary skill would understand to be convolutions of the source image of numeral ‘7’ originating from the display 52. Accordingly, the convolved images from all 16 sub- optics kernels can be captured simultaneously when the single mode fiber 10 is replaced with the display 52 that emits source images. Generally, the required meta-optics 100 can be inverse-designed if a target PSF is known. In some embodiments, the Gerchberg-Saxton (GS) phase retrieval method may be used to design phase masks that correspond to the optimized convolutional kernels.
  • GS Gerchberg-Saxton
  • each sub-optics may be designed to have a PSF which resembles the convolutional kernel. Since electronic convolutional kernels include both positive and negative values, each kernel may be separated into positive and negative parts and meta- optics can be designed for each PSF. All sixteen sub-optics (corresponding to 8 positive and 8 negative convolutional kernels) may be fabricated on a single substrate, allowing convolution from 16 different kernels in a single image capture. In some embodiments, each kernel meta-optic is 470 ⁇ m ⁇ 470 ⁇ m and has a center-to-center distance of 705 ⁇ m.
  • the kernel meta-optics may be arranged in two rows of eight meta- optics for a total footprint of 5.6 mm ⁇ 1.4 mm.
  • the positive and negative images can be computationally subtracted after being acquired by the camera 200 to produce the net convolution.
  • FIG.5B shows phase maps and scanning electron microscope (SEM) images of exemplary meta-optics corresponding to the positive and negative convolutional kernels in accordance with embodiments of the present technology.
  • the phase maps and SEM images are illustrated for a positive kernel P4 and a negative kernel N4, but this should be understood as a sample choice for illustrative purposes, and other positive and negative kernels can also be characterized by their phase maps and SEM images.
  • the phase maps are implemented using silicon nitride pillars which are 750 nm tall and having varied widths resulting in their different relative phase delay. Scanning electron microscope (SEM) images of the fabricated optics are also shown for comparison, highlighting the fabrication quality.
  • FIG.5C shows the positive and negative parts of an example convolutional kernel (left) when an electronic kernel is used, and the corresponding PSF simulation results (middle) and PSF experimental results (right) when kernels of the meta-optics 100 are used in accordance with embodiments of the present technology.
  • Exemplary electronic convolutional kernels are shown in the leftmost column for two sample kernels P4 and N4.
  • the electronic convolutional kernels (left column) represent the ground truth PSFs for comparison with the optically-implemented kernels of the meta-optics 100.
  • the middle column represents the simulated PSFs of the meta-optics that may be obtained using angular spectrum propagation or another numerical simulation.
  • the rightmost column represents experimentally measured PSF for the kernels of meta-optics 100.
  • a comparison of the simulated PSFs in the middle column and the experimental PSFs in the rightmost column shows a relatively accurate match, confirming the fabrication accuracy.
  • FIG.5D shows simulated electronic output (left) and experimental optical convolved output (right) for the example kernels in accordance with embodiments of the present technology.
  • a calibration layer 130 may be utilized, as, for example, shown in FIG. 4C above.
  • the calibration layer 130 is configured for assigning weights to the convolved images, where the weights are determined by comparing electronically convolved images, as in FIG. 4B, with the captured convolved images of FIG. 4C.
  • this calibration layer 130 can be fine-tuned with only 10% of the training dataset thus ensuring that the computational backend does not need to be retrained.
  • the electronic backend of the hybrid network 1003 may include the calibration layer 130 followed by the original compressed electronic backend 400. Furthermore, the number of multiply-accumulate (MAC) operations in the entire hybrid network 1003 can be reduced by over two orders of magnitude through compressing multiple convolutional layers 300 of the teacher network 1001 into a single layer 300 of the student network 1002 to facilitate implementation of the optical convolutional layer by the meta-optics 100 of the hybrid network 1003.
  • the classification accuracy of the compressed hybrid CNN 1003 may still be very high. For example, comparison of the three CNN architectures (multi-layer electronic network 1001, compressed electronic network 1002, and hybrid optical-electronic network 1003) is shown in Table 1 below.
  • the multi- layer electronic network 1001 (a modified AlexNet in the illustrated case) achieves classification accuracy exceeding 98% on both training and testing datasets.
  • the number of MAC operations of this network is 688 million with 8-bit precision.
  • the compressed all- electronic network 1002 achieves greater than 96% accuracy; this slight decline reflects the inherent challenges of compressing multiple layers into a single layer. Primarily due to the compression of the convolution layers, the number of MAC operations is reduced to 228,672 with the compressed electronic network 1002.
  • the hybrid network 1003 which integrates the optical convolution layer with the calibration layer and a single fully connected layer electronic backend, experimentally achieves classification accuracy of 93.9% ( ⁇ 0.25%) and 93.4% ( ⁇ 0.22%) on the training and testing datasets, respectively, and requires only 85,824 MAC operations, which is 0.01% and 37% of that required for AlexNet and the compressed electronic networks, respectively.
  • T able 1 Classification Results
  • FIG. 6A is a of a convolutional neural network (CNN) for image classification based on all-electronic multi-layered CNN in accordance with embodiments of the present technology.
  • CNN convolutional neural network
  • the original CNN includes five convolutional layers 300 and three max pooling (also referred to as electronic) layers 400 at the front, followed by three fully-connected layers 400 at the end, while nonlinear activation functions ”ReLU” may be placed in each layer.
  • the source images being processed come from the image display 52 and the camera 200.
  • the camera 200 can segregate the acquired image into red (R), green (G), and blue (B) datasets.
  • F IG. 6B is a schematic view of a CNN for image classification based on all- electronic compressed (student) CNN in accordance with embodiments of the present technology.
  • F IG. 6C is a schematic view of a CNN for image classification based on a hybrid CNN which combines optical meta-optic front end 100 and electronic backend 400 without needing an electronic convolutional layer.
  • This architecture includes a single convolutional layer based on meta-optics 100 without needing an electronic convolution layer.
  • the convolution layer is realized using an array of kernels on the meta-optics 100, where each kernel performs a separate convolution for each color channel (e.g., each kernel performing separate convolutions for green, blue, and red light).
  • each kernel performs a separate convolution for each color channel (e.g., each kernel performing separate convolutions for green, blue, and red light).
  • two fully-connected electronic layers 400 are used based on knowledge distillation, but in other embodiments other numbers of fully- connected electronic layers 400 may be used, including a single fully-connected electronic layer.
  • the compression of an original CNN into the meta-optics 100 is a basis for realizing the optical encoder scheme, there are practical trade-offs to consider.
  • the physical size of the sensor (camera) 200 and meta-optics 100 limit the number of kernels and kernel size.
  • an optimal number and size of the kernels while compressing the original CNN during the CIFAR-10 classification task is 16 kernels of 7 ⁇ 7 size, resulting in the training and testing accuracy from the compressed all-digital CNN being 76.24 ⁇ 0.31% and 75.90 ⁇ 0.30%, respectively. Therefore, the kernels operate to produce convolution results that are 2-dimensional convolution layers. Since the CIFAR-10 images have three channel information corresponding to red (R), green (G), and blue (B) – the inventive hybrid CNN may have 16 individual 7 ⁇ 7 kernels for each color, for a total of 48 kernels.
  • each kernel is physically separated into its positive and negative parts, resulting in a total of 32 meta-optics kernels, constituting 16 positive and 16 negative polychromatic kernels, each kernel operating over the RGB wavelengths (i.e., each kernel operating for R, G, and B wavelenengths).
  • Another design parameter can be based on determining how many pixels on the camera should represent one pixel of the PSF. This design parameter is sometimes referred to as the enlargement factor.
  • the ground- truth PSF which is a 7 ⁇ 7 matrix corresponds to 14 ⁇ 14 pixels on the camera. While a large enlargement factor ensures less alignment error, the signal intensity on each camera pixel becomes lower, resulting in a lower signal-to-noise ratio.
  • the electronic backend 400 may include software that is a trained and reconfigurable artificial intelligence (AI).
  • FIG.7A is an isometric view of a meta-optic scatterer (silicon nitride pillar) 110 on a quartz substrate 115 in accordance with embodiments of the present technology.
  • the scatterers 110 are made of silicon nitride on a quartz substrate to ensure high transparency in the visible wavelength. That is, the scatterers 110 are designed to operate within the entire light spectrum interest, thus making the scatterers 110 and, by extension, the meta-optics 100 suitable for executing a polychromatic convolution in a single step.
  • the polychromatic operation of the meta-optics 100 can be well represented by the R, G, and B components of the light.
  • wavelengths ( ) of 450, 532, and 635 for RGB colors may be utilized based on the availability of the laser diodes.
  • the illustrated scatterer 110 is characterized by a height t and width w, but other shapes of the scatterers are also possible in different embodiments.
  • FIG.7B shows graphs of relative phase shift (left vertical axes) and transmission (right vertical axes) of the uniform array of pillars (also referred to as the nanoposts or scatterers) with respect to the pillar width, w, for different RGB wavelengths in accordance with embodiments of the present technology.
  • the horizontal axis indicates width (w) of the pillar 110.
  • the phase shifts as experienced by the light passing through the scatterer 110 are shown in radians on the left vertical axis.
  • the solid line next to the dotted line is the fitted proxy function of the phase shift with respect to the w.
  • the transmission coefficients (right vertical axis) and phase shifts from silicon nitride pillars for RGB wavelengths are shown as a function of the pillar width, , at the fixed height of 800 nm, obtained by rigorous coupled-wave analysis (RCWA).
  • RCWA rigorous coupled-wave analysis
  • the first term corresponds to a general phase shift from a dielectric waveguide, where and are the effective refractive index and height of the silicon nitride pillars.
  • the second term only having the variance, corresponds to a correction term in the Gaussian shape with , , and as a fitting parameters.
  • Lastly, corresponds to a phase shift offset, making the (0) 0.
  • FIG.7C illustrates design flow of the polychromatic meta-optics optimization in accordance with embodiments of the present technology.
  • the illustrated design flow for the polychromatic RGB meta-optics produces optimized PSFs for individual RGB colors.
  • a meta-optics 100 parameterized by an arbitrary two- dimensional pillar width map, ( , ) we can extract three separate phase maps using the proxy functions .
  • FIG.8A illustrates fabricated optical encoder (meta-optics 100) having 16 positive convolutional kernels 152, 16 negative convolutional kernels 154, and 5 alignment metalenses in accordance with embodiments of the present technology.
  • the alignment metalenses 5 are focusing light at the focal plane to aid in the alignment (e.g., tilt, rotation, and distance) between the meta-optics and the camera.
  • different numbers of the convolutional and alignment lenses may be used.
  • FIG.8B is a schematics of a measurement setup for polychromatic PSFs in accordance with embodiments of the present technology.
  • beams of coherent light at different wavelengths e.g., at individual RGB wavelengths
  • a pinhole 50 of 25 diameter creates an approximate point source and the positions of the optics, i.e., pinhole, meta-optics, and camera, remain the same while changing the laser diodes to produce different wavelenghts (e.g., B at 450nm, G at 532 nm, R at 635 nm).
  • Different PSFs can be captured by the camera 200 for further processing.
  • FIG.8C illustrates ground-truth and measured RGB PSFs for a particular polychromatic kernel in accordance with embodiments of the present technology.
  • the knowledge distillation algorithm is designed to compress neural networks.
  • AlexNet pre-trained teacher network
  • the student network comprises only a single convolutional layer coupled with a backend, which consists of a single fully connected layer and a linear calibration layer.
  • AlexNet was the foundational model that successfully addressed the ImageNet datase; and second, compared to more complex networks like ResNet-18 or VGG-16, the five-layer AlexNet is more accessible and easier to implement optically.
  • the knowledge distillation algorithm may include two types of losses: student loss and temperature loss. Student loss minimizes the discrepancy between the student network’s predictions and the ground truth labels.
  • FIG.8D is a schematics of meta-optics convolved image measurement setup in accordance with embodiments of the present technology.
  • the polychromatic optical encoder is tested for the CIFAR-10 dataset.
  • FIG.8E shows electronically convolved images (upper row) and optically convolved images (lower row) of a particular CIFAR-10 image in individual RGB colors in accordance with embodiments of the present technology.
  • the meta-optically convolved image loses some of the high resolution components, likely due to the imperfect fabrication and alignment errors which are already recognizable from the PSF measurements, as well as because of the spectral overlap between RGB color pixels of the camera.
  • imperfections may be overcome-able based on the computational backend that is robust against such discrepancy, as we do average pooling of the convolved image into 6 ⁇ 6 size.
  • the additional fully-connected calibration layer 130 assigns the weights of each kernel and color, dealing with the discrepancy between optical/digital systems (e.g., normalization, scaling, translation, rotation, tilt, noise). Therefore, the calibration layer 130 allows the use of pre-trained digital backend while incurring minimal computational cost.
  • convolutional meta-optics 100 implements particular convolutional kernels which come from compressed CNN for CIFAR-10 data. Unlike electronic neural networks, optical implementations are difficult to modify once they are fabricated. This necessitates different convolutional meta-optics for different dataset, e.g., CIFAR100, ImageNet, etc. However, the inventors have found that the convolutional layer that is optimized for CIFAR-10 can be readily adapted to classify another dataset, i.e., High-10, with a transfer learning process. Therefore, in some embodiments, an additional fully-connected layer, called a “transfer learning layer” is added in between the former fully-connected layers and convolutional layer.
  • transfer learning layer is added in between the former fully-connected layers and convolutional layer.
  • the transfer learning layer By training the transfer learning layer, we can fit the other dataset, i.e., High-10, to the hybrid CNN which is pre-optimized for a particular dataset, i.e., CIFAR-10, without re-training the former fully-connected layers.
  • the High-10 image dataset has 224 ⁇ 224 size and polychromatic (RGB) information.
  • RGB polychromatic
  • a calibration function is added into the calibration layer 130 to remap the optical convolution outputs to align with those of the previously trained backend.
  • This approach aims to refine the experimental outputs to align more closely with the pre-designed network.
  • the training may be limited to only 20% of the available data, ensuring that the model remains efficient. Representative results are shown in Table 2 below.
  • Table 2 Classification Results on CIFAR-10 Dataset Network Architecture Train accuracy (%) Test accuracy (%) set classification tasks based on all-electronic multi-layered CNN, all-electronic compressed (student) CNN, and a hybrid CNN which combines a meta-optiocs front end and electronic backend in accordance with embodiments of the present technology. Even though there are slight differences between the optical and digital convolution results (FIG. 8E), after introducing the calibration layer 130, the inventive technology can achieve similar accuracy (less than 5% loss) for both the training and testing datasets. Additionally, the hybrid approach illustrated in FIG.9C significantly reduces computational costs which can be represented by the number of multiply-accumulate (MAC) operations.
  • MAC multiply-accumulate

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Biomedical Technology (AREA)
  • Biophysics (AREA)
  • General Physics & Mathematics (AREA)
  • General Health & Medical Sciences (AREA)
  • Evolutionary Computation (AREA)
  • Data Mining & Analysis (AREA)
  • Molecular Biology (AREA)
  • Computing Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Mathematical Physics (AREA)
  • Software Systems (AREA)
  • Artificial Intelligence (AREA)
  • Neurology (AREA)
  • Optics & Photonics (AREA)
  • Image Analysis (AREA)

Abstract

Compressed meta-optical encoder for image classification, and associated systems and methods are presented. In one embodiment, a convolutional neural network (CNN) system for image classification includes a meta-optics having a plurality of convolutional kernels. Each convolutional kernel includes a plurality of nanostructures. The plurality of kernels are configured for convolving source images by applying their respective point spread functions (PSFs) to the source images. The system also includes an image capturing camera configured for capturing convolved images, and an electronic backend software configured for classifying the source images based on the convolved images.

Description

COMPRESSED META-OPTICAL ENCODER FOR IMAGE CLASSIFICATION CROSS-REFERENCE TO RELATED APPLICATION This application claims the benefit of U.S. Provisional Application No.63/568722, filed March 22, 2024, and U.S. Provisional Application No. 63/721961, filed November 18, 2024; the entire disclosures of which are hereby incorporated by reference. STATEMENT OF GOVERNMENT LICENSE RIGHTS This invention was made with Government support under Grant Nos. EFRI- BRAID-2223495, NSF-ECCS-2127235, NNCI-1542101 and NNCI-2025489 awarded by the National Science Foundation. The Government has certain rights in the invention. BACKGROUND Visual information plays a crucial role in human response, particularly in situations where reaction time is limited to a few tens to hundreds of milliseconds. Though the human brain has efficiency far exceeding that of any other human-made computing systems, it still cannot process the entire collected visual data due to its massive amount of information. Most likely, our brain performs preliminary visual processing to extract essential features for efficient and rapid interpretation without handling the entire visual data. With the dramatic development of artificial intelligence (AI), computers can process the visual information like human eyes, thanks to artificial neural network (ANN), enabling computer/machine vision. Despite impressive progress, realtime inference remains very challenging even with faster processors and more efficient algorithms. For example, in a flying object (i.e., habitat drones) on-site data processing is plagued by severe heating, battery capacity and weight handling challenges. Utilizing cloud based systems for storing and processing the data poses challenges associated with data security and additional data transfer latency. 3915-P1339WO.UW -1- Accordingly, methods and systems for improved classification and processing of visual information are still needed. SUMMARY This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This summary is not intended to identify key features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. Convolutional neural networks (CNNs), for example, AlexNet, represent a significant milestone in image classification, recognition, and tracking. CNNs are composed of several electronics convolutional layers that adaptively learn spatial representations from input images. While powerful, the electronics convolution operation is computationally expensive, leading to high latency and power consumption. In fact, it has been estimated that about 80% of the total runtime of CNNs is used performing convolution operations. Reducing this latency and power consumption may include free- space optical systems as a solution. Beyond power consumption and latency reduction, optical information processing may be characterized by high bandwidth, spatial parallelism, and low-loss transmission. However, a challenge to all CNN approaches is that non- linear layers are interspersed with the linear layers. For example, AlexNet, a known baseline structure in computer vision challenges, consists of five convolutional layers followed by three fully connected layers (electrical layers), employing ReLU activations. It is known that without the nonlinear layers, the classification accuracy of the CNN may drop by about 20%. Generally, the nonlinear layers cannot be implemented using simple lens like optics; to implement them optically, some physical nonlinearity must be introduced, for instance by using an atomic vapor cell or image intensifier. On the other hand, hybrid approaches involving repeated transduction of the signal (e.g., optical images) to perform linear operations in optics and non-linear operations in electronics provide little benefit due to large latency and power consumption in signal transduction. Implementing only one of many required convolution operations does not provide significant benefit in terms of speed and latency. On the other hand, if image classification is performed without using convolution, such an approach may be implement multiple linear layers in one optical frontend, however at the cost of being computationally expensive to train and being applicable only to the physical system for which they were specifically designed. With some embodiments of inventive technology, a hybrid CNN consisting of a single optical convolution layer with an electronic single fully connected layer is experimentally demonstrated. To overcome the limitation from the absence of optical nonlinearity, we apply knowledge distillation to remove the nonlinear layers, and compress multiple layers to a single linear layer. Knowledge distillation (also referred to as ‘compression’ from a ‘teacher’ network to a ‘student’ network) may circumvent the need for nonlinearity without a significant reduction in the performance by transferring knowledge from a larger, pre-trained network (the ‘teacher’ network) to a more compact network (the ‘student’ network). In some embodiments, AlexNet is the teacher network and a single convolutional layer coupled with a single fully connected layer is the student network. In some embodiments, inventive technology uses a hybrid meta-optical platform, where an optical frontend is based on a single meta-optic performing the convolution operation, followed by an electronic backend which contains a linear calibration layer and a fully connected layer. In this way, the most computationally expensive operation is performed optically to leverage the benefits of optical computing, namely high spatial bandwidth and low power consumption. The use of a single meta-optic drastically simplifies the experimental setup and provides compact and efficient geometry. The optical frontend of the network may be realized using inverse-designed meta- optics. The meta-optic may include arrays of sub-wavelength scatterers which act as phase masks, imparting spatially-coded phase shifts to incident light. In some embodiments, the meta-optics is designed to realize a phase mask which performs the desired convolutional steps of the CNN by engineering the PSF. In some embodiments, we validate the performance of the designed optics using incoherent green light illumination from a light emitting diode, centered at 525 nm. The experiments demonstrate the classification accuracy of the entire hybrid CNN for the Modified National Institute of Standards and Technology (MNIST) dataset. In some embodiments, inventive technology achieves a classification accuracy of 93%, which is comparable to the 96% accuracy achieved using an all-electronic network compressed using knowledge distillation. By compressing multiple convolutional layers into a single layer and implementing this layer optically, the number of multiply-accumulation (MAC) operations in the entire network may be reduced by four orders of magnitude. In some embodiments, a polychromatic optical encoder with PSF-engineered meta- optic is used to classify the Canadian Institute For Advanced Research in 10 classes (CIFAR-10) dataset. This makes the inventive system compact and fully compatible with the conventional optical imaging systems. This ability to generalize the frontend as a passive meta-optic is helpful for any ANNs as it enhances their versatility, efficiency, and robustness. A network that generalizes well can be applied to different tasks without extensive re-training, saving time, reducing costs for meta-surface fabrications, and conserving computational resources for real-world AI applications. In one embodiment, a convolutional neural network (CNN) system for image classification is disclosed. The system includes a meta-optics having a plurality of convolutional kernels, each convolutional kernel having a plurality of nanostructures. The plurality of kernels are configured for convolving source images by applying their respective point spread functions (PSFs) to the source images. An image capturing camera is configured for capturing convolved images. An electronic backend software configured for classifying the source images based on the convolved images. In one embodiment, convolving images includes multi-wavelength kernel convolutions that are executed simultaneously. In one embodiment, the multi-wavelength kernel convolutions include convolutions of a red (R) wavelength, a green (G) wavelength, and a blue (B) wavelength of light. In one embodiment, the plurality of kernels are distributed over a monolithic meta- optics. In one embodiment, the PSFs of the plurality of kernels are at least in part obtained from a knowledge distillation that includes compressing operations of multiple electronic convolutional layers into an operation of one electronic convolutional layer. In one embodiment, the PSFs are determined by simulation prior to manufacturing of the meta-optics. In one embodiment, the electronic backend software includes a calibration layer configured for compensating fabrication and misalignment errors of the meta-optics. In one embodiment, the calibration layer is configured for assigning weights to the convolved images, where the weights are determined by comparing the convolved images obtained by the meta-optics with the convolved images obtained by electronic convolution. In one embodiment, the nanostructures are made of silicon nitride on a quartz substrate. In one embodiment, the electronic backend software is a trained and reconfigurable artificial intelligence (AI). In one embodiment, the convolved images are equivalent to an output of 2- dimensional convolutional layers. In one embodiment, the system also includes a source of light configured to emit in a visible spectrum. In one embodiment, a method for convolutional neural network (CNN) image classification includes illuminating a meta-optics having a plurality of convolutional kernels with source images. Each convolutional kernel includes a plurality of nanostructures. The method also includes convolving the source images by applying point spread functions (PSFs) of the convolutional kernels to the source images; and capturing convolved images by an image capturing camera. Based on the convolved images, the method classifies the source images with an electronic backend software. DESCRIPTION OF THE DRAWINGS The foregoing aspects and many of the attendant advantages of this invention will become more readily appreciated as the same become better understood by reference to the following detailed description, when taken in conjunction with the accompanying drawings, where: FIGURE 1 is a scanning electron micrograph image of a meta-optics in accordance with an embodiment of the present technology; FIGURES 2A-2C illustrate several views of meta-optic’s nanoposts in accordance with embodiments of the present technology; FIGURE 3 illustrates operation of a meta-optic according to embodiments of the present technology; FIGURE 4A is a schematic view of a convolutional neural network (CNN) for image classification based on all-electronic multi-layered CNN in accordance with embodiments of the present technology; FIGURE 4B is a schematic view of a CNN for image classification based on all- electronic compressed CNN in accordance with embodiments of the present technology; FIGURE 4C is a schematic view of a CNN for image classification based on a hybrid CNN which combines optical meta-optic front end with electronic backend in accordance with embodiments of the present technology; FIGURE 5A illustrates a point spread function (PSF) measurement setup using a monochromatic point light source (left) and optical convolution testing using a micro-LED display (right) in accordance with embodiments of the present technology; FIGURE 5B shows phase maps and scanning electron microscope (SEM) images of exemplary meta-optics corresponding to sample positive and negative convolutional kernels in accordance with embodiments of the present technology; FIGURE 5C shows the positive and negative parts of an example electronic convolutional kernel (left) and the corresponding meta-optics PSF simulation (middle) and meta-optics PSF experiment (right) in accordance with embodiments of the present technology; FIGURE 5D shows simulated electronic output (left) and optical experiment (right) convolved output for the example kernel in accordance with embodiments of the present technology; FIGURE 6A is a schematic view of a convolutional neural network (CNN) for image classification based on all-electronic multi-layered CNN in accordance with embodiments of the present technology; FIGURE 6B is a schematic view of a CNN for image classification based on all- electronic compressed (student) CNN in accordance with embodiments of the present technology; FIGURE 6C is a schematic view of a CNN for image classification based on a hybrid CNN which combines optical meta-optic front end and electronic backend in accordance with embodiments of the present technology; FIGURE 7A is an isometric view of a meta-optic scatterer (silicon nitride pillar) on a quartz substrate in accordance with embodiments of the present technology; FIGURE 7B shows graphs of relative phase shift and transmission of the uniform array of pillars with respect to the pillar width, w, for different RGB wavelengths in accordance with embodiments of the present technology; FIGURE 7C illustrates design flow of the polychromatic meta-optics optimization in accordance with embodiments of the present technology; FIGURE 8A illustrates fabricated optical encoder having 16 positive and 16 negative convolutional kernels and 5 alignment metalenses in accordance with embodiments of the present technology; FIGURE 8B is a schematics of polychromatic PSFs measurement setup in accordance with embodiments of the present technology; FIGURE 8C illustrates ground-truth and measured RGB PSFs for a particular polychromatic kernel in accordance with embodiments of the present technology; FIGURE 8D is a schematics of meta-optic convolved image measurement setup in accordance with embodiments of the present technology; FIGURE 8E shows electronically convolved images (upper row) and optically (lower row) convolved images of a particular CIFAR-10 image in individual RGB colors in accordance with embodiments of the present technology; FIGURE 9A show a confusion matrix of CIFAR-10 dataset classification tasks based on all-electronic multi-layered CNN in accordance with embodiments of the present technology; FIGURE 9B show a confusion matrix of CIFAR-10 dataset classification tasks based on all-electronic compressed (student) CNN in accordance with embodiments of the present technology; and FIGURE 9C show a confusion matrix of CIFAR-10 dataset classification tasks based on a hybrid CNN which combines and optical meta-optic front end and electronic backend in accordance with embodiments of the present technology. DETAILED DESCRIPTION While illustrative embodiments have been illustrated and described, it will be appreciated that various changes can be made therein without departing from the spirit and scope of the invention. Before explaining at least one embodiment of the presently disclosed and/or claimed inventive concept(s) in detail, it is to be understood that the presently disclosed and/or claimed inventive concept(s) is not limited in its application to the details of construction and the arrangement of the components or steps or methodologies set forth in the following description. The presently disclosed and/or claimed inventive concept(s) is capable of other embodiments or of being practiced or carried out in various ways. Also, it is to be understood that the phraseology and terminology employed herein is for the purpose of description and should not be regarded as limiting. Unless otherwise defined herein, technical terms used in connection with the presently disclosed and/or claimed inventive concept(s) shall have the meanings that are commonly understood by those of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. All patents, published patent applications, and non-patent publications mentioned in the specification are indicative of the level of skill of those skilled in the art to which the presently disclosed and/or claimed inventive concept(s) pertains. All patents, published patent applications, and non-patent publications referenced in any portion of this application are herein expressly incorporated by reference in their entirety to the same extent as if each individual patent or publication was specifically and individually indicated to be incorporated by reference. All of the articles and/or methods disclosed herein can be made and executed without undue experimentation in light of the present disclosure. While the articles and methods of the presently disclosed and/or claimed inventive concept(s) have been described in terms of preferred embodiments, it will be apparent to those skilled in the art that variations may be applied to the articles and/or methods and in the steps or in the sequence of steps of the method described herein without departing from the concept, spirit, and scope of the presently disclosed and/or claimed inventive concept(s). As utilized in accordance with the present disclosure, the following terms, unless otherwise indicated, shall be understood to have the following meanings. The use of the word “a” or “an” when used in conjunction with the term “comprising” may mean “one”, but it is also consistent with the meaning of “one or more”, “at least one”, and “one or more than one”. The use of the term “or” is used to mean “and/or” unless explicitly indicated to refer to alternatives only if the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives “and/or”. Throughout this application, the term “about” is used to indicate that a value includes the inherent variation of error for the quantifying device, the method being employed to determine the value, or the variation that exists among the study subjects. For example, but not by way of limitation, when the term “about” is utilized, the designation value may vary by plus or minus twelve percent, or eleven percent, or ten percent, or nine percent, or eight percent, or seven percent, or six percent, or five percent, or four percent, or three percent, or two percent, or one percent. The use of the term “at least one” will be understood to include one as well as any quantity more than one, including but not limited to, 2, 3, 4, 5, 10, 15, 20, 30, 40, 50, 100, etc. The term “at least one” may extend up to 100 or 1000 or more, depending on the term to which it is attached; in addition, the quantities of 100/1000 are not to be considered limiting, as lower or higher limits may also produce satisfactory results. In addition, the use of the term “at least one of X, Y, and Z” will be understood to include X alone, Y alone, and Z alone, as well as any combination of X, Y, and Z. The use of ordinal number terminology (i.e., “first”, “second”, “third”, “fourth”, etc.) is solely for the purpose of differentiating between two or more items and is not meant to imply any sequence or order or importance to one item over another or any order of addition, for example. As used herein, the words “comprising” (and any form of comprising, such as “comprise” and “comprises”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “includes” and “include”) or “containing” (and any form of containing, such as “contains” and “contain”) are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. The term “or combinations thereof” as used herein refers to all permutations and combinations of the listed items preceding the term. For example, “A, B, C, or combinations thereof” is intended to include at least one of: A, B, C, AB, AC, BC, or ABC and, if order is important in a particular context, also BA, CA, CB, CBA, BCA, ACB, BAC, or CAB. Continuing with this example, expressly included are combinations that contain repeats of one or more item or term, such as BB, AAA, AB, BBC, AAABCCCC, CBBAAA, CABABB, and so forth. The skilled artisan will understand that typically there is no limit on the number of items or terms in any combination, unless otherwise apparent from the context. In the context of this disclosure, the terms “about,” “approximately,” “generally” and similar mean +/- 5% of the stated value. FIG.1 is an optical image of a meta-optics (also referred as a metalens or meta- optic encoder) in accordance with an embodiment of the present technology. Illustrated meta-optics 100 includes a number of nanostructures (also referred to as nanoposts or scatterers) 110 that are carried by a substrate (also referred to as a carrier) 115. The nanostructures 110 may be nanoscale structures that are generally cylindrical or rectangular, and are characterized by one or more characteristic scales (e.g., cylinder diameter d, width w, height t, etc.). In some embodiments, the nanostructures 110 may have different sizes, as illustrated in FIG.1. In different embodiments, the meta-optics 100 may be manufactured by the process described below. In some embodiments, during the manufacturing of the meta-optics 100, a 600 nm layer of silicon nitride is first deposited via plasma-enhanced chemical vapor deposition (PECVD) on a quartz substrate, followed by spin-coating with a high-performance positive electron beam resist (e.g., ZEP-520A). An 8 nm Au/Pd charge dissipation layer is then sputtered followed by subsequent exposure to an electron-beam lithography system (e.g., JEOL JBX6300FS). The Au/Pd layer may then be removed with a thin film etchant (e.g., type TFA gold etchant), and the samples may be developed in amyl acetate. In some embodiments, to form an etch mask, 50 nm of aluminum is evaporated and lifted off via sonication in methylene chloride, acetone, and isopropyl alcohol. The samples are then dry etched using a CHF3 and SF6 chemistry and the aluminum is removed by immersion in AD-10 photoresist developer. In other embodiments, other manufacturing processes are possible. FIGS.2A-2C illustrate several views of nanoposts in accordance with embodiments of the present technology. FIG. 2A is an isometric view of a nanopost 110 that is carried by a substrate 115. The illustrated nanopost 110 is cylindrical, but in other embodiments the nanopost 110 may have other shapes, for example, an elliptical cross-section, a square cross-section, a rectangular cross-section or other cross-sectional shape that maintain center-to-center spacing at a sub-wavelength value. FIG.2B is a top view of two adjacent nanoposts that are separated by a distance “p” (pitch). Only two nanoposts are illustrated in FIG.2B for simplicity. However, for a practical meta-optics 100, many more nanoposts are distributed over the substrate 115. FIG. 2C is a side view of a nanopost 110 that is carried by a substrate 115. In some embodiments, the nanoposts (scatterers, nanostructures) 110 are made of silicon nitride due to its broad transparency window and CMOS compatibility. The illustrated nanoposts 110 are characterized by a height “t” and diameter “d”. In some embodiments, the values of “d” may range from about 100 nm to about 300 nm. Generally, the value of “t” (height) is constant (within the limits of manufacturing tolerance) for all diameters “d” for a given metalens. In some embodiments, the values of “t” may range from about 500 nm to about 800 nm. The nanoposts (scatterers) may be polarization-insensitive cylindrical nanoposts 110 arranged in a square lattice on a quartz substrate 115. The phase shift mechanism of these nanoposts arises from an ensemble of oscillating modes within the nanoposts that couple amongst themselves at the top and bottom interfaces of the post. By adjusting the diameter “d” of the nanoposts, the modal composition varies, modifying the transmission coefficient through the nanoposts. FIG.3 illustrates operation of meta-optics 100. In operation, an image is projected onto the meta-optics 100 as an optical input. The meta-optics 100 convolves the input image, and produces an output image that is projected-on and acquired-by an optical sensor 180 (e.g., a digital camera). The convolution operation of the meta-optics 100 modifies the output image by the point spread function (PSF) of the meta-optics, which, in turn, may be designed to optimize a task of image classification. FIG.4A is a schematic view of a convolutional neural network (CNN) 1001 for image classification based on all-electronic multi-layered CNN in accordance with embodiments of the present technology. FIG.4B is a schematic view of a CNN for image classification based on compressed all-electronic CNN 1002 in accordance with embodiments of the present technology. FIG.4C is a schematic view of a CNN 1003 for image classification based on a hybrid CNN which combines an optical meta-optic front end with electronic backend in accordance with embodiments of the present technology. In different embodiments, the CNNs 1001 and 1002 take input 11 that is a digitized image (e.g., an image acquired by a digital camera, coupled with a capability of assigning numerical values to individual pixels of the digital camera like a digital image of a scene where the values represented at each pixel of the image correspond to intensity values of an imaged scene or an image wherein the values represented at each pixel of the image correspond to intensity values of an imaged scene). Such input is then electronically processed by multiple convolutional layers 300 and multiple fully connected layers 400 of the teacher network 1001, or by a single convolutional layer 300 and a single fully connected layer 400 of the student network 1002. The classification system in FIG. 4B (student network 1002) can be derived by compressing multiple convolution layers 300 of FIG. 4A (teacher network 1001) into a single convolution layer 300 of FIG.4B. By this process of knowledge distillation, a need for nonlinearity is circumvented without a significant reduction in the performance by transferring knowledge from a larger, pre-trained network (the “teacher” network) of FIG. 4A to a more compact network (the “student” network) of FIG.4B. The convolution layers 300 are electronic convolution layers, thus the teacher network 1001 and student network 1003 are both fully electronic networks. In some embodiments, a modified AlexNet, denoted AlexNet-Mod can be used as the teacher network of FIG. 4A, whereas a single convolutional layer 300 coupled with a single fully connected layer400 can be used as the student network of FIG.4B. FIG. 4C illustrates a hybrid meta-optics platform 1003, where an optical frontend based on a single meta-optics 100 performs the linear convolution operation of the light signal coming from the light source 10 (e.g., display showing a source image to be classified, a laser, a light emitting diode, a digital display with a representation of a digital image, a reflective or transmissive target with a representation of an object or scene, an ambient scene illuminated by sunlight or man-made sources containing objects to be detected or classified, etc.), followed by an electronic backend which contains a linear calibration layer 130 and a fully connected layer 400. In such a way, the most computationally expensive operation is performed optically by the meta-optics 100 to leverage the benefits of optical computing, namely high spatial bandwidth and low power consumption. In at least some embodiments, the use of a single meta-optic layer 100 drastically simplifies the experimental setup and provides a compact geometry for the system 1003. The optical frontend 100 of the network 1003 may be realized using inverse- designed meta-optics. In some embodiments, the meta-optics 100 includes arrays of sub- wavelength scatterers which act as phase masks, imparting spatially-coded phase shifts to incident light. In operation, these spatially-coded phase shifts to incident light impart a point spread function (PSF) onto the input image as the input image is convolved by the meta-optics 100. Therefore, the meta-optics 100 is designed to apply a phase mask which performs the desired convolutional steps of the CNN by engineering the required PSF. In some embodiments, such meta-optics 100 is fabricated and experimentally validated using incoherent green light illumination, e.g., light illumination from a light emitting diode, centered at 525 nm. However, in other embodiments, other sources of light are also possible, for example, lasers or light emitting diodes operating at different wavelengths. FIG.5A is a point spread function (PSF) measurement setup using a monochromatic point light source (left) and optical convolution using a micro-LED display (right) in accordance with embodiments of the present technology. As explained above with respect to FIGs. 4A and 4B, compressed network 1002 with a single convolutional layer can be obtained by distilling knowledge from a teacher network 1001 (e.g., AlexNet) to a student network of the desired structure. The knowledge distillation approach assumes that the teacher network is already trained and performs the desired task with high accuracy. For example, AlexNet as the teacher network 1001 may achieve over 98% classification accuracy for both training and testing MNIST datasets over repeated trials. A sample AlexNet consists of five electronical convolutional layers 300 and three fully connected electronic layers 400 (total of 8 layers) which can be compressed to the student system 1002 having one electronical convolutional layer 300 and one fully connected layer 400 (two layers total). Additionally, AlexNet teacher network 1001 may employ nonlinear activation functions (ReLU) to optimize performance. In the sample hybrid network 1003 shown in FIG.5A, the meta-optics 100 (i.e., the compressed convolutional layer) is divided into positive and negative kernels, for example 8 positive and 8 negative kernels that are characterized by their individual PSFs, as further explained below. In some embodiments, each kernel may be 6 × 6 pixels in size (as captured on the pixels of the camera 200), but other pixel numbers allocated to each positive/negative kernel as well as other numbers of kernels are also possible in different embodiments. Stated differently, the illustrated meta-optic 100 contains 16 different sub-optics (or kernels), which can be understood as segments of one monolithic meta-optic 100, the kernels being spatially distributed within a single convolution layer to operate in parallel for the classification tasks. The images acquired by the camera 200 may be forwarded to the fully connected (electronic backend) layer 400 for classification. The left-hand side illustration of the CNN in FIG. 5A includes a source of the light 10 (e.g., a fiber coupled LED or laser) that provides PSF-determining images to the camera 200 and further to the fully connected (electronic back-end) layer 400. The right-hand side illustration of the CNN in FIG. 5A includes a real image (in the illustrated sample case, an image of a digit ‘7’) that projects on the meta-optics 100 and further onto the camera 200. Thus, the left-hand side setup of the CNN 1003 in FIG.5A may be used for determining or testing the real PSFs (i.e., 8 PSFs for positive kernels and 8 PSFs for negative kernels) of a given meta-optics 100 using, for example, a monochromatic source of light 10. To verify the performance of the meta-optics, the PSF of each sub-optic kernel can be measured using a single-mode fiber 10 as a point light source. The PSFs from all 16 sub-optics can be simultaneously captured by the two-dimensional CMOS camera 200. The right-hand side illustration of the CNN 1003 in FIG. 5A produces convolved images captured by the camera 200, where the captured images represent f(x, y)*PSF(x,y) for the 16 kernels (8 positive and 8 negative), which a person of ordinary skill would understand to be convolutions of the source image of numeral ‘7’ originating from the display 52. Accordingly, the convolved images from all 16 sub- optics kernels can be captured simultaneously when the single mode fiber 10 is replaced with the display 52 that emits source images. Generally, the required meta-optics 100 can be inverse-designed if a target PSF is known. In some embodiments, the Gerchberg-Saxton (GS) phase retrieval method may be used to design phase masks that correspond to the optimized convolutional kernels. Specifically, each sub-optics may be designed to have a PSF which resembles the convolutional kernel. Since electronic convolutional kernels include both positive and negative values, each kernel may be separated into positive and negative parts and meta- optics can be designed for each PSF. All sixteen sub-optics (corresponding to 8 positive and 8 negative convolutional kernels) may be fabricated on a single substrate, allowing convolution from 16 different kernels in a single image capture. In some embodiments, each kernel meta-optic is 470 μm × 470 μm and has a center-to-center distance of 705 μm. In some embodiments, the kernel meta-optics may be arranged in two rows of eight meta- optics for a total footprint of 5.6 mm × 1.4 mm. In operation, the positive and negative images can be computationally subtracted after being acquired by the camera 200 to produce the net convolution. FIG.5B shows phase maps and scanning electron microscope (SEM) images of exemplary meta-optics corresponding to the positive and negative convolutional kernels in accordance with embodiments of the present technology. In particular, the phase maps and SEM images are illustrated for a positive kernel P4 and a negative kernel N4, but this should be understood as a sample choice for illustrative purposes, and other positive and negative kernels can also be characterized by their phase maps and SEM images. In some embodiments, the phase maps are implemented using silicon nitride pillars which are 750 nm tall and having varied widths resulting in their different relative phase delay. Scanning electron microscope (SEM) images of the fabricated optics are also shown for comparison, highlighting the fabrication quality. FIG.5C shows the positive and negative parts of an example convolutional kernel (left) when an electronic kernel is used, and the corresponding PSF simulation results (middle) and PSF experimental results (right) when kernels of the meta-optics 100 are used in accordance with embodiments of the present technology. Exemplary electronic convolutional kernels are shown in the leftmost column for two sample kernels P4 and N4. The electronic convolutional kernels (left column) represent the ground truth PSFs for comparison with the optically-implemented kernels of the meta-optics 100. The middle column represents the simulated PSFs of the meta-optics that may be obtained using angular spectrum propagation or another numerical simulation. The rightmost column represents experimentally measured PSF for the kernels of meta-optics 100. A comparison of the simulated PSFs in the middle column and the experimental PSFs in the rightmost column shows a relatively accurate match, confirming the fabrication accuracy. FIG.5D shows simulated electronic output (left) and experimental optical convolved output (right) for the example kernels in accordance with embodiments of the present technology. Due to the constraints on physically realizable PSFs, there may be notable differences between the ground truth PSFs (leftmost column) and experimentally measured PSFs (rightmost column). To correct for these differences, as well as to account for noise and misalignments in the optical system, a calibration layer 130 may be utilized, as, for example, shown in FIG. 4C above. In some embodiments, the calibration layer 130 is configured for assigning weights to the convolved images, where the weights are determined by comparing electronically convolved images, as in FIG. 4B, with the captured convolved images of FIG. 4C. In some embodiments, this calibration layer 130 can be fine-tuned with only 10% of the training dataset thus ensuring that the computational backend does not need to be retrained. Therefore, in some embodiments the electronic backend of the hybrid network 1003 may include the calibration layer 130 followed by the original compressed electronic backend 400. Furthermore, the number of multiply-accumulate (MAC) operations in the entire hybrid network 1003 can be reduced by over two orders of magnitude through compressing multiple convolutional layers 300 of the teacher network 1001 into a single layer 300 of the student network 1002 to facilitate implementation of the optical convolutional layer by the meta-optics 100 of the hybrid network 1003. The classification accuracy of the compressed hybrid CNN 1003 may still be very high. For example, comparison of the three CNN architectures (multi-layer electronic network 1001, compressed electronic network 1002, and hybrid optical-electronic network 1003) is shown in Table 1 below. The multi- layer electronic network 1001 (a modified AlexNet in the illustrated case) achieves classification accuracy exceeding 98% on both training and testing datasets. The number of MAC operations of this network is 688 million with 8-bit precision. The compressed all- electronic network 1002 achieves greater than 96% accuracy; this slight decline reflects the inherent challenges of compressing multiple layers into a single layer. Primarily due to the compression of the convolution layers, the number of MAC operations is reduced to 228,672 with the compressed electronic network 1002. The hybrid network 1003, which integrates the optical convolution layer with the calibration layer and a single fully connected layer electronic backend, experimentally achieves classification accuracy of 93.9% (± 0.25%) and 93.4% (± 0.22%) on the training and testing datasets, respectively, and requires only 85,824 MAC operations, which is 0.01% and 37% of that required for AlexNet and the compressed electronic networks, respectively. These test results are shown in Table 1 below. Table 1: Classification Results FIG. 6A is a of a convolutional neural network (CNN) for image classification based on all-electronic multi-layered CNN in accordance with embodiments of the present technology. In the illustrated embodiment, the original CNN (AlexNet) includes five convolutional layers 300 and three max pooling (also referred to as electronic) layers 400 at the front, followed by three fully-connected layers 400 at the end, while nonlinear activation functions ”ReLU” may be placed in each layer. The source images being processed come from the image display 52 and the camera 200. The camera 200 can segregate the acquired image into red (R), green (G), and blue (B) datasets. FIG. 6B is a schematic view of a CNN for image classification based on all- electronic compressed (student) CNN in accordance with embodiments of the present technology. Here, the AlexNet is compressed into one or two convolutional layers 300 and two fully-connected layers 400 using knowledge distillation method, which reduces the complexity of the architecture with a minimal compromise in accuracy. FIG. 6C is a schematic view of a CNN for image classification based on a hybrid CNN which combines optical meta-optic front end 100 and electronic backend 400 without needing an electronic convolutional layer. Here, we demonstrate a polychromatic optical encoder with PSF-engineered meta-optics 100 to classify sample source images. This architecture includes a single convolutional layer based on meta-optics 100 without needing an electronic convolution layer. The convolution layer is realized using an array of kernels on the meta-optics 100, where each kernel performs a separate convolution for each color channel (e.g., each kernel performing separate convolutions for green, blue, and red light). In the illustrated embodiment, two fully-connected electronic layers 400 are used based on knowledge distillation, but in other embodiments other numbers of fully- connected electronic layers 400 may be used, including a single fully-connected electronic layer. While the compression of an original CNN into the meta-optics 100 is a basis for realizing the optical encoder scheme, there are practical trade-offs to consider. On one hand, the physical size of the sensor (camera) 200 and meta-optics 100 limit the number of kernels and kernel size. On the other hand, small kernel size or small number of kernels may fail to classify the data effectively. The inventors have found that, in some embodiments, an optimal number and size of the kernels while compressing the original CNN during the CIFAR-10 classification task is 16 kernels of 7 × 7 size, resulting in the training and testing accuracy from the compressed all-digital CNN being 76.24 ± 0.31% and 75.90 ± 0.30%, respectively. Therefore, the kernels operate to produce convolution results that are 2-dimensional convolution layers. Since the CIFAR-10 images have three channel information corresponding to red (R), green (G), and blue (B) – the inventive hybrid CNN may have 16 individual 7 × 7 kernels for each color, for a total of 48 kernels. In some embodiments of the inventive technology, it is possible to design a single meta-optics 100 that produces three different PSFs (i.e., convolutional kernels) for RGB wavelengths. Since it may be difficult to create positive and negative weights on the camera in terms of light intensity at the same time, in some embodiments each kernel is physically separated into its positive and negative parts, resulting in a total of 32 meta-optics kernels, constituting 16 positive and 16 negative polychromatic kernels, each kernel operating over the RGB wavelengths (i.e., each kernel operating for R, G, and B wavelenengths). Another design parameter can be based on determining how many pixels on the camera should represent one pixel of the PSF. This design parameter is sometimes referred to as the enlargement factor. For example, when the enlargement factor is 2, the ground- truth PSF which is a 7 × 7 matrix corresponds to 14 × 14 pixels on the camera. While a large enlargement factor ensures less alignment error, the signal intensity on each camera pixel becomes lower, resulting in a lower signal-to-noise ratio. To determine the optimal enlargement factor, several meta-optics with different enlargement factors were tested for a particular kernel. Based on these tests, a generally useful enlargement factor of 2 was adopted for a meta-optic made of 3200 × 3200 scatterers. In some embodiments, the electronic backend 400 may include software that is a trained and reconfigurable artificial intelligence (AI). FIG.7A is an isometric view of a meta-optic scatterer (silicon nitride pillar) 110 on a quartz substrate 115 in accordance with embodiments of the present technology. In some embodiments, the scatterers 110 are made of silicon nitride on a quartz substrate to ensure high transparency in the visible wavelength. That is, the scatterers 110 are designed to operate within the entire light spectrum interest, thus making the scatterers 110 and, by extension, the meta-optics 100 suitable for executing a polychromatic convolution in a single step. In many practical applications, the polychromatic operation of the meta-optics 100 can be well represented by the R, G, and B components of the light. For example, wavelengths ( ) of 450, 532, and 635 for RGB colors may be utilized based on the availability of the laser diodes. The illustrated scatterer 110 is characterized by a height t and width w, but other shapes of the scatterers are also possible in different embodiments. FIG.7B shows graphs of relative phase shift (left vertical axes) and transmission (right vertical axes) of the uniform array of pillars (also referred to as the nanoposts or scatterers) with respect to the pillar width, w, for different RGB wavelengths in accordance with embodiments of the present technology. The horizontal axis indicates width (w) of the pillar 110. The phase shifts as experienced by the light passing through the scatterer 110 are shown in radians on the left vertical axis. The solid line next to the dotted line is the fitted proxy function of the phase shift with respect to the w. The transmission coefficients (right vertical axis) and phase shifts from silicon nitride pillars for RGB wavelengths are shown as a function of the pillar width, , at the fixed height of 800 nm, obtained by rigorous coupled-wave analysis (RCWA). In order to effectively model the wavelength-dependent effects of the meta-optics in a gradient descent-based optimization method, a fast and differentiable function is required to map between the pillar width and imparted phase. We define a proxy function inspired by the approximate phase shift of a dielectric waveguide with corrective factors that are fit to the RCWA simulation results. To calculate the phase shift ( , , and ) for RGB wavelengths with respect to the pillar width, , the proxy function can be defined as: ( ) = 2 / + exp(( ) / ) . (1) The first term corresponds to a general phase shift from a dielectric waveguide, where and are the effective refractive index and height of the silicon nitride pillars. The second term, only having the variance, corresponds to a correction term in the Gaussian shape with , , and as a fitting parameters. Lastly, corresponds to a phase shift offset, making the (0) = 0. This proxy function does not model the resonance- induced phase variations; however, one may not want to use those phase variations due to reduced amplitude and these resonances are expected to be less prominent in the fabricated devices due to the sidewall roughness. FIG.7C illustrates design flow of the polychromatic meta-optics optimization in accordance with embodiments of the present technology. In particular, the illustrated design flow for the polychromatic RGB meta-optics produces optimized PSFs for individual RGB colors. For a meta-optics 100 parameterized by an arbitrary two- dimensional pillar width map, ( , ), we can extract three separate phase maps using the proxy functions . We can then propagate the electromagnetic field using angular spectrum method to simulate the PSFs at the focal plane, e.g., 2.4 away from the meta- optics. At the focal plane, we can compare the computational ground-truth PSFs defined by the convolutional kernels obtained using knowledge distillation ( , ) and optically- simulated PSFs ( , ) at each RGB channels, where the channel-dependent losses are defined by the sum of squares of differences in each pixels: = , | , ( , ) , ( , )| . (2) Next, we can optimize the map of two-dimensional pillar width, i.e., meta-optics, for minimizing the net loss which defined as a root mean square of the losses at three different colors using the Adam optimizer in TensorFlow: min | / ( , ) | || . (3) Losses for all the kernels at three different colors can also be calculated. FIG.8A illustrates fabricated optical encoder (meta-optics 100) having 16 positive convolutional kernels 152, 16 negative convolutional kernels 154, and 5 alignment metalenses in accordance with embodiments of the present technology. The alignment metalenses 5 are focusing light at the focal plane to aid in the alignment (e.g., tilt, rotation, and distance) between the meta-optics and the camera. In different embodiments, different numbers of the convolutional and alignment lenses may be used. FIG.8B is a schematics of a measurement setup for polychromatic PSFs in accordance with embodiments of the present technology. By using different the laser diodes, beams of coherent light at different wavelengths (e.g., at individual RGB wavelengths) can be sequantially directed onto the camera 200 through the meta-optics 100 to experimentally characterize the polychromatic PSFs. In some embodiments, a pinhole 50 of 25 diameter creates an approximate point source and the positions of the optics, i.e., pinhole, meta-optics, and camera, remain the same while changing the laser diodes to produce different wavelenghts (e.g., B at 450nm, G at 532 nm, R at 635 nm). Different PSFs can be captured by the camera 200 for further processing. FIG.8C illustrates ground-truth and measured RGB PSFs for a particular polychromatic kernel in accordance with embodiments of the present technology. To quantitatively analyze the difference between the ground-truth (electronic) and experimentally measured (meta-optics) PSFs, we can define a cosine similarity ( ) as: = intensity profiles of the PSF for RGB wavelengths, respectively. The calculated for RGB wavelengths are about 0.88, 0.56, and 0.81, respectively. The quntitative discrepancy can be attributed partially to the fabrication and measurement imperfections. Additionaly, not all the polychromatic PSFs are physically realizable as the phases at different wavelengths are not completely independent. Creating more physically-realizable PSFs via co-designing the optical frontend and computational backend, also termed as end-to-end design, instead of replacing the convolutional layer with optics, may increase the . However, including the meta– optical simulation in the end-to-end design may result in a local optimum, and the fabrication/measurement imperfections may still be present. Nevertheless, in many embodiments, the computational backend is robust against such discrepancy in the PSFs, therefore correcting for these errors by introducing an additional fully-connected calibration layer 130 in the digital backend. Typically, the knowledge distillation algorithm is designed to compress neural networks. Here, we propose using knowledge distillation to transfer the generalized knowledge from a larger, pre-trained teacher network, AlexNet, to a more compact CNN, also referred to as the “student network." Specifically, the student network comprises only a single convolutional layer coupled with a backend, which consists of a single fully connected layer and a linear calibration layer. In addition, we can selected AlexNet as a teacher network for two main reasons: first, AlexNet was the foundational model that successfully addressed the ImageNet datase; and second, compared to more complex networks like ResNet-18 or VGG-16, the five-layer AlexNet is more accessible and easier to implement optically. The knowledge distillation algorithm may include two types of losses: student loss and temperature loss. Student loss minimizes the discrepancy between the student network’s predictions and the ground truth labels. The softmax function can be used to compute: = ( ) ( ) (5) where represents the student logits after the last fully connected layer. Temperature loss, on the other hand, optimizes the discrepancy between the student network’s predictions and the teacher network’s predictions. Knowledge distillation incorporates a softening parameter, , known as the distillation temperature for the teacher probabilities. Thus, we can compute such loss as: = ( / ) ( / ) (6) Finally, the total loss is ( , ) = ( , ) + (1 ) ; = , ; = loss function, is the Kullback-Leibler (KL) Divergence loss function. In different embodiments, other key hyperparameters might impact the meta- optiocs hybrid CNN system. First, in most CNNs, such as ResNet-18 and AlexNet, there are multiple convolutional layers, each with more than 200 kernels to extract useful features and maintain generalization across various datasets. While some pruning strategies show that using 1% of the parameters can achieve similar accuracy, applying these algorithms to optical neural networks are non-trivial. Most pruning methods still retain ANN structures, which suffer from misalignments that are difficult to eliminate. Therefore, compressing into shallower layers with more kernels is preferred. However, each meta-surface has a physical size that limits the number of kernels it can contain. To address this limitation, we can use multiple cameras and multiple meta-surfaces to increase the number of kernels, thereby improving the classification accuracy and generalization of the hybrid CNN. FIG.8D is a schematics of meta-optics convolved image measurement setup in accordance with embodiments of the present technology. Here, the polychromatic optical encoder is tested for the CIFAR-10 dataset. By replacing the pinhole with an organic light- emitting diode (OLED) display 52 that produces source images, we can convolve the CIFAR-10 images with the pre-characterized PSFs of the meta-optics 100. The displayed image size may be adjusted according to the convolutional kernel size and the enlargement factor on the camera. As explained above, in some embodiments the multiple kernels of the meta-optics 100 are each capable of processing different wavelengths of light. FIG.8E shows electronically convolved images (upper row) and optically convolved images (lower row) of a particular CIFAR-10 image in individual RGB colors in accordance with embodiments of the present technology. In many embodiments, the meta-optically convolved image loses some of the high resolution components, likely due to the imperfect fabrication and alignment errors which are already recognizable from the PSF measurements, as well as because of the spectral overlap between RGB color pixels of the camera. However, such imperfections may be overcome-able based on the computational backend that is robust against such discrepancy, as we do average pooling of the convolved image into 6 × 6 size. Furthermore, the additional fully-connected calibration layer 130 assigns the weights of each kernel and color, dealing with the discrepancy between optical/digital systems (e.g., normalization, scaling, translation, rotation, tilt, noise). Therefore, the calibration layer 130 allows the use of pre-trained digital backend while incurring minimal computational cost. In some embodiments, convolutional meta-optics 100 implements particular convolutional kernels which come from compressed CNN for CIFAR-10 data. Unlike electronic neural networks, optical implementations are difficult to modify once they are fabricated. This necessitates different convolutional meta-optics for different dataset, e.g., CIFAR100, ImageNet, etc. However, the inventors have found that the convolutional layer that is optimized for CIFAR-10 can be readily adapted to classify another dataset, i.e., High-10, with a transfer learning process. Therefore, in some embodiments, an additional fully-connected layer, called a “transfer learning layer” is added in between the former fully-connected layers and convolutional layer. By training the transfer learning layer, we can fit the other dataset, i.e., High-10, to the hybrid CNN which is pre-optimized for a particular dataset, i.e., CIFAR-10, without re-training the former fully-connected layers. For example, the High-10 image dataset has 224 × 224 size and polychromatic (RGB) information. To use the CNN that was optimized for CIFAR-10 data with the High-10 data, we can resize the High-10 images to 32 × 32 size, same as the CIFAR-10 data. Without applying a transfer learning method, the training and testing accuracy is expected to be around 40%. However, after transfer learning, a much higher training and testing accuracy (66%) is achievable on the High-10 data with the convolutional layer and fully-connected layers that is used for the CIFAR-10 dataset. The inventors have further experimentally verified that this approach works in the inventive hybrid optical/digital CNN using the same convolutional meta-optics that we used for the CIFAR-10 data electronic backend structure with one additional fully-connected layer. The average training and testing experiment accuracy of the High-10 data are similar (less than 5% loss) to the compressed all-electronic CNN, which is about the same of the CIFAR-10 case. As previously discussed, optical fabrication and alignment noise are unavoidable in meta-surface kernels. These include scaling, translation, rotation, image aberration, and optical noises. To address this issue, a calibration function is added into the calibration layer 130 to remap the optical convolution outputs to align with those of the previously trained backend. Specifically, a fully connected layer 130 may be used for the calibration function and corresponding loss function is defined as: = min( ( / , )) (8) This approach aims to refine the experimental outputs to align more closely with the pre-designed network. In some embodiments, oo prevent overfitting, the training may be limited to only 20% of the available data, ensuring that the model remains efficient. Representative results are shown in Table 2 below. Table 2: Classification Results on CIFAR-10 Dataset Network Architecture Train accuracy (%) Test accuracy (%) set classification tasks based on all-electronic multi-layered CNN, all-electronic compressed (student) CNN, and a hybrid CNN which combines a meta-optiocs front end and electronic backend in accordance with embodiments of the present technology. Even though there are slight differences between the optical and digital convolution results (FIG. 8E), after introducing the calibration layer 130, the inventive technology can achieve similar accuracy (less than 5% loss) for both the training and testing datasets. Additionally, the hybrid approach illustrated in FIG.9C significantly reduces computational costs which can be represented by the number of multiply-accumulate (MAC) operations. From the original CNN to compressed CNN, we can reduce the computational load, which is represented by a number of MAC operations, by a factor of 1,400, while the computational effort is reduced further by a factor of 17 after replacing a convolutional layer with meta-optics 100. From the foregoing, it will be appreciated that specific embodiments of the technology have been described herein for purposes of illustration, but that various modifications may be made without deviating from the disclosure. Moreover, while various advantages and features associated with certain embodiments have been described above in the context of those embodiments, other embodiments may also exhibit such advantages and/or features, and not all embodiments need necessarily exhibit such advantages and/or features to fall within the scope of the technology. Accordingly, the disclosure can encompass other embodiments not expressly shown or described herein.

Claims

CLAIMS What is claimed is: 1. A convolutional neural network (CNN) system for image classification, the system comprising: a meta-optics comprising a plurality of convolutional kernels, each convolutional kernel comprising a plurality of nanostructures, wherein the plurality of kernels are configured for convolving source images by applying their respective point spread functions (PSFs) to the source images; an image capturing camera configured for capturing convolved images; and an electronic backend software configured for classifying the source images based on the convolved images. 2. The system of claim 1, wherein convolving images comprises multi- wavelength kernel convolutions that are executed simultaneously. 3. The system of claim 2, wherein the multi-wavelength kernel convolutions comprise convolutions of a red (R) wavelength, a green (G) wavelength, and a blue (B) wavelength of light. 4. The system of claim 1, wherein the plurality of kernels are distributed over a monolithic meta-optics. 5. The system of claim 1, wherein the PSFs of the plurality of kernels are at least in part obtained from a knowledge distillation that comprises compressing operations of multiple electronic convolutional layers into an operation of one electronic convolutional layer. convolving the source images by applying point spread functions (PSFs) of the convolutional kernels to the source images; capturing convolved images by an image capturing camera; and based on the convolved images, classifying the source images with an electronic backend software. 14. The method of claim 13, wherein convolving images comprises executing multi-wavelength kernel convolutions simultaneously. 15. The method of claim 14, wherein the multi-wavelength kernel convolutions comprise convolutions of a red (R) wavelength, a green (G) wavelength, and a blue (B) wavelength of light. 16. The method of claim 13, wherein the plurality of kernels are distributed over a monolithic meta-optics. 17. The method of claim 13, further comprising designing the PSFs of the plurality of kernels by compressing operations of multiple electronic convolutional layers into an operation of one electronic convolutional layer. 18. The method of claim 17, further comprising determining PSFs by simulation prior to manufacturing of the meta-optics. 18. The method of claim 17, further comprising compensating fabrication and misalignment errors of the meta-optics by a calibration layer. 19. The method of claim 18, wherein the calibration layer is configured for assigning weights to the convolved images, wherein the weights are determined by 31 comparing the convolved images obtained by the meta-optics with the convolved images obtained by an electronic convolution. 20. The method of claim 13, wherein the electronic backend software is a trained and reconfigurable artificial intelligence (AI). 21. The method of claim 13, wherein the convolved images are equivalent to an output of 2-dimensional convolutional layers. 32 comparing the convolved images obtained by the meta-optics with the convolved images obtained by an electronic convolution.
20. The method of claim 13, wherein the electronic backend software is a trained and reconfigurable artificial intelligence (Al).
21. The method of claim 13, wherein the convolved images are equivalent to an output of 2-dimensional convolutional layers.
PCT/US2025/020565 2024-03-22 2025-03-19 Compressed meta-optical encoder for image classification Pending WO2025199235A1 (en)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
US202463568722P 2024-03-22 2024-03-22
US63/568,722 2024-03-22
US202463721961P 2024-11-18 2024-11-18
US63/721,961 2024-11-18

Publications (1)

Publication Number Publication Date
WO2025199235A1 true WO2025199235A1 (en) 2025-09-25

Family

ID=97140197

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2025/020565 Pending WO2025199235A1 (en) 2024-03-22 2025-03-19 Compressed meta-optical encoder for image classification

Country Status (1)

Country Link
WO (1) WO2025199235A1 (en)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114361151A (en) * 2021-12-31 2022-04-15 深圳市穗晶光电股份有限公司 Novel optical sensing chip packaging structure combined with super lens
US20230230195A1 (en) * 2018-05-14 2023-07-20 Tempus Labs, Inc. Determining Biomarkers from Histopathology Slide Images

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20230230195A1 (en) * 2018-05-14 2023-07-20 Tempus Labs, Inc. Determining Biomarkers from Histopathology Slide Images
CN114361151A (en) * 2021-12-31 2022-04-15 深圳市穗晶光电股份有限公司 Novel optical sensing chip packaging structure combined with super lens

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
MARTIN BARKEY; REBECCA BUCHNER; ALWIN WESTER; STEFANIE D. PRITZL; MAKSIM MAKARENKO; QIZHOU WANG; THOMAS WEBER; DIRK TRAUNER; STEFA: "Pixelated high-Q metasurfaces for in-situ biospectroscopy and AI-enabled classification of lipid membrane photoswitching dynamics", ARXIV.ORG, 29 August 2023 (2023-08-29), pages 1 - 22, XP091601135 *
QUAN LIU; HANYU ZHENG; BRANDON T. SWARTZ; HO HIN LEE; ZUHAYR ASAD; IVAN KRAVCHENKO; JASON G. VALENTINE; YUANKAI HUO: "Digital Modeling on Large Kernel Metamaterial Neural Network", ARXIV.ORG, 21 July 2023 (2023-07-21), pages 1 - 17, XP091571346 *
WANG XIHUI: "Thin-film Nitrate Sensor Performance Prediction Based on pre-processed sensor images", THESIS, 1 December 2023 (2023-12-01), pages 1 - 129, XP093360505 *

Similar Documents

Publication Publication Date Title
US12086717B2 (en) Devices and methods employing optical-based machine learning using diffractive deep neural networks
US12443838B2 (en) Diffractive deep neural networks with differential and class-specific detection
CN110441271B (en) High-resolution deconvolution method and system of light field based on convolutional neural network
US11640040B2 (en) Simultaneous focal length control and achromatic computational imaging with quartic metasurfaces
US12506945B2 (en) Neural nano-optics for high-quality thin lens imaging
EP3007432B1 (en) Image acquisition device and image acquisition method
CN107615022A (en) Imaging device with image dispersion for creating spatially encoded images
US20230296516A1 (en) Ai-driven signal enhancement of sequencing images
KR20230127277A (en) Imaging device and optical element
US20250078291A1 (en) Meta-Optic Accelerators for Machine Vision and Related Methods
US12505511B2 (en) AI-driven enhancement of motion blurred sequencing images
KR20230127278A (en) Imaging device and optical element
CN118401811A (en) Spectroscopic device and terminal equipment with spectroscopic device and working method
US20230401436A1 (en) Scale-, shift-, and rotation-invariant diffractive optical networks
CN116824356A (en) Method and system for extracting and classifying spatial elevation spectral features of multi-source remote sensing images
WO2024025901A1 (en) Optical sensing with nonlinear optical neural networks
US10823945B2 (en) Method for multi-color fluorescence imaging under single exposure, imaging method and imaging system
Blank et al. Diffraction lens in imaging spectrometer
Majumdar et al. Transferable polychromatic optical encoder for neural networks
EP4703785A1 (en) A light source for a fourier ptychography microscopy system
Pan et al. Efficient restoration of diffraction-induced space-variant blur in stacked microlens array scanning imaging system
Vašinka et al. Universal super-resolution framework for imaging of quantum dots
WO2024018081A1 (en) Method for processing digital images of a microscopic sample and microscope system
WO2024218294A1 (en) Construction of a digital image depicting a sample
CN117850726A (en) Method, system and device for convolution calculation using DOE photonic crystal film

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25774296

Country of ref document: EP

Kind code of ref document: A1