WO2025096551A1 - Photonic tensor core devices and systems - Google Patents

Photonic tensor core devices and systems Download PDF

Info

Publication number
WO2025096551A1
WO2025096551A1 PCT/US2024/053575 US2024053575W WO2025096551A1 WO 2025096551 A1 WO2025096551 A1 WO 2025096551A1 US 2024053575 W US2024053575 W US 2024053575W WO 2025096551 A1 WO2025096551 A1 WO 2025096551A1
Authority
WO
WIPO (PCT)
Prior art keywords
optical
coupled
signal
modulators
matrix
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/US2024/053575
Other languages
French (fr)
Inventor
Zhaoran Huang
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Rensselaer Polytechnic Institute
Original Assignee
Rensselaer Polytechnic Institute
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Rensselaer Polytechnic Institute filed Critical Rensselaer Polytechnic Institute
Publication of WO2025096551A1 publication Critical patent/WO2025096551A1/en
Anticipated expiration legal-status Critical
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/06Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons
    • G06N3/067Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons using optical means
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06EOPTICAL COMPUTING DEVICES
    • G06E1/00Devices for processing exclusively digital data
    • G06E1/02Devices for processing exclusively digital data operating upon the order or content of the data handled
    • G06E1/04Devices for processing exclusively digital data operating upon the order or content of the data handled for performing computations using exclusively denominational number representation, e.g. using binary, ternary, decimal representation
    • G06E1/045Matrix or vector computation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06EOPTICAL COMPUTING DEVICES
    • G06E3/00Devices not provided for in group G06E1/00, e.g. for processing analogue or hybrid data
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06EOPTICAL COMPUTING DEVICES
    • G06E3/00Devices not provided for in group G06E1/00, e.g. for processing analogue or hybrid data
    • G06E3/008Matrix or vector computation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/06Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons
    • G06N3/067Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons using optical means
    • G06N3/0675Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons using optical means using electro-optical, acousto-optical or opto-electronic means

Definitions

  • Photonic computing has emerged as a promising technology for high-performance and energy-efficient computing, particularly in computation-intensive artificial intelligence (Al) tasks.
  • Integrated photonic tensor core (PTC) designs have demonstrated ultra-fast photonic analog linear operation acceleration.
  • Coherent PTCs that leverage interference and diffraction include Mach- Zender interferometer (MZI) arrays, butterfly-style meshes, auto-designed photonic circuits, coupler-crossbar array, star-coupler-based design, and metalens-based diffractive PTCs.
  • MZI Mach- Zender interferometer
  • WDM wavelength-division multiplexing
  • MRR microring
  • PCM Phase-change material
  • the efficiency of PTCs for use with general edge Al can be considered from the perspectives of versatility, dynamic reprogrammability, and domain-specific customization.
  • Versatility, or universality is one of the important features of photonic Al hardware to accelerate a variety of deep neural network (DNN) workloads.
  • DNN deep neural network
  • a versatile/generic photonic accelerator based on universal optical linear units is capable of realizing general matrix multiplication (GEMM) and thus directly implementing a wide spectrum of pretrained digital DNNs.
  • GEMM general matrix multiplication
  • Many specialized linear units are not applicable to generic tensor computation since they restrict their matrix expressivity to a subspace of specialized matrices for higher hardware efficiency.
  • optical computing hardware demonstrations are based on standard foundry process design kit (PDK) elements, which are designed for optical communications and not optimized for analog neuromorphic computing.
  • PDK foundry process design kit
  • circuit level customization is important for reducing the long-lasting analog-to-digital and optical-to-electrical conversion bottlenecks.
  • the architecture topology and dataflow also need to be customized to fully leverage the temporal locality to reduce data movement cost and maximize hardware sharing.
  • the devices include sets of optical modulators for receiving optical signals and encoding matrix values onto the optical signals.
  • a set of dot product engines combines the encoded optical signals corresponding to the matrix values to provide product photocurrent signals, which are then converted to digital electric signals.
  • the optical modulators comprise Mach- Zehnder modulators that encode each optical signal via a temporal electrical signal provided to each modulator.
  • FIG. 1 shows a schematic view of an embodiment of a photonic tensor core device according to the present technology
  • FIG. 2 shows a schematic view of an alternative embodiment of a photonic tensor core device according to the present technology
  • FIG. 4 shows a schematic view of an embodiment of a temporal integrator according to the present technology
  • FIG. 5 shows a schematic view of an embodiment of a computing system according to the present technology.
  • FIG. 6A shows a schematic top-view of a modulator according to an embodiment of the present technology
  • FIG. 6B shows a schematic cross-section view of the modulator of FIG. 6 A.
  • a time-multiplexed dynamic photonic tensor accelerator design for Al acceleration is provided.
  • Some embodiments include ultracompact, slow-light electro-optic modulators for input operand encoding.
  • Some embodiments include hierarchical partial product accumulation with lightweight capacitive temporal integration modules.
  • Some embodiments include multi-core architecture to maximize sharing of data input/readout circuitry.
  • slow-light MZI modulators (SL-MZM) with enhanced light-matter interaction for size and power reduction are used.
  • the SL-MZM phase shifter length is from about 150 pm to about 200 pm with a footprint about 10 greater than Si MRR while an order of magnitude smaller than a typical foundry offered Si Mach- Zehnder modulator (MZM) PDK element.
  • MZM Mach- Zehnder modulator
  • this SL-MZM is thermally robust, with no thermal tuning/locking circuit needed, and can also tolerate large manufacturing variations.
  • some embodiments of the present technology simplify the spectral multi -wavelength encoding by employing high-speed temporal encoding, eliminating the need for complex dispersion-engineered broadband device designs such as Si modulators, optical power splitters and directional couplers as well as remove WDM MUX/DEMUX overhead.
  • the present technology relates to matrix multiplication, which is an important linear operation for various information processing workloads.
  • Some embodiments of the present technology perform matrix-matrix multiplication.
  • Each vector dot-product operation can be mapped to a dynamic dot-product engine, explained below in reference to FIG. 3.
  • Multiple dot-product engines can form an array structure, i.e., a tensor core, to realize parallel matrix-matrix multiplication.
  • the first and second sets of optical modulators 101, 102 comprise Mach-Zehnder modulators.
  • the modulators comprise slow-light Mach- Zehnder modulators.
  • Silicon-based (Si) modulators are used that utilize the carrier plasma effect for on-chip PTC. In the embodiment shown, which achieves a dot-product operation for matrices of the size of K x K requires 2K modulators for signal conversion.
  • the physical dimension of the Si modulators is a design parameter that impacts the scalability of matrix operation.
  • digital electrical signals carrying the matrix information are encoded onto the optical signals represented by the amplitude and phase by the optical modulators.
  • E in be the electric field of the optical signal to the optical modulator
  • the electric field of the modulator output can be expressed as Ein cos ri, allowing broadband mapping of both positive and negative values.
  • a ID dielectric photonic crystal waveguide is used as the slow-light-enabled compact modulator.
  • the footprint of the modulator array is significantly reduced.
  • an Si slow-light MZM (SL-MZM) with a phase shifter length (LPS) of 150 pm is used.
  • the phase shifter length is in the range from about 150 pm to about 200 pm.
  • the SL-MZM is fabricated under a multi-project wafer (MPW) run, for complete foundry compatibility.
  • MPW multi-project wafer
  • FIG. 6a shows a schematic top-view diagram of an embodiment of a modulator 626.
  • the modulator 626 comprises a Y-splitter 632 with even power splitting between a grating arm 627 and a reference arm 628 followed by a Y-combiner 633 to recombine the split signals.
  • the arms include heaters 634 to help compensate for asymmetries.
  • FIG. 6B shows a schematic cross section view of the modulator 626.
  • Both the electrical bandwidth and the linearity of a Si modulator help determine the high- bit resolution at a high computing clock frequency.
  • the speed of a Si SL-MZM is limited by its RC time constant and photon lifetime.
  • the phase shifter length is reduced in an SL-MZM, the total capacitance decreases.
  • the measured SL-MZM junction capacitance was approximately -0.75 pF.
  • the embodiment of a device 100 in FIG. 1 further comprises a first set 104 of K 1 x K optical power splitters for splitting the encoded signal corresponding to the matrix values ... x KN each coupled to one of the first set 101 of K optical modulators and coupled to K dot product engines 103.
  • This embodiment of the device 100 further comprises a second set 105 of K 1 x K optical power splitters for splitting the encoded signal corresponding to the matrix values yn ⁇ VNK eac h coupled to one of the second set 102 of K optical modulators and coupled to K dot product engines 103.
  • the design in this embodiment includes (K - I) 2 waveguide crossings 106.
  • the 1 further comprises a multimode interferometer 107, which is coupled to each of the optical modulators 101 and 102 for providing the optical signal.
  • the multimode interferometer 107 acts as a 1 x 2K splitter to divide the incoming light signal into 2K optical signals to be processed by the modulators 101, 102. Accordingly, the multimode interferometer 107 is configured to be coupled to a light source 109.
  • Some embodiments of PTCs according to the present technology that utilize optical wave phase and amplitude in time-domain processing use a monochromatic light source for optical signal processing.
  • Some embodiments utilize o-band operation instead of c-band components.
  • O- band operation offers several advantages such as a smaller optical mode volume in Si/SiCh waveguide structure, higher mode confinement with tighter bending radius and >1.5 higher in Ge photodetector responsivity.
  • a laser module that is disposed “off-chip” is used. That is, the light source is separate from the device 100.
  • an on-chip laser diode is used as the light source.
  • an on-chip III-V integrated laser diode is used.
  • a high-power, monolithic o-band laser is used, having output power of about 150 mW. In some embodiments, a moderate laser power of 100 mW is used.
  • FIG. 2 shows an alternative embodiment of a photonic tensor core device 200. Similar to the device 100, device 200 includes a first set 201 of K optical modulators, a second set 202 of K of optical modulators, and a set 203 of K x K dot product engines. In this embodiment, however, the coupling arrangement between the optical modulators and the dot product engines is different. Device 200 comprises a plurality of first sets 204 of K - 1 optical power splitters, each set 204’, 204”, 204’” is coupled in series to one of the first set 201 of K optical modulators and each coupled to a dot product engine.
  • Device 200 also comprises a plurality of second sets 205 of K- 1 optical power splitters, each set 205’, 205”, 205’” coupled in series to one of the second set 202 of K optical modulators and each coupled to a dot product engine.
  • the splitting ratios of the first and second sets 204 and 205 of optical power splitters are set at 1.(K - 1), 1.(K - 2), . . . , 1 : 1 along the series of splitters. This is a series of uneven splitters intended to reduce the number of waveguide crossings 206 as compared to the design of device 100.
  • FIG. 3 shows a schematic diagram of an optical dot-product engine 303 according to some embodiments of the present technology.
  • the engine 303 comprises a 2 x 2 optical power splitter 309, comprising a first arm 310 and a second arm 311.
  • the first arm 310 is coupled to one of the first set (e.g., 101, 201, not shown in FIG. 3) of K modulators and the second arm 311 is coupled to one of the second set (e.g., 102, 202, not shown in FIG. 3) of K modulators.
  • the engine 303 further comprises a 7t/2 phase shifter 312 coupled to the second arm 311 of the optical power splitter; a first photodiode 313 coupled to the first arm 310 of the optical power splitter; and a second photodiode 314 coupled to the second arm 311 of the optical power splitter.
  • the first and second photodiodes are configured to convert each encoded optical signal 315, 316 to a photocurrent signal which combine to form the product photocurrent signal 317.
  • a 2 * 2 50:50 optical power splitter is used to generate interference between the optical signals from two input arms.
  • the embodiment shown in FIG. 3 utilizes a directional coupler structure to generate the 50:50 power splitting, but, in other embodiments, multimode interferometers (MMI) are used.
  • MMI multimode interferometers
  • a temporal integrator 318 is coupled to each dot product engine 303 for converting the product photocurrent signal 317 to an integrated signal.
  • a schematic diagram of a temporal integrator 418 is shown in FIG. 4.
  • the integrator 418 comprises at least two capacitors 419 and 420 connected in parallel to the dot product engine for receiving the product photocurrent signal.
  • the capacitors 419 and 420 are arranged as a flipped capacitor pair.
  • multiple flipped capacitor pairs are connected in parallel to achieve a symmetric circuit topology.
  • input 421 receives the product photocurrent signal.
  • the integrator 418 further comprises one or more field effect transistors connected in parallel to the capacitors.
  • two n-channel FETs 422 and two p-channel FETs 423 are shown, connected in parallel with the capacitors.
  • ten 40 nm n-channel and ten 40 nm p-channel FETs are connected in parallel with the capacitors to help with current driving capability for reset within a single baud time period.
  • integrators with an op-amp and a capacitive feedback loop are used.
  • the devices 100 and/or 200 further comprise at least one analog-to- digital converter associated with one or more of the temporal integrators for converting the integrated signal to a digital signal.
  • FIG. 5 shows a schematic diagram of a computing system 524 according to another embodiment of the present technology.
  • the system 524 comprises a monochromatic light source 508.
  • the source 508 is a laser.
  • the system also includes one or more photonic tensor core devices 500 for multiplying & K x N portion of a first matrix A, comprising M rows and N columns, by a N x K portion of a second matrix F, comprising N rows and Q columns.
  • Each device 500 includes an optical power splitter 507 that receives light from the monochromatic light source 508 and splits the light into 2K optical signals.
  • Each device also comprises a first set 501 of K optical modulators each configured to receive an optical signal and to encode a matrix value ... x KN from matrix X onto the optical signal and a second set 502 of K optical modulators each configured to receive an optical signal and to encode a matrix value yn ⁇ YNK f rom matrix Y onto the optical signal.
  • the devices 500 include a set of 2K dot product engines, where each dot product engine is coupled to one of the first set of modulators and to one of the second set of modulators, and each dot product engine is configured to combine the optical signals corresponding to each of the matrix values % X1 ... x KN and y lx ... y NK and provide a product photocurrent signal corresponding to elements z xl ... z KK of a result matrix Z.
  • the system also includes an encoding circuitry 525 that is configured to generate an electrical signal corresponding to each matrix value to send to each of the first and second optical modulators for encoding each optical signal with each matrix value.
  • digital electrical signals carrying the matrix information are converted to optical signals represented by the amplitude and phase by the optical modulators, which comprise, e.g., slow-light Mach-Zehnder modulators.
  • the optical modulators comprise, e.g., slow-light Mach-Zehnder modulators.
  • letE/n be the electric field of the optical signal to the optical modulator
  • the electric field of the modulator output can be expressed as Ein cos 0, allowing broadband mapping of both positive and negative values.
  • a set of integrators 518 integrate the photocurrent signals and a trans-impedance amplifier (TIA) and analog-to-digital converter (ADC) converts them to electronic digital signals.
  • TIA trans-impedance amplifier
  • ADC analog-to-digital converter
  • each PTC is a crossbar of K x K dot-product engines that can finish a T x 1 times 1 x K vector outer product at each time step, T.
  • the X matrix of size M x N is partitioned into M/K horizontal strips, each with a size K -' N
  • the matrix Y of size N x Q is partitioned into Q/K vertical strips, each with a size of N x K.
  • Each cycle therefore, includes (1) feeding one vector into the PTC, (2) reading out the outer product results as photocurrent, (3) converting it to the electronic domain, and (4) accumulating partial product.
  • first element functioning as a male
  • second element functioning as a female
  • first element functioning as a female
  • second element functioning as a male
  • first and second elements are configured to mate with, fit with or otherwise interlock with each other.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Computing Systems (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Biophysics (AREA)
  • Mathematical Physics (AREA)
  • Biomedical Technology (AREA)
  • Health & Medical Sciences (AREA)
  • Computational Linguistics (AREA)
  • Neurology (AREA)
  • Artificial Intelligence (AREA)
  • Data Mining & Analysis (AREA)
  • Evolutionary Computation (AREA)
  • General Health & Medical Sciences (AREA)
  • Molecular Biology (AREA)
  • General Engineering & Computer Science (AREA)
  • Software Systems (AREA)
  • Pure & Applied Mathematics (AREA)
  • Mathematical Analysis (AREA)
  • Algebra (AREA)
  • Optical Modulation, Optical Deflection, Nonlinear Optics, Optical Demodulation, Optical Logic Elements (AREA)

Abstract

Photonic tensor cores that utilize optical modulators to perform matrix multiplication. Optical signals are encoded with matrix values via electrical signals applied by the optical modulators. Dot product engines combine the encoded optical signals and provide product photocurrent signals for conversion to digital electrical signals.

Description

PHOTONIC TENSOR CORE DEVICES AND SYSTEMS
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63/546,305 filed October 30, 2023 and U.S. Provisional Patent Application No. 63/713,086 filed October 29, 2024, both of which are incorporated by reference as if disclosed herein in their entireties.
BACKGROUND
[0002] Photonic computing has emerged as a promising technology for high-performance and energy-efficient computing, particularly in computation-intensive artificial intelligence (Al) tasks. Integrated photonic tensor core (PTC) designs have demonstrated ultra-fast photonic analog linear operation acceleration. Coherent PTCs that leverage interference and diffraction include Mach- Zender interferometer (MZI) arrays, butterfly-style meshes, auto-designed photonic circuits, coupler-crossbar array, star-coupler-based design, and metalens-based diffractive PTCs. In addition, to leverage the wavelength-division multiplexing (WDM) technique, there are incoherent multi -wavelength PTCs, including: microring (MRR) weight banks, Phase-change material (PCM) crossbar arrays, and micro-comb-based computing engines.
[0003] The efficiency of PTCs for use with general edge Al can be considered from the perspectives of versatility, dynamic reprogrammability, and domain-specific customization. Versatility, or universality, is one of the important features of photonic Al hardware to accelerate a variety of deep neural network (DNN) workloads. A versatile/generic photonic accelerator based on universal optical linear units is capable of realizing general matrix multiplication (GEMM) and thus directly implementing a wide spectrum of pretrained digital DNNs. Many specialized linear units are not applicable to generic tensor computation since they restrict their matrix expressivity to a subspace of specialized matrices for higher hardware efficiency.
[0004] Real-time, efficient input tensor encoding with low reconfiguration costs (z.e., reprogrammability) is also advantageous for photonic computing. High weight encoding costs due to the high complexity of matrix decomposition required to encode weights can restrict photonic computing designs to only support weight-static linear operations, e.g., fully connected (FC) layers and convolutional (CONV) layers, where weights are pretrained and pre-encoded into the device/circuit transmissions. However, advanced Al models, based on attention operations where both matrix multiplication operands are dynamic, full-range, and general tensors, cannot be efficiently mapped to weight-static PTCs. [0005] Domain-specific hardware customization is a third advantageous feature for efficient, scalable PTCs. At the device level, many optical computing hardware demonstrations are based on standard foundry process design kit (PDK) elements, which are designed for optical communications and not optimized for analog neuromorphic computing. At the circuit level, customization is important for reducing the long-lasting analog-to-digital and optical-to-electrical conversion bottlenecks. At the architecture level, due to the lack of optical memory, the large spatial footprint of photonic circuits, and the high digital memory access cost, the architecture topology and dataflow also need to be customized to fully leverage the temporal locality to reduce data movement cost and maximize hardware sharing.
[0006] What are needed, therefore, are versatile, reprogrammable, and customizable photonic computing designs to help process advanced Al tasks.
SUMMARY
[0007] Aspects of the present technology are directed to photonic tensor core devices for performing matrix multiplication. In some embodiments, the devices include sets of optical modulators for receiving optical signals and encoding matrix values onto the optical signals. In some embodiments, a set of dot product engines combines the encoded optical signals corresponding to the matrix values to provide product photocurrent signals, which are then converted to digital electric signals. In some embodiments, the optical modulators comprise Mach- Zehnder modulators that encode each optical signal via a temporal electrical signal provided to each modulator.
[0008] Other aspects of the present technology are directed to computing systems for performing matrix multiplication that comprise a light source, one or more photonic tensor core devices, and encoding circuitry configured to generate electrical signals corresponding to each matrix value to send to optical modulators for encoding onto optical signals.
[0009] Additional aspects and features of the present technology will be apparent to those of skill in the art upon consideration of the following description and the attached drawings.
BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings show embodiments of the disclosed subject matter for the purpose of illustrating the technology. However, it should be understood that the present application is not limited to the precise arrangements and instrumentalities shown in the drawings, wherein:
[0011] FIG. 1 shows a schematic view of an embodiment of a photonic tensor core device according to the present technology; [0012] FIG. 2 shows a schematic view of an alternative embodiment of a photonic tensor core device according to the present technology;
[0013] FIG. 3 shows a schematic view of an embodiment of an optical dot-product engine according to the present technology;
[0014] FIG. 4 shows a schematic view of an embodiment of a temporal integrator according to the present technology;
[0015] FIG. 5 shows a schematic view of an embodiment of a computing system according to the present technology.
[0016] FIG. 6A shows a schematic top-view of a modulator according to an embodiment of the present technology;
[0017] FIG. 6B shows a schematic cross-section view of the modulator of FIG. 6 A.
DETAILED DESCRIPTION
[0018] The following discussion relates to various embodiments of photonic sensor core systems and devices. It will be understood that the herein described versions are examples that embody certain inventive concepts as detailed herein. To that end, other variations and modifications will be readily apparent to those of sufficient skill.
[0019] In some embodiments of the present technology, a time-multiplexed dynamic photonic tensor accelerator design for Al acceleration is provided. Some embodiments include ultracompact, slow-light electro-optic modulators for input operand encoding. Some embodiments include hierarchical partial product accumulation with lightweight capacitive temporal integration modules. Some embodiments include multi-core architecture to maximize sharing of data input/readout circuitry. In some embodiments, slow-light MZI modulators (SL-MZM) with enhanced light-matter interaction for size and power reduction are used. In some embodiments, the SL-MZM phase shifter length is from about 150 pm to about 200 pm with a footprint about 10 greater than Si MRR while an order of magnitude smaller than a typical foundry offered Si Mach- Zehnder modulator (MZM) PDK element. In some embodiments, this SL-MZM is thermally robust, with no thermal tuning/locking circuit needed, and can also tolerate large manufacturing variations. In contrast to a multi wavelength dynamic PTC design, some embodiments of the present technology simplify the spectral multi -wavelength encoding by employing high-speed temporal encoding, eliminating the need for complex dispersion-engineered broadband device designs such as Si modulators, optical power splitters and directional couplers as well as remove WDM MUX/DEMUX overhead. [0020] The present technology relates to matrix multiplication, which is an important linear operation for various information processing workloads. Some embodiments of the present technology perform matrix-matrix multiplication. Consider two input matrices, matrix X with M x N dimension and matrix Y with N x Q dimension:
Figure imgf000006_0001
[0021] The matrix multiplication Z = X ■ Y is:
Figure imgf000006_0002
[0022] The resulting Z is an M x Q matrix; and its c/th row, Z>th column element zab is obtained by calculating the dot-product of the c/th row vector of X and Z>th column vector of Y, i.e.,
Figure imgf000006_0003
[0023] Each vector dot-product operation can be mapped to a dynamic dot-product engine, explained below in reference to FIG. 3. Multiple dot-product engines can form an array structure, i.e., a tensor core, to realize parallel matrix-matrix multiplication.
[0024] Referring now to FIG. 1, a first embodiment of a photonic tensor core device 100 for multiplying a K N portion of a first matrix X, comprising M rows and N columns, by a N x K portion of a second matrix Y, comprising N rows and Q columns, is shown schematically. In this embodiment, the device comprises a first set 101 of K optical modulators each configured to receive an optical signal and to encode a matrix value ... xKN from matrix X onto the optical signal. The device further comprises a second set 102 of K optical modulators each configured to receive an optical signal and to encode a matrix value yxl ... yNK from matrix T onto the optical signal. The optical signals are provided by a light source, such as a laser, as described below.
[0025] A set 103 of K x K dot product engines, each coupled to one of the first set of modulators 101 and to one of the second set of modulators 102 is also included in this embodiment. Each dot product engine 103 is configured to combine the encoded optical signals corresponding to each of the matrix values x ... xKN and y ... yNK and to provide product photocurrent signals corresponding to elements zxl ... zKK of a result matrix Z. The result matrix is the result of the multiplication of the portions of the X and T matrices. In this embodiment, each matrix value is encoded to each optical signal via a temporal electrical signal corresponding to each matrix value provided to each modulator.
[0026] In some embodiments, the first and second sets of optical modulators 101, 102 comprise Mach-Zehnder modulators. In some embodiments, the modulators comprise slow-light Mach- Zehnder modulators. In some embodiments, Silicon-based (Si) modulators are used that utilize the carrier plasma effect for on-chip PTC. In the embodiment shown, which achieves a dot-product operation for matrices of the size of K x K requires 2K modulators for signal conversion. In some embodiments, the physical dimension of the Si modulators is a design parameter that impacts the scalability of matrix operation.
[0027] In some embodiments, digital electrical signals carrying the matrix information are encoded onto the optical signals represented by the amplitude and phase by the optical modulators. For example, let Ein be the electric field of the optical signal to the optical modulator, then the electric field of the modulator output can be expressed as Ein cos ri, allowing broadband mapping of both positive and negative values.
[0028] In some embodiments, a ID dielectric photonic crystal waveguide, specifically a rectangular-shaped Bragg grating, is used as the slow-light-enabled compact modulator. In some such embodiments, the footprint of the modulator array is significantly reduced. In some embodiments, an Si slow-light MZM (SL-MZM) with a phase shifter length (LPS) of 150 pm is used. In some embodiments, the phase shifter length is in the range from about 150 pm to about 200 pm. In some embodiments, the SL-MZM is fabricated under a multi-project wafer (MPW) run, for complete foundry compatibility.
[0029] FIG. 6a shows a schematic top-view diagram of an embodiment of a modulator 626. The modulator 626 comprises a Y-splitter 632 with even power splitting between a grating arm 627 and a reference arm 628 followed by a Y-combiner 633 to recombine the split signals. In this embodiment, the arms include heaters 634 to help compensate for asymmetries. FIG. 6B shows a schematic cross section view of the modulator 626. In this embodiment, the gratings 629 have a width (Wg) of 1.5 pm and a period of 282 nm, and Hf = 220 nm, Wc = 400nm, and Hs = 110 nm. To encode the input optical signal with a matrix value, voltages are applied through the vias 631. [0030] Reflection occurring at different junctions within the modulator device, optical absorption due to carriers in waveguides, propagation loss in the Bragg grating phase shifter due to increased group indices, and mode mismatch at the Bragg grating waveguide interfaces are the primary factors contributing to the modulator insertion loss. In one embodiment, the measured total modulator insertion loss is ~6.4 dB for Lps = 150 pm.
[0031] Both the electrical bandwidth and the linearity of a Si modulator help determine the high- bit resolution at a high computing clock frequency. Operating under reverse bias, the speed of a Si SL-MZM is limited by its RC time constant and photon lifetime. In some embodiments, the PN junctions are doped at an elevated level (ranging from 1018 =cm3 to 1019 =cm3) to enhance the carrier plasma effect. As the phase shifter length is reduced in an SL-MZM, the total capacitance decreases. In some embodiments, the measured SL-MZM junction capacitance was approximately -0.75 pF. Depending on the doping level in the connecting Si bar from the ridge waveguide to the via contacts, the intrinsic resistance of a SL-MZM according to some embodiments ranges from 5 to 10 Q. The estimated RC time-limited electrical bandwidth of a SL- MZM is thus in the hundreds of GHz in some embodiments. The slow-light effect can be viewed as a traveling wave resonant in its propagation direction, with the optical bandwidth determined by the Q-factor of the resonator. For the rectangular Bragg grating-shaped slow-light, an optical bandwidth of approximately -26 GHz is estimated for some embodiments. In some embodiments, dispersion engineering techniques such as phase-shifted Bragg grating, dispersion compensation, and line-shift photonic crystal waveguide are used to reduce the dispersion-induced bandwidth penalty.
[0032] The embodiment of a device 100 in FIG. 1 further comprises a first set 104 of K 1 x K optical power splitters for splitting the encoded signal corresponding to the matrix values ... xKN each coupled to one of the first set 101 of K optical modulators and coupled to K dot product engines 103. This embodiment of the device 100 further comprises a second set 105 of K 1 x K optical power splitters for splitting the encoded signal corresponding to the matrix values yn ■■■ VNK each coupled to one of the second set 102 of K optical modulators and coupled to K dot product engines 103. As shown in FIG. 1, the design in this embodiment includes (K - I)2 waveguide crossings 106. [0033] The embodiment of FIG. 1 further comprises a multimode interferometer 107, which is coupled to each of the optical modulators 101 and 102 for providing the optical signal. The multimode interferometer 107 acts as a 1 x 2K splitter to divide the incoming light signal into 2K optical signals to be processed by the modulators 101, 102. Accordingly, the multimode interferometer 107 is configured to be coupled to a light source 109.
[0034] Some embodiments of PTCs according to the present technology that utilize optical wave phase and amplitude in time-domain processing use a monochromatic light source for optical signal processing. Some embodiments utilize o-band operation instead of c-band components. O- band operation offers several advantages such as a smaller optical mode volume in Si/SiCh waveguide structure, higher mode confinement with tighter bending radius and >1.5 higher in Ge photodetector responsivity.
[0035] In some embodiments, such as the embodiment in FIG. 1, a laser module that is disposed “off-chip” is used. That is, the light source is separate from the device 100. In other embodiments, an on-chip laser diode is used as the light source. In some embodiments, an on-chip III-V integrated laser diode is used. In some embodiments, a high-power, monolithic o-band laser is used, having output power of about 150 mW. In some embodiments, a moderate laser power of 100 mW is used.
[0036] FIG. 2 shows an alternative embodiment of a photonic tensor core device 200. Similar to the device 100, device 200 includes a first set 201 of K optical modulators, a second set 202 of K of optical modulators, and a set 203 of K x K dot product engines. In this embodiment, however, the coupling arrangement between the optical modulators and the dot product engines is different. Device 200 comprises a plurality of first sets 204 of K - 1 optical power splitters, each set 204’, 204”, 204’” is coupled in series to one of the first set 201 of K optical modulators and each coupled to a dot product engine. Device 200 also comprises a plurality of second sets 205 of K- 1 optical power splitters, each set 205’, 205”, 205’” coupled in series to one of the second set 202 of K optical modulators and each coupled to a dot product engine. In this embodiment, the splitting ratios of the first and second sets 204 and 205 of optical power splitters are set at 1.(K - 1), 1.(K - 2), . . . , 1 : 1 along the series of splitters. This is a series of uneven splitters intended to reduce the number of waveguide crossings 206 as compared to the design of device 100.
[0037] FIG. 3 shows a schematic diagram of an optical dot-product engine 303 according to some embodiments of the present technology. The engine 303 comprises a 2 x 2 optical power splitter 309, comprising a first arm 310 and a second arm 311. In this embodiment, the first arm 310 is coupled to one of the first set (e.g., 101, 201, not shown in FIG. 3) of K modulators and the second arm 311 is coupled to one of the second set (e.g., 102, 202, not shown in FIG. 3) of K modulators. The engine 303 further comprises a 7t/2 phase shifter 312 coupled to the second arm 311 of the optical power splitter; a first photodiode 313 coupled to the first arm 310 of the optical power splitter; and a second photodiode 314 coupled to the second arm 311 of the optical power splitter. In this embodiment, the first and second photodiodes are configured to convert each encoded optical signal 315, 316 to a photocurrent signal which combine to form the product photocurrent signal 317.
[0038] Thus, in this embodiment, a 2 * 2 50:50 optical power splitter is used to generate interference between the optical signals from two input arms. The embodiment shown in FIG. 3 utilizes a directional coupler structure to generate the 50:50 power splitting, but, in other embodiments, multimode interferometers (MMI) are used.
[0039] In some embodiments, a temporal integrator 318 is coupled to each dot product engine 303 for converting the product photocurrent signal 317 to an integrated signal. A schematic diagram of a temporal integrator 418 is shown in FIG. 4. In some embodiments, the integrator 418 comprises at least two capacitors 419 and 420 connected in parallel to the dot product engine for receiving the product photocurrent signal. In this embodiment, the capacitors 419 and 420 are arranged as a flipped capacitor pair. In some embodiments, multiple flipped capacitor pairs are connected in parallel to achieve a symmetric circuit topology. In this embodiment, input 421 receives the product photocurrent signal. In some embodiments, the integrator 418 further comprises one or more field effect transistors connected in parallel to the capacitors. As shown, two n-channel FETs 422 and two p-channel FETs 423 are shown, connected in parallel with the capacitors. In some embodiments, ten 40 nm n-channel and ten 40 nm p-channel FETs are connected in parallel with the capacitors to help with current driving capability for reset within a single baud time period. In other embodiments, integrators with an op-amp and a capacitive feedback loop are used.
[0040] In some embodiments, the devices 100 and/or 200 further comprise at least one analog-to- digital converter associated with one or more of the temporal integrators for converting the integrated signal to a digital signal.
[0041] FIG. 5 shows a schematic diagram of a computing system 524 according to another embodiment of the present technology. The system 524 comprises a monochromatic light source 508. In some embodiments, the source 508 is a laser. The system also includes one or more photonic tensor core devices 500 for multiplying & K x N portion of a first matrix A, comprising M rows and N columns, by a N x K portion of a second matrix F, comprising N rows and Q columns. Each device 500 includes an optical power splitter 507 that receives light from the monochromatic light source 508 and splits the light into 2K optical signals. Each device also comprises a first set 501 of K optical modulators each configured to receive an optical signal and to encode a matrix value ... xKN from matrix X onto the optical signal and a second set 502 of K optical modulators each configured to receive an optical signal and to encode a matrix value yn ■■■ YNK from matrix Y onto the optical signal.
[0042] Further, similar to the devices described above, the devices 500 include a set of 2K dot product engines, where each dot product engine is coupled to one of the first set of modulators and to one of the second set of modulators, and each dot product engine is configured to combine the optical signals corresponding to each of the matrix values %X1 ... xKN and ylx ... yNK and provide a product photocurrent signal corresponding to elements zxl ... zKK of a result matrix Z.
[0043] The system also includes an encoding circuitry 525 that is configured to generate an electrical signal corresponding to each matrix value to send to each of the first and second optical modulators for encoding each optical signal with each matrix value. As described above, digital electrical signals carrying the matrix information are converted to optical signals represented by the amplitude and phase by the optical modulators, which comprise, e.g., slow-light Mach-Zehnder modulators. For example, letE/n be the electric field of the optical signal to the optical modulator, then the electric field of the modulator output can be expressed as Ein cos 0, allowing broadband mapping of both positive and negative values.
[0044] In this embodiment, for each tile of PTC, a set of integrators 518 integrate the photocurrent signals and a trans-impedance amplifier (TIA) and analog-to-digital converter (ADC) converts them to electronic digital signals.
[0045] In the system 524 shown in FIG. 5, a total of R tiles of C photonic tensor core devices are included, and each PTC is a crossbar of K x K dot-product engines that can finish a T x 1 times 1 x K vector outer product at each time step, T. In this embodiment, then, the X matrix of size M x N is partitioned into M/K horizontal strips, each with a size K -' N, and the matrix Y of size N x Q is partitioned into Q/K vertical strips, each with a size of N x K.
[0046] Each cycle, therefore, includes (1) feeding one vector into the PTC, (2) reading out the outer product results as photocurrent, (3) converting it to the electronic domain, and (4) accumulating partial product.
[0047] The parts, components, and structural elements described herein can be combined into an integral or unitary, one-piece objects, or such parts, components, and structural elements can be distinct, removable items that are attachable to each other.
[0048] In the foregoing description, certain components or elements may have been described as being configured to mate with each other. For example, an embodiment may be described as a first element (functioning as a male) configured to be inserted into a second element (functioning as a female). It should be appreciated that an alternate embodiment includes the first element (functioning as a female) configured to receive the second element (functioning as a male). In either such embodiment, the first and second elements are configured to mate with, fit with or otherwise interlock with each other.
[0049] Although the technology has been described and illustrated with respect to embodiments thereof, it should be understood by those skilled in the art that the foregoing and various other changes, omissions and additions may be made therein and thereto, without parting from the spirit and scope of the present technology.

Claims

CLAIMS What is claimed is:
1. A photonic tensor core device for multiplying K N portion of a first matrix A, comprising M rows and N columns, by a N x K portion of a second matrix K, comprising N rows and Q columns, the device comprising: a first set of K optical modulators each configured to receive an optical signal and to encode a matrix value x ... xKN from matrix A onto the optical signal; a second set of K optical modulators each configured to receive an optical signal and to encode a matrix value yxl ... yNK from matrix K onto the optical signal; and a set of K x K dot product engines, wherein each dot product engine is coupled to one of the first set of modulators and to one of the second set of modulators, and each dot product engine is configured to combine the encoded optical signals corresponding to each of the matrix values x ... xKN and yxl ... yNK and to provide product photocurrent signals corresponding to elements zxl ... zKK of a result matrix Z; wherein each matrix value is encoded to each optical signal via a temporal electrical signal corresponding to each matrix value provided to each modulator.
2. The device of claim 1, wherein each optical modulator comprises a slow-light Mach- Zehnder modulator.
3. The device of claim 2, wherein each optical modulator has a phase shifter length in the range from about 150 pm to about 200 pm.
4. The device of claim 2, further comprising a ID dielectric photonic crystal waveguide.
5. The device of claim 1, further comprising: a first set of A 1 x K optical power splitters for splitting the encoded signal corresponding to the matrix values x ... xKN each coupled to one of the first set of K optical modulators and coupled to K dot product engines; and a second set of K 1 x K optical power splitters for splitting the encoded signal corresponding to the matrix values ylx ... yNK each coupled to one of the second set of K optical modulators and coupled to K dot product engines.
6. The device of claim 1, further comprising: a plurality of first sets of K - 1 optical power splitters, each set coupled in series to one of the first set of K optical modulators and each coupled to a dot product engine; and a plurality of second sets of K- 1 optical power splitters, each set coupled in series to one of the second set of K optical modulators and each coupled to a dot product engine; wherein the splitting ratios of the first and second sets of optical power splitters are set at 1.(K - 1), 1 K - 2), . . . , 1 : 1 along the series of splitters.
7. The device of claim 1, wherein each of the dot product engines comprises: a 2 x 2 optical power splitter, comprising a first arm and a second arm, wherein the first arm is coupled to one of the first set of K modulators and the second arm is coupled to one of the second set of K modulators; a 7t/2 phase shifter coupled to the second arm of the optical power splitter; a first photodiode coupled to the first arm of the optical power splitter; and a second photodiode coupled to the second arm of the optical power splitter; wherein the first and second photodiodes are configured to convert each encoded optical signal to a photocurrent signal which combine to form a product photocurrent signal.
8. The device of claim 7, further comprising a temporal integrator coupled to each dot product engine for converting the product photocurrent signal to an integrated signal, comprising: at least two capacitors connected in parallel to the dot product engine for receiving the product photocurrent signal; one or more field effect transistors connected in parallel to the capacitors.
9. The device of claim 1, further comprising a multimode interferometer coupled to each of the optical modulators for providing the optical signal and wherein the multimode interferometer is configured to be coupled to a light source.
10. The device of claim 8, further comprising an anal og-to-digi tai converter associated with one or more of the temporal integrators for converting the integrated signal to a digital signal.
11. A computing system, comprising: a monochromatic light source; one or more photonic tensor core devices for multiplying aK N portion of a first matrix X, comprising AT rows and N columns, by a N x K portion of a second matrix T, comprising N rows and Q columns, comprising: an optical power splitter that receives light from the monochromatic light source and splits the light into 2K optical signals; a first set of K optical modulators each configured to receive an optical signal and to encode a matrix value
Figure imgf000015_0001
... xKN from matrix X onto the optical signal; a second set of K optical modulators each configured to receive an optical signal and to encode a matrix value yxl ... yNK from matrix T onto the optical signal; and a set of 2K dot product engines, wherein each dot product engine is coupled to one of the first set of modulators and to one of the second set of modulators, and each dot product engine is configured to combine the optical signals corresponding to each of the matrix values x ... xKN and yxl ... yNK and provide a product photocurrent signal corresponding to elements zxl ... zKK of a result matrix Z; an encoding circuitry configured to generate an electrical signal corresponding to each matrix value to send to each of the first and second optical modulators for encoding each optical signal with each matrix value.
12. The system of claim 11, wherein each optical modulator comprises a slow-light Mach- Zehnder modulator.
13. The device of claim 12, wherein each optical modulator has a phase shifter length in the range from about 150 pm to about 200 pm.
14. The device of claim 12, further comprising a ID dielectric photonic crystal waveguide.
15. The system of claim 11, wherein the device further comprises: a first set of K 1 x K optical power splitters for splitting the encoded signal corresponding to the matrix values x ... xKN each coupled to one of the first set of K optical modulators and coupled to K dot product engines; and a second set of K 1 x K optical power splitters for splitting the encoded signal corresponding to the matrix values yxl ... yNK each coupled to one of the second set of K optical modulators and coupled to K dot product engines.
16. The system of claim 11, wherein the device further comprises: a plurality of first sets of K - 1 optical power splitters, each set coupled in series to one of the first set of K optical modulators and each coupled to a dot product engine; a plurality of second sets of K- 1 optical power splitters, each set coupled in series to one of the second set of K optical modulators and each coupled to a dot product engine; wherein the splitting ratios of the first and second sets of optical power splitters are set at 1.(K - 1), 1 K - 2), . . . , 1 : 1 along the series of splitters.
17. The system of claim 11, wherein the device further comprises a multimode interferometer coupled to each of the optical modulators for providing the optical signal and wherein the multimode interferometer is configured to be coupled to a light source.
18. The system of claim 11, wherein each of the dot product engines comprises: a 2 * 2 optical power splitter, comprising a first arm and a second arm, wherein the first arm is coupled to one of the first set of K modulators and the second arm is coupled to one of the second set of K modulators; a 7t/2 phase shifter coupled to the second arm of the optical power splitter; a first photodiode coupled to the first arm of the optical power splitter; and a second photodiode coupled to the second arm of the optical power splitter; wherein the first and second photodiodes are configured to convert each encoded optical signal to a photocurrent signal which combine to form the product photocurrent signal.
19. The system of claim 18, wherein the device further comprises a temporal integrator coupled to each dot product engine for converting the product photocurrent signal to an integrated signal, comprising: at least two capacitors connected in parallel to the dot product engine for receiving the product photocurrent signal; one or more field effect transistors connected in parallel to the capacitors.
20. The system of claim 19, wherein the device further comprises an analog-to-digital converter associated with one or more of the temporal integrators for converting the integrated signal to a digital signal.
PCT/US2024/053575 2023-10-30 2024-10-30 Photonic tensor core devices and systems Pending WO2025096551A1 (en)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
US202363546305P 2023-10-30 2023-10-30
US63/546,305 2023-10-30
US202463713086P 2024-10-29 2024-10-29
US63/713,086 2024-10-29

Publications (1)

Publication Number Publication Date
WO2025096551A1 true WO2025096551A1 (en) 2025-05-08

Family

ID=95581452

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2024/053575 Pending WO2025096551A1 (en) 2023-10-30 2024-10-30 Photonic tensor core devices and systems

Country Status (1)

Country Link
WO (1) WO2025096551A1 (en)

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20190370644A1 (en) * 2018-06-04 2019-12-05 Lightmatter, Inc. Convolutional layers for neural networks using programmable nanophotonics
WO2022256905A1 (en) * 2021-06-11 2022-12-15 Huawei Technologies Canada Co., Ltd. System and method for optically performing computations using a photonic modulator
WO2023043712A1 (en) * 2021-09-14 2023-03-23 University Of Pittsburgh - Of The Commonwealth System Of Higher Education Systems and methods for coherent photonic crossbar arrays

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20190370644A1 (en) * 2018-06-04 2019-12-05 Lightmatter, Inc. Convolutional layers for neural networks using programmable nanophotonics
WO2022256905A1 (en) * 2021-06-11 2022-12-15 Huawei Technologies Canada Co., Ltd. System and method for optically performing computations using a photonic modulator
WO2023043712A1 (en) * 2021-09-14 2023-03-23 University Of Pittsburgh - Of The Commonwealth System Of Higher Education Systems and methods for coherent photonic crossbar arrays

Similar Documents

Publication Publication Date Title
TWI767877B (en) Optoelectronic processing system
Al-Qadasi et al. Scaling up silicon photonic-based accelerators: Challenges and opportunities
Filipovich et al. Silicon photonic architecture for training deep neural networks with direct feedback alignment
Farmakidis et al. Integrated photonic neuromorphic computing: opportunities and challenges
Peserico et al. Integrated photonic tensor processing unit for a matrix multiply: a review
CN112912900B (en) Optoelectronic computing system
CN112424796B (en) Optoelectronic computing system
TWI806042B (en) Optoelectronic processing apparatus, system and method
Mehrabian et al. PCNNA: A photonic convolutional neural network accelerator
CN111882052B (en) Photon convolution neural network system
JP2022523209A (en) Hybrid analog / digital matrix processor
CN111561953B (en) On-chip Optical Matrix-Vector Multiplier Based on Wavelength Division Multiplexing and Balanced Detection
Rahimi Kari et al. Realization of an integrated coherent photonic platform for scalable matrix operations
CN120525014A (en) Computing systems and computing devices
CN113496281B (en) Optoelectronic computing system
Zhang et al. Tempo: Efficient time-multiplexed dynamic photonic tensor core for edge ai with compact slow-light electro-optic modulator
US11700078B2 (en) Systems and methods for utilizing photonic degrees of freedom in a photonic processor
CN114326923B (en) Optical matrix vector multiplier based on polarization rotating beam splitter
CN114488650A (en) Silicon-based photonic integrated chip
Gu et al. All-integrated multidimensional optical sensing with a photonic neuromorphic processor
Kovaios et al. Scaling photonic neural networks: A silicon photonic GeMM leveraging a Time-Space multiplexed Xbar
WO2025096551A1 (en) Photonic tensor core devices and systems
Filipovich et al. Monolithic silicon photonic architecture for training deep neural networks with direct feedback alignment
CN118368023B (en) All-optical reconfigurable silicon-based photonic neural network chip based on wavelength division multiplexing
CN117642659A (en) Photon Computing Platform

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24886772

Country of ref document: EP

Kind code of ref document: A1