WO2025188570A1 - Stochastic optical computing architecture with reconfigurable logic and accumulation - Google Patents
Stochastic optical computing architecture with reconfigurable logic and accumulationInfo
- Publication number
- WO2025188570A1 WO2025188570A1 PCT/US2025/017914 US2025017914W WO2025188570A1 WO 2025188570 A1 WO2025188570 A1 WO 2025188570A1 US 2025017914 W US2025017914 W US 2025017914W WO 2025188570 A1 WO2025188570 A1 WO 2025188570A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- optical
- logic
- logic gate
- resonance frequency
- signals
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/06—Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons
- G06N3/067—Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons using optical means
-
- G—PHYSICS
- G02—OPTICS
- G02F—OPTICAL DEVICES OR ARRANGEMENTS FOR THE CONTROL OF LIGHT BY MODIFICATION OF THE OPTICAL PROPERTIES OF THE MEDIA OF THE ELEMENTS INVOLVED THEREIN; NON-LINEAR OPTICS; FREQUENCY-CHANGING OF LIGHT; OPTICAL LOGIC ELEMENTS; OPTICAL ANALOGUE/DIGITAL CONVERTERS
- G02F1/00—Devices or arrangements for the control of the intensity, colour, phase, polarisation or direction of light arriving from an independent light source, e.g. switching, gating or modulating; Non-linear optics
- G02F1/01—Devices or arrangements for the control of the intensity, colour, phase, polarisation or direction of light arriving from an independent light source, e.g. switching, gating or modulating; Non-linear optics for the control of the intensity, phase, polarisation or colour
- G02F1/015—Devices or arrangements for the control of the intensity, colour, phase, polarisation or direction of light arriving from an independent light source, e.g. switching, gating or modulating; Non-linear optics for the control of the intensity, phase, polarisation or colour based on semiconductor elements having potential barriers, e.g. having a PN or PIN junction
- G02F1/0151—Devices or arrangements for the control of the intensity, colour, phase, polarisation or direction of light arriving from an independent light source, e.g. switching, gating or modulating; Non-linear optics for the control of the intensity, phase, polarisation or colour based on semiconductor elements having potential barriers, e.g. having a PN or PIN junction modulating the refractive index
-
- G—PHYSICS
- G02—OPTICS
- G02F—OPTICAL DEVICES OR ARRANGEMENTS FOR THE CONTROL OF LIGHT BY MODIFICATION OF THE OPTICAL PROPERTIES OF THE MEDIA OF THE ELEMENTS INVOLVED THEREIN; NON-LINEAR OPTICS; FREQUENCY-CHANGING OF LIGHT; OPTICAL LOGIC ELEMENTS; OPTICAL ANALOGUE/DIGITAL CONVERTERS
- G02F3/00—Optical logic elements; Optical bistable devices
- G02F3/02—Optical bistable devices
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06E—OPTICAL COMPUTING DEVICES
- G06E1/00—Devices for processing exclusively digital data
- G06E1/02—Devices for processing exclusively digital data operating upon the order or content of the data handled
- G06E1/04—Devices for processing exclusively digital data operating upon the order or content of the data handled for performing computations using exclusively denominational number representation, e.g. using binary, ternary, decimal representation
- G06E1/045—Matrix or vector computation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06E—OPTICAL COMPUTING DEVICES
- G06E3/00—Devices not provided for in group G06E1/00, e.g. for processing analogue or hybrid data
- G06E3/001—Analogue devices in which mathematical operations are carried out with the aid of optical or electro-optical elements
- G06E3/005—Analogue devices in which mathematical operations are carried out with the aid of optical or electro-optical elements using electro-optical or opto-electronic means
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/044—Recurrent networks, e.g. Hopfield networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/048—Activation functions
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/06—Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons
- G06N3/063—Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons using electronic means
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/06—Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons
- G06N3/067—Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons using optical means
- G06N3/0675—Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons using optical means using electro-optical, acousto-optical or opto-electronic means
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
- G06N3/084—Backpropagation, e.g. using gradient descent
-
- B—PERFORMING OPERATIONS; TRANSPORTING
- B82—NANOTECHNOLOGY
- B82Y—SPECIFIC USES OR APPLICATIONS OF NANOSTRUCTURES; MEASUREMENT OR ANALYSIS OF NANOSTRUCTURES; MANUFACTURE OR TREATMENT OF NANOSTRUCTURES
- B82Y10/00—Nanotechnology for information processing, storage or transmission, e.g. quantum computing or single electron logic
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N10/00—Quantum computing, i.e. information processing based on quantum-mechanical phenomena
- G06N10/40—Physical realisations or architectures of quantum processors or components for manipulating qubits, e.g. qubit coupling or qubit control
Definitions
- aspects of the present disclosure relate to optical computing architecture with reconfigurable logic and accumulation.
- Optical computing has emerged as a promising technology for accelerating workloads like artificial intelligence.
- Optics provide inherent advantages in speed and parallelism over conventional electronics.
- Optical signals can be switched at sub-picosecond speeds, enabling high bandwidth density.
- multiplexing multiple wavelengths onto a single waveguide provides extensive parallelism. This allows high throughput optical matrix operations by imparting weights at each wavelength.
- existing optical computing architectures face substantial precision and efficiency challenges.
- optical accelerators To reach sufficient precision for an artificial intelligence system, optical accelerators often split digitally encoded values into shorter analog bit slices. Bit serial processing then recombines sliced results. However, throughput substantially declines due to the need for sequential slice-wise operations. Bit slicing also requires area-inefficient duplicate weighting structures. Overall efficiency suffers from the difficulty of precision scaling. While certain accelerators have adopted digital optics, they incur high area overheads rendering them impractical. So both digital and analog approaches face efficiency obstacles, reflected across optical matrix multiply architectures.
- an optical logic gate includes a single microring resonator having an input port and an output port; one or more input terminals connected to apply logic signals to the microring resonator; and a tuning component to shift a resonance frequency of the microring resonator wherein a transmission between the input port and output port of the microring resonator performs a logic operation on signals applied to the one or more input terminals.
- the substrate comprises a plurality of bitline groups in a DRAM bank, each bitline group comprising a set of bitlines corresponding to an input stochastic bit vector; sense amplifiers connected to the bitlines; stochastic-to-analog converter circuits connected between the sense amplifiers and analog lines of each bitline group; analog- to-unary converter circuits connected between the analog lines and the bitlines of each bitline group; and unary -to-binary converter circuits connected to the bitlines of each bitline group.
- Some embodiments include an optical logic gate comprising: a single microring resonator; input terminals connected to apply logic signals to the microring resonator; a tuning component to shift a resonance frequency of the microring resonator; and control logic configured to tune the resonance frequency to set a logic function of the optical logic gate by adjusting the tuning component.
- Some embodiments include an accumulation circuit comprising: a photodetector configured to receive a plurality of optical signals over a plurality of wavelengths and convert the optical signals to electrical pulses; an integration circuit connected to integrate the electrical pulses from the photodetector over time to produce an output voltage level proportional to a count of electrical pulses; and an output terminal to provide the output voltage level as an accumulated result of the plurality of optical signals.
- Some embodiments are directed to a computing system comprising: an optical source providing a plurality of optical wavelength channels; an array of optical logic gates, each optical logic gate operable to: receive a subset of the optical wavelength channels and two multi-bit operands, logically combine the multi-bit operands in a bitwise manner by selectively attenuating the subset of optical wavelength channels, and output the selectively attenuated subset as an optical output signal; accumulation circuits each connected to receive optical output signals from a group of the optical logic gates and configured to integrate the optical output signals over time to produce a voltage level representing a sum of the logically combined multi-bit operands; and analog-to-digital conversion circuitry connected to convert the voltage levels into digital results.
- FIG. 1A depicts an optical computing architecture (e.g., aggregation, modulation, modulation (AAM) vector dot product core (VDPC)) for performing vector matrix multiplications in accordance with examples of the present disclosure.
- optical computing architecture e.g., aggregation, modulation, modulation (AAM) vector dot product core (VDPC)
- FIG. IB depicts an optical computing architecture (e.g., modulation, aggregation, modulation (MAM) vector dot product core (VDPC)) for performing vector matrix multiplications in accordance with examples of the present disclosure.
- FIG. 1C depicts details of a summation element in accordance with examples of the present disclosure.
- FIG. 2 depicts details of a VDPC organization, in accordance with examples of the present disclosure.
- FIG. 3A depicts details of an Optical Stochastic Multiplier (OSM), in accordance with examples of the present disclosure.
- OSM Optical Stochastic Multiplier
- FIG. 3B depicts details of an Optical AND Gate (OAG), in accordance with examples of the present disclosure.
- OAG Optical AND Gate
- FIG. 3C depicts a cross-section view of an Optical AND Gate (OAG), in accordance with examples of the present disclosure.
- OAG Optical AND Gate
- FIGS. 3D-3I depict example transmission spectra for different logic-gate and complementary logic-gate functions corresponding to extracted transmission spectra for drop and through ports of an MRR, in accordance with examples of the present disclosure.
- FIG. 4 depicts additional details of the photo charge accumulator, in accordance with examples of the present disclosure.
- FIG. 5 illustrates a system-level implementation of a Stochastic Computing based Optical Neural Network Accelerator (SCONNA), in accordance with examples of the present disclosure.
- SCONNA Stochastic Computing based Optical Neural Network Accelerator
- FIGS. 5A-5H depict example timing diagrams, in accordance with examples of the present disclosure.
- FIG. 6 depicts details of a XNOR-Bitcount based Binary Neural Network Accelerator (OXBNN), in accordance with examples of the present disclosure.
- FIG. 6 depicts an example method for performing stochastic-to-binary conversion of input operands, in accordance with examples of the present disclosure.
- FIG. 7 depicts details of an example PCA circuit, in accordance with examples of the present disclosure.
- FIG. 8 depicts a system-level implementation of an OXBNN accelerator, in accordance with examples of the present disclosure.
- aspects of the present disclosure provide apparatuses, methods, processing systems, and computer-readable mediums for stochastic optical computing architecture with reconfigurable logic and accumulation.
- Optical computing has emerged as a promising technology for accelerating workloads like artificial intelligence.
- existing optical computing architectures face substantial precision and efficiency challenges.
- Many prior optical accelerators rely on analog optics, restricting scalability due to the tradeoff between precision and parallel wavelength channels. This highlights the need for an optimized optical computing architecture tailored to overcome these limitations.
- a technical problem includes that both digital and analog optical computing approaches face efficiency obstacles, reflected across optical matrix multiply architectures. Combined with limited power, these factors restrict the bit precision, with 4-6 bits typical in prior architectures.
- aspects of the present disclosure provide a stochastic computing based optical architecture with reconfigurable logic gates and specialized accumulation circuits.
- Such embodiments may take the form of a data center server and/or other computing architecture.
- the architecture encodes data as streams of random bits corresponding to probabilities between zero and one. Logic operations between stochastic bit streams effectively multiply their probabilities using simple optical logic. Moreover, encoding data as stochastic bit streams relaxes precision requirements.
- the architecture also includes polymorphic optical logic gates enabling different logical functions for improved efficiency through area reuse.
- the accumulation circuits integrate streams of optical bits over time to produce a voltage level result.
- some embodiments may be configured utilizing the optical logic gates and the accumulation circuits to enable bitwise processing layers in an artificial intelligence system. This architecture provides an optimized solution customized for stochastic optical computing to help fully realize its acceleration potential.
- FIG. 1A depicts an optical computing architecture (e.g., aggregation modulation modulation (AAM) vector dot product core (VDPC)) for performing vector matrix multiplications.
- An array 102 of laser diodes 104 emit light at a plurality of wavelengths i through N, generating N optical wavelength channels.
- An aggregation block 108 aggregates the generated optical wavelength channels into a single photonic waveguide through dense wavelength division multiplexing (DWDM) (using an Ax 1 multiplexer 106) and then splits the optical power of these N wavelength channels equally into M separate waveguides (using a l x M splitter 110).
- DWDM dense wavelength division multiplexing
- a modulation block also referred to as DIV block 112 employs M arrays of microring resonators (MRRs) 114 (one array per waveguide, with each array having N MRRs; each array referred to as DIV element) to imprint M DIVs of N points each onto the N'M wavelength channels by modulating the analog power amplitudes of the wavelength channels.
- MRRs microring resonators
- DKV block 116 employs another M arrays of MRRs (one array per waveguide, with each array having N MRRs; each array referred to as DKVelement) to further modulate the V x M wavelength channels with DKVs, so that the analog power amplitudes of the individual wavelength channels then represent the point-wise products of the utilized DKVs 116 and DIVs 112.
- MRRs one array per waveguide, with each array having N MRRs; each array referred to as DKVelement
- a summation block (SB) 120 employs M summation elements (SEs) 122, with each SE 122 having two balanced photodiodes (PDs) upon which the point- wise-product-modulated N wavelength channels are incident to produce an output current that is proportional to the result of the vector dot product (VDP) operation between the corresponding DKV 116 block and DIV 112.
- the array 102 of laser diodes 104 and SB 120 may be positioned at the two ends of the VDPC, with the aggregation, modulation (DIV block 112), and modulation (DKV block 116) blocks placed in between them.
- the MRR-based VDPC organizations may be classified as a MAM (modulation, aggregation, modulation) or AMM (aggregation, modulation, modulation) VDPC organization.
- the AMM VDPC organization positions the aggregation block 108 after the array 102 of laser diodes 104, and then the DIV block 112 followed by the DKV 116 modulation block.
- the MAM VDPC in FIG. IB positions the DIV block 124 first after the array 102 of laser diodes 104, and then positions the aggregation block 108 followed by the DKV block 116.
- the MAM DIV block 124 is structurally different from the AMM DIV block 112.
- the MAM DIV block 124 employs one MRR 126 per waveguide, and as a result, it can imprint one DIV with N points onto the N wavelength channels.
- This one DIV (e.g., one array per waveguide, each waveguide having one MRR) is shared among all DKVs 118 in the MAM VDPC 130, whereas each DKV 118 can have a different DIV corresponding to it in the AMM VDPC 100.
- M generally is equal N for many MAM and AMM VDPC.
- VDPE VDP element
- FIG. 1C depicts details of a summation element 122 in accordance with examples of the present disclosure.
- the summation element 122 can include a pair of balanced photodiodes 152 and 154 (e.g., balanced photodiode pair) coupled to a trans-impedance amplifier 156.
- the trans impedance amplifier 156 receives an input current from the balanced photodiodes 152 and 154 and converts the input current to an output voltage.
- FIG. 2 depicts a VDPC organization in accordance with examples of the present disclosure. Similar to the VDPCs of analog optical accelerators, the example VDPC depicted in FIG. 2 also implements multiple VDP operations in parallel. Thus, an array 202 of total V single-wavelength laser diodes (EDs) are used, with each LD sourcing optical power of amount at a distinct wavelength Xi. The total power from all N EDs (Xi to XN) is multiplexed into a single photonic waveguide through wavelength division multiplexing (WDM) using the wavelength division multiplexor 204. These multiplexed wavelengths are split (using a splitter 206) into M input waveguide arms (IWAs).
- WDM wavelength division multiplexing
- Each IWA receives N- wavelength optical power and guides it to a VDPE (e.g., 208).
- VDPE e.g., 208
- Each VDPE 207 includes at least three components: a cascade of N Optical Stochastic Multipliers (OSMs) 208A-208N; (ii) a bank of fdter MRRs (e.g., 214A- 214N); and (iii) a Photo-Charge Accumulator (PCA) pair.
- OSMs N Optical Stochastic Multipliers
- PCA Photo-Charge Accumulator
- Each OSM (e.g., 208) performs stochastic multiplication between an input bitstream I (e.g., 210) (corresponding to a point in an V-point DIV) and weight bit-stream W (e.g., 212) (corresponding to a point in an V-point DKV).
- Each OSM 208 receives its bit-streams 7210 and from its corresponding peripherals at a supported bitrate (BR).
- Bitstream IP 212 provides a weight value along with a sign bit.
- Bitstream I 210 provides an RELU-activated output value from the previous CNN layer, without a sign bit as RELU has a nonnegative output.
- Each OSM 208 performs a bitwise logical AND operation between the I 210 and IT 212 bit-streams to produce a resultant optical bit-stream that represents the stochastic multiplication between the 7210 and W 220 bit-streams.
- each fdter MRR (e.g., 214A, 214B, 214N) operates on a distinct optical bit-stream Xi.
- Each fdter MRR receives the sign bit from the peripheral W 212 of its corresponding OSM 208.
- the sign bit operates the filter to steer the incoming optical bit-stream Xi to the output waveguide arm OWA (if the sign bit is ‘0’) or OWA’ (if the sign bit is ‘ 1’).
- OWA and OWA’ of a VDPE guide the optical bit-streams, carrying the stochastic multiplication results, to PC As 216A/216B.
- a PC A 216 can include a circuit that collects the optical bit-streams (i.e., stochastic multiplication results) from its corresponding OWA (or OWA’) and generates the accumulation result in a binary format as further described below.
- the accumulated result represents an inner product between optical vector inputs.
- the OWA coupled PCA e.g., 216A
- combines with the corresponding OWA’ coupled PCA e.g., 216B
- an OSM 300 can include peripherals 302 and an optical logic gate, such as an optical ‘AND’ Gate (OAG) 304.
- the optical logic gate maintains operating parameters while reconfigured between logic functions.
- the operating parameters include optical power, bit precision, bit rate, and footprint.
- the peripherals 302 convert a binary input value 4 (e.g., 306) and binary weight value H'h (308) retrieved from a buffer 310 into unipolar stochastic bitstreams I and W.
- the OAG 304 performs a multiplication-equivalent bitwise AND operation between the stochastic bit- streams I and W.
- the peripherals 302 of the OSM 300 can use a lookup table 312 and serializers 314A/314B to generate a combination of unipolar stochastic bit-streams 1314A and W 314B.
- Two unipolar stochastic bitstreams 316A and 316B for their eventual error-free multiplication using an AND gate, can be generated in combination with each other, so that they are uncorrelated, i.e., the marginal probability of one bit-stream (i.e., / 316 A or W 316B) is equal to its conditional probability given the other bit-stream (i.e., I given W or W given I).
- the marginal probability of one bit-stream i.e., / 316 A or W 316B
- I conditional probability given the other bit-stream
- all combinations of uncorrelated bit-streams I and W can be generated a priori (offline) using a unipolar circuit, and then stored as bit-streams in the bit-vector (bitparallel) format in the lookup table 316.
- each entry in the lookup table 316 stores a combination of uncorrelated bit- vectors Iv and Wv.
- the OSM 300 utilizes a unique identifier for each combination of binary values lb 306 and Wb 308 (that are accessed from a buffer 310, such as a scratchpad memory by performing an XOR- based hash function Ib®Wb (e.g.,).
- Ib®Wb XOR- based hash function
- the OSM 300 uses a Ib®Wb value to fetch the desired combination of Iv and Wv from the lookup table 316.
- the OSM 300 pushes the Iv and Wv through dedicated high-speed serializers 320A/320B, to generate bit-streams 1314A and W 314B.
- the stochastic bit-streams I and W, generated by the peripherals 302 of the OSM 300, are then fed to the OAG 304 via highspeed drivers 322A/322B for stochastic multiplication. Design aspects of the OAG 304 is further described in FIG. 3B.
- the OAG 304 can be an add-drop microring resonator (MRR) 326 having two operand terminals (realized as embedded PN-junctions 324A/324B) that can take two stochastic bit-streams I 314A and W 314B as inputs at a predefined bitrate (BR).
- MRR microring resonator
- the MRR’s temperature can be increased using the integrated microheater 328 to consequently tune its operand-independent resonance from its fabrication-defined initial position y to its programmed position r
- the MRR’s 326 resonance passband electro-refractively moves to an operand-driven position.
- the drop-port 330 transmission (T(kin)) of the MRR 326 provides bit-wise logical AND operation between the inputs 1314A and W 314B.
- the passbands of the MRR 326 for different operand inputs and temperature conditions are further depicted in FIGS. 3D-3I.
- FIG. 3B depicts additional details of a microring (MRR) 332 based optical stochastic multiplier (OSM) 304A in accordance with examples of the present disclosure.
- MRR microring
- OSM optical stochastic multiplier
- the MRR 332 can accept both input operands electrically, and its drop-port (through-port) optical response can be thermo-optically programmed to make it dynamically follow the truth table of different logic functions, such as AND, OR and XOR (NAND, NOR, and XNOR), at different times.
- the control logic configures the optical logic gate to perform XOR logic by controlling the tuning component to shift the resonance frequency to a minimum transmission position relative to an optical input wavelength.
- E-0 circuits built using the MRR 332 can address the above-described shortcomings by providing (1) the ability of all-electrical application of the input operands, (2) compactness through a single-MRR structure of the E-0 gate, and (3) high flexibility through the introduced polymorphism, and consequently, low idle time and improved suitability for use with SIMD/MIMD/SA based architectures.
- the MRR 332 is similar to an add-drop MRR, but includes four quarter-sized phase-shifting sections embedded therein. In examples, two quarter-sized sections of the MRR 332 are two PN junctions 324A/324B which are operated in the forward bias condition, whereas the remaining two quarter-sized sections integrate micro- heaters 328A/328B.
- a cross-section 334 of a PN-junction based section of the MRR 332 is depicted in FIG. 3C and includes a ridge waveguide 336 with an embedded lateral PN junction 324B, fabricated on a buried oxide layer 338.
- Example dimensions of the P-type regions 340, 342, and 344 and N-type regions, 346, 348, and 350 and corresponding example carrier concentrations are also depicted in FIG. 3C.
- the PN junction based sections (e.g., 324A/324B) of the MRR 332 work at the input terminals where the input logic signals/operand bits are applied.
- the microheaters 328A/328B integrated sections of the MRR 332 work as programming terminals that are used to program the MRR to perform specific logic-gate functions. Applying a voltage to the microheaters 328A/328B based programming terminals can increase the temperature of the MRR 332, which in turn can shift (red shift) the resonance of the MRR 332 towards a longer wavelength. This is because of the thermo-optic effect in silicon.
- the operand-independent MRR resonance i.e., the programmed MRR resonance
- the electrical input logic signals or input operand bits are applied to the PN junctions (e.g., 324A/324B) based input terminals of the MRR 332.
- the resonance of the MRR 332 shifts (blue shifts) towards a shorter wavelength depending on the combination of the applied input operand bits.
- the through-port 352 and drop-port 354 optical responses of the MRR 332 follow a truth-table of a logic-gate function for which the MRR 332 is programmed. In this manner, the MRR 332 can perform different logic-gate functions at different times. At any given time, the through-port 352 optical response of the MRR 332 follows logical complement of the drop-port 354 optical response. Therefore, AND, OR and XOR functions can be realized (one function at a time) at the drop-port 354 of the MRR 332. Concurrently, the through-port 352 of the MRR 332 can provide complementary logic gate functions such as NAND, NOR and XNOR as discussed below.
- FIGS. 3D-3I depict example transmission spectra for different logic-gate and complementary logic-gate functions corresponding to extracted transmission spectra for drop and through ports of an MRR 332. More specifically, the extracted transmission spectra for different values of the detuning of the operand-independent MRR resonance position K with respect to the input wavelength kin are depicted in FIGS. 3D-3I. As previously described, the detuning values correspond to different logic-gate functions that the MRR 332 can perform. In addition, transmission spectra for different combinations of the input operand bits are also depicted in FIGS. 3D-3I.
- Transmission spectra corresponding to logic-gate functions AND, OR, and XOR are depicted in FIGS. 3D, 3F, and 3H, respectively. These transmission spectra are drop-port transmission spectra (Torentzian lineshape passbands). Similarly, the transmission spectra corresponding to complementary logic-gate functions NAND, NOR, and XNOR are depicted in in FIGS. 3E, 3G, and 31, respectively. These transmission spectra are through-port transmission spectra (inverse Torentzian lineshape passbands). As depicted in FIG. 3D, the drop port and through port transmission exhibits two clearly distinguishable levels.
- the full transmission range at the drop port and through port of the MRR 332 is divided into two areas, in which the lower part of the full transmission range is indicated as logic 0, whereas the upper part is indicated as logic 1. If the drop port (DT(Xin)) and through port (TT(Xin)) transmission at Xin falls in the lower part of the full transmission range, then it is referred to as logic ‘0’ transmission. On the other hand, if the drop port and through port transmission at X m falls in the upper part of the full transmission range, then it is referred to as logic ‘ 1 ’ transmission. However, the vertical spans of the two distinguishable transmission levels differ between the drop port and through port.
- SOMA optical modulation amplitude
- an example 0.9 V voltage (3.52 mW power) is applied to the programming terminals of the MRR 332, shifting the resonance from the initial position, r
- the applied bit-combination (I, W) is (0,0)
- the resonance position of the MRR stays at K and the drop port transmission at Xin provides logic ‘0’ level (the bottom dot on the Y-axis).
- the applied bit-combination (7, W) is (0,1) or (1,0)
- the position of the MRR resonance changes, but the blueshift is the same for both (0,1) and (1,0) bit combinations, and the drop port transmission at A m still remains at logic ‘0’ level (the top dot on the Y-axis).
- the applied bit-combination (7, IF) is (1,1)
- the MRR resonance undergoes a larger blueshift, and the position of the passband with respect to A m changes.
- shifting the resonance frequency includes performing AND logic on the signals applied to the one or more input terminals when the resonance frequency provides a peak transmission at the predefined optical input wavelength.
- this AND function at the drop port of the MRR 332 corresponds to NAND function at the through port of the MRR 332 as illustrated in FIG. 3D).
- the MRR 332 can be reconfigured to implement OR (NOR) and XOR (XNOR) gate functions as well, by applying a suitable voltage to the programming terminals of the MRR 332 to set the relative position of K with respect to in as shown in FIGS. 3E, 3G, and 31.
- OR OR
- XNOR XOR
- FIG. 4 depicts additional details of the photo charge accumulator 216 of FIG. 2 in accordance with examples of the present disclosure.
- the stochastic multiplication bit-streams generated by OSMs 208 are guided to a PCA 216, where they are accumulated to generate a binary output value equivalent to the VDP result.
- the PCA 216 may include has two stages: (i) a stochastic-to-analog conversion stage; and (ii) an analog-to-binary conversion stage.
- the stochastic-to-analog stage employs a photodetector 402 and two TIR circuits 404A and 404B, whereon one TIR circuit remains redundant, enabled by the demux 406 and mux 408.
- the photodetector 402 generates a current pulse for each optical logic ‘ 1’ incident upon it.
- This current pulse accumulates a certain amount of charge on the capacitor of the active TIR circuit (e.g., the circuit with Cl capacitor); as a result, the capacitor accrues an analog voltage level.
- the total accumulated charge (and thus, the accrued analog voltage level) on the active capacitor (e.g., Cl) is proportional to the total number of 1’s in the incident bitstreams.
- the number of 1’s that can be accumulated in such manner might be limited, as the charge across the capacitor of TIR circuit 404B can saturate.
- the active capacitor e.g., Cl
- capacitor C2 of the redundant TIR circuit e.g., 404A
- the output analog voltage computed by the stochastic-to-analog conversion stage represents the unipolar unsealed addition of the stochastic bit-streams.
- the analog-to-binary stage of the PCA circuit employs an analog-to-digital converter ADC 410. This binary value is the VDP result.
- FIG. 5 illustrates a system-level implementation of a Stochastic Computing based Optical Neural Network Accelerator (SCONNA) 500 in accordance with examples of the present disclosure.
- the SCONNA 500 may include a global memory 502 for storing CNN parameters, and a preprocessing and mapping unit 504 for decomposing the tensors into DIVs/DKVs and mapping them onto VDPEs.
- the SCONNA 500 can include a mesh of tiles 506A-506N coupled to routers 510, and this mesh network facilitates parameter communication among tiles 506A-506N with network interfaces 508.
- Each tile 506 can include one or more SCONNA VDPCs 514A-514N interconnected (via H-tree network) with output buffer 522, activation 518, and pooling 520 units.
- each tile 506 can include a sum reduction network 516 configured to accumulate a number of partial sum values.
- FIG. 6 depicts details of a XNOR-Bitcount based Binary Neural Network Accelerator (OXBNN) in accordance with examples of the present disclosure.
- Binary Neural Networks (BNNs) are increasingly preferred over full-precision Convolutional Neural Networks (CNNs) to reduce the memory and computational requirements of inference processing with minimal accuracy drop.
- BNNs convert CNN model parameters to 1-bit precision, allowing inference of BNNs to be processed with simple XNOR and bitcount operations. This makes BNNs amenable to hardware acceleration.
- the OXBNN architecture includes an XNOR-Bitcount Processing Core (XPC) 600, which may include an array 602 of total N single-wavelength laser diodes (EDs), with each LD sourcing optical power of P amount at a distinct wavelength
- XPC XNOR-Bitcount Processing Core
- EDs total N single-wavelength laser diodes
- WDM wavelength division multiplexing
- the optical power containing all these A wavelengths is split (e.g., at splitter 606) into M input waveguides, each of which connects to an XNOR-Bitcount Processing Element (XPE) 620.
- An XPC 600 can include a total of A/XPEs 620.
- an XPE 620 in the OXBNN architecture can include: (i) an array 608 of a total of N Optical XNOR Gates (OXGs) (e.g., 616A-616N) that generates an XNOR vector (or an XNOR vector slice) containing N optical bits, and (ii) a Photo-Charge Accumulator (PCA) 612A that performs a bitcount on the generated XNOR vector (or XNOR vector slice).
- the value A here which is equal to the number of wavelengths and number of OXGs 616 per XPE 620, is referred to as the size of the XPE 620.
- an array of a total of N OXGs 616 couples to an input waveguide (e.g., 601A, 601B, 601N) as depicted in FIG. 6.
- Each OXG 616 operates upon a unique wavelength traversing the input waveguide (e.g., 601).
- Each OXG 616 in the array electrically receives two binary operands (i.e., input bit and weight bit w ) from its corresponding drivers (not shown).
- Each OXG 616 in the array produces one bit of the resultant XNOR vector slice, and it imprints this bit on its corresponding Ai (by modulating the optical transmission at - to be consequently guided to the bitcount circuit (i.e., PCA 612) via the output waveguide 610A-610N.
- the PCA 612 receives the N individual optical bits of the X-bit XNOR vector slice concurrently on X distinct wavelengths.
- the PCA 612 performs bitcount on these optical bits.
- This processing step from the bit parallel application of the binary input and weight vector slices at the electrical input terminals of the array of N OXGs 616 to the generation of the bitcount result by the PCA 612, occur with low latency because of the lightspeed operation of the XPE 620.
- This processing step mapped on an XPE 620 can be referred to as a PASS and the corresponding latency as r.
- the XPE 620 can produce one bitcount result for one XNOR vector slice in every single PASS with T latency.
- the XPE 620 can achieve very high processing throughput by completing one PASS every T period.
- multiple input and weight vector slices ⁇ I- , 1 2 , ⁇ , I a ⁇ and ⁇ Wi , VP 2 , ... , W a can be applied to the array of OXGs 616 of an XPE 620
- the Optical XNOR Gate (OXG) 616 depicted in FIG. 6 includes an add-drop microring resonator (MRR), which has two operand terminals (realized as embedded PN- junctions) that can take two operand bits z and w as inputs for a predefined time-width (usually a little less than the T period).
- FIG. 3H depicts the passbands of the MRR for different operand inputs and temperature conditions.
- the MRR’s temperature can be increased using the integrated microheater 328 (e.g., FIG. 3A), to consequently tune its operand-independent resonance from its fabrication-defined initial position operand-independent MRR resonance position r
- the MRR’s resonance passband electro-refractively moves to an operand-driven position.
- the through-port transmission (T(kin)) of the MRR provides bit-wise logical XNOR operation between the input bits z and w.
- the XNOR vector bits generated by an array of OXGs are guided to a PCA circuit (e.g., PCA circuit 702 of FIG. 7), where a bitcount is performed on the XNOR vector bits to generate an output result.
- the PCA circuit 702 employs a photodetector 704 and two time integrating receiver (TIR) circuits 706A and 706B, where one of the TIR1 and TIR2 circuits remains redundant, enabled by the demux 708 and mux 710.
- the photodetector 704 generates a current pulse for each optical logic ‘ 1 ’ incident upon it.
- the amplitude of a current pulse generated for an optical logic ‘0’ remains under the noise limit; therefore, a logic ‘0’ remains statistically undetected.
- the current pulse generated by an optical logic ‘ 1’ accumulates a certain statistically significant amount of charge on the capacitor of the active TIR circuit (e.g., the circuit with Cl capacitor); as a result, the TIR circuit 706B outputs a detectable analog voltage level.
- the total accumulated charge on the active capacitor (e.g., Cl) grows proportionally to the total number of optical ‘ 1’s that are incident.
- the number of ‘ l’s that can be accumulated in such a manner might be limited, as the output of the TIR circuit 706B might saturate.
- the ongoing accumulation phase ends and the bitcount result (i.e., the final TIR output voltage) is passed through a comparator 712 to generate the activation value for a next BNN layer.
- a discharge of the active capacitor e.g., Cl
- capacitor Cl is discharging, the redundant TIR2 circuit 706 A with capacitor C2 mitigates the discharge latency by allowing a continuation of a concurrent bitcount.
- FIG. 8 depicts a system-level implementation 800 of an OXBNN accelerator.
- the implementation 800 includes a global memory 802 that stores BNN parameters and a preprocessing and mapping unit 804.
- the implementation 800 includes a mesh network of tiles 806A-806N.
- Each tile 806 includes a network interface 808A-N and, as an example, 4 XPCs 814A-814N interconnected (via H-tree) with an output buffer 822 and pooling unit(s) 820.
- a first aspect includes an optical logic gate comprising: a microring resonator having an input port and an output port; one or more input terminals connected to apply logic signals to the microring resonator; and a tuning component to shift a resonance frequency of the microring resonator wherein a transmission between the input port and the output port of the microring resonator performs a logic operation on signals applied to the one or more input terminals.
- a second aspect includes the first aspect, wherein shifting the resonance frequency sets a transmission level of the output port relative to a predefined optical input wavelength to the input port.
- a third aspect includes the first aspect and/or the second aspect, wherein shifting the resonance frequency includes performing AND logic on the signals applied to the one or more input terminals when the resonance frequency provides a peak transmission at the predefined optical input wavelength.
- a fourth aspect includes any of the first aspect through the third aspect, wherein shifting the resonance frequency includes performing OR logic on the signals applied to the one or more input terminals when the resonance frequency provides an intermediate transmission level at the predefined optical input wavelength.
- a fifth aspect includes any of the first aspect through the fourth aspect, wherein shifting the resonance frequency includes performing XNOR logic on the signals applied to the one or more input terminals when the resonance frequency provides minimum transmission at the predefined optical input wavelength.
- a sixth aspect includes any of the first aspect through the fifth aspect, wherein the one or more input terminals comprise embedded PN junctions for carrier density modulation to shift the resonance frequency.
- a seventh aspect includes any of the first aspect through the sixth aspect, wherein the tuning component to shift the resonance frequency comprises a microheater disposed adjacent to the microring resonator.
- An eighth aspect includes any of the first aspect through the seventh aspect, wherein the microring resonator, the one or more input terminals, and the tuning component are integrated on a silicon photonic integrated circuit.
- a ninth aspect includes any of the first aspect through the eighth aspect, wherein the one or more input terminals modulate a refractive index of the microring resonator based on applied logic signals.
- a tenth aspect includes any of the first aspect through the ninth aspect, wherein the logic operation on the signals applied to the one or more input terminals comprises AND, OR, XOR, NAND, NOR, or XNOR logic.
- An eleventh aspect includes an optical logic gate comprising: a microring resonator; input terminals connected to apply logic signals to the microring resonator; a tuning component to shift a resonance frequency of the microring resonator; and control logic configured to tune the resonance frequency to set a logic function of the optical logic gate by adjusting the tuning component.
- a twelfth aspect includes the eleventh aspect, wherein the control logic configures the optical logic gate to perform AND logic by controlling the tuning component to shift the resonance frequency to a peak transmission position relative to an optical input wavelength.
- a thirteenth aspect includes the eleventh aspect and/or the twelfth aspect, wherein the control logic configures the optical logic gate to perform OR logic by controlling the tuning component to shift the resonance frequency to an intermediate transmission position relative to an optical input wavelength.
- a fourteenth aspect includes any of the eleventh aspect through the thirteenth aspect, wherein the control logic configures the optical logic gate to perform XOR logic by controlling the tuning component to shift the resonance frequency to a minimum transmission position relative to an optical input wavelength.
- a fifteenth aspect includes any of the eleventh aspect through the fourteenth aspect, wherein the control logic dynamically reconfigures the logic function over time by adjusting the tuning component to shift the resonance frequency of the microring resonator.
- a sixteenth aspect includes any of the eleventh aspect through the fifteenth aspect, wherein dynamic reconfiguration enables time-division multiplexing of logical operations.
- a seventeenth aspect includes any of the eleventh aspect through the sixteenth aspect, wherein the tuning component comprises microheaters to thermally tune a refractive index of the microring resonator.
- An eighteenth aspect includes any of the eleventh aspect through the seventeenth aspect, wherein the control logic controls electrical power supplied to the microheaters to shift the resonance frequency by a predefined detuning amount associated with a target logic function.
- a nineteenth aspect includes any of the eleventh aspect through the eighteenth aspect, wherein the input terminals comprise PN junctions.
- a twentieth aspect includes any of the eleventh aspect through the nineteenth aspect, wherein the optical logic gate maintains operating parameters while reconfigured between logic functions.
- a twenty-first aspect includes any of the eleventh aspect through the twentieth aspect, wherein the operating parameters include optical power, bit precision, bit rate, and footprint.
- a twenty-second aspect includes any of the eleventh aspect through the twenty-first aspect, wherein the control logic receives commands identifying a target logic function and controls the tuning component to configure the optical logic gate with the target logic function.
- a twenty -third aspect includes any of the eleventh aspect through the twenty-second aspect, wherein the input terminals modulate a refractive index of the microring resonator based on applied logic signals.
- a twenty-fourth aspect includes any of the eleventh aspect through the twenty -third aspect, wherein the control logic and the tuning component are integrated with the microring resonator on a photonic integrated circuit.
- a twenty-fifth aspect that includes an accumulation circuit comprising: a photodetector configured to receive a plurality of optical signals over a plurality of wavelengths and convert the plurality of optical signals to electrical pulses; an integration circuit connected to integrate the electrical pulses from the photodetector over time to produce an output voltage level proportional to a count of electrical pulses; and an output terminal to provide the output voltage level as an accumulated result of the plurality of optical signals.
- a twenty-sixth aspect includes the twenty-fifth aspect, wherein the output voltage level represents a sum of incident optical pulse amplitudes within the plurality of optical signals.
- a twenty-seventh aspect includes the twenty-fifth aspect and/or the twenty-sixth aspect, wherein the integration circuit comprises one or more capacitors charged by the electrical pulses.
- a twenty-eighth aspect includes any of the twenty-fifth aspect through the twentyseventh aspect, further comprising switching logic to alternate integration across the one or more capacitors.
- a twenty -ninth aspect includes any of the twenty-fifth aspect through the twentyeighth aspect, wherein the photodetector concurrently receives the plurality of optical signals without wavelength filtering.
- a thirtieth aspect includes any of the twenty-fifth aspect through the twenty-ninth aspect, wherein the plurality of optical signals encode probabilistic values as streams of optical bits.
- a thirty-first aspect includes any of the twenty-fifth aspect through the thirtieth aspect, wherein the output voltage level represents a multiplication of the probabilistic values encoded in the plurality of optical signals when computed as a function of optical bit logical products.
- a thirty-second aspect includes any of the twenty-fifth aspect through the thirty- first aspect, wherein the output voltage level saturates upon receiving a threshold number of optical pulses.
- a thirty-third aspect includes any of the twenty-fifth aspect through the thirty- second aspect, further comprising a reset switch connected across integration capacitors to enable discharge upon saturation.
- a thirty-fourth aspect includes any of the twenty-fifth aspect through the thirty- third aspect, wherein the photodetector, the integration circuit and the output terminal are integrated on an integrated circuit.
- a thirty-fifth aspect includes any of the twenty-fifth aspect through the thirty-fourth aspect, wherein the photodetector comprises a balanced photodiode pair connected to the integration circuit.
- a thirty-sixth aspect includes any of the twenty-fifth aspect through the thirty-fifth aspect, further comprising a transimpedance amplifier coupling the photodetector to the integration circuit.
- a thirty-seventh aspect includes any of the twenty-fifth aspect through the thirtysixth aspect, wherein the accumulated result represents an inner product between optical vector inputs.
- a thirty-eighth aspect includes any of the twenty-fifth aspect through the thirtyseventh aspect, wherein the plurality of optical signals are received over at least 50 wavelength channels with at least 0.7 nm channel spacing.
- a thirty-ninth aspect includes a computing system comprising: an optical source providing a plurality of optical wavelength channels; an array of optical logic gates, each optical logic gate operable to: receive a subset of the plurality of optical wavelength channels and two multi-bit operands, logically combine the two multi-bit operands in a bitwise manner by selectively attenuating the subset of the plurality of optical wavelength channels, and output the selectively attenuated subset as an optical output signal; accumulation circuits each connected to receive optical output signals from a group of the array of optical logic gates and configured to integrate the optical output signals over time to produce a voltage level representing a sum of the logically combined multi-bit operands; and analog-to-digital conversion circuitry connected to convert the voltage level into digital results.
- a fortieth aspect includes the thirty -ninth aspect, wherein the array of optical logic gates each include a single microring resonator structure.
- a forty-first aspect includes the thirty -ninth aspect and/or the fortieth aspect, wherein the array of optical logic gates are dynamically reconfigurable to selectively perform AND, OR, XOR, NAND, NOR, or XNOR logical operations on the two multi-bit operands.
- a forty-second aspect includes any of the thirty -ninth aspect through the forty-first aspect, wherein the two multi-bit operands are provided to the array of optical logic gates in a stochastic bit stream format encoding probabilistic values between 0 and 1.
- a forty -third aspect includes any of the thirty-ninth aspect through the forty-second aspect, wherein a length of the stochastic bit streams sets a precision of probabilistic values.
- a forty-fourth aspect includes any of the thirty-ninth aspect through the forty -third aspect, wherein the accumulation circuits are each configured to concurrently receive the optical output signals over multiple wavelengths without wavelength filtering.
- a forty-fifth aspect includes any of the thirty -ninth aspect through the forty-fourth aspect, further comprising control logic connected to reconfigure functions of both the array of optical logic gates and the accumulation circuits over time.
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Biophysics (AREA)
- Biomedical Technology (AREA)
- Mathematical Physics (AREA)
- Computing Systems (AREA)
- Molecular Biology (AREA)
- General Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Software Systems (AREA)
- General Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Data Mining & Analysis (AREA)
- Evolutionary Computation (AREA)
- Nonlinear Science (AREA)
- Optics & Photonics (AREA)
- Neurology (AREA)
- Pure & Applied Mathematics (AREA)
- Algebra (AREA)
- Mathematical Analysis (AREA)
- Optical Modulation, Optical Deflection, Nonlinear Optics, Optical Demodulation, Optical Logic Elements (AREA)
- Logic Circuits (AREA)
Abstract
Systems and method directed to stochastic optical computing architecture with reconfigurable logic and accumulation are disclosed. In one embodiment, an optical logic gate includes a single microring resonator having an input port and an output port; one or more input terminals connected to apply logic signals to the microring resonator; and a tuning component to shift a resonance frequency of the microring resonator, wherein a transmission between the input port and output port of the microring resonator performs a logic operation on signals applied to the one or more input terminals.
Description
PCT APPLICATION FOR:
STOCHASTIC OPTICAL COMPUTING ARCHITECTURE WITH RECONFIGURABLE LOGIC AND ACCUMULATION
STOCHASTIC OPTICAL COMPUTING ARCHITECTURE WITH RECONFIGURABLE LOGIC AND ACCUMULATION
GOVERNMENT SUPPORT
[0001] This invention was made with government support under 2139167 from the National Science Foundation. The Government has certain rights to the invention.
BACKGROUND
Field
[0002] Aspects of the present disclosure relate to optical computing architecture with reconfigurable logic and accumulation.
Description of Related Art
[0003] Optical computing has emerged as a promising technology for accelerating workloads like artificial intelligence. Optics provide inherent advantages in speed and parallelism over conventional electronics. Optical signals can be switched at sub-picosecond speeds, enabling high bandwidth density. Furthermore, multiplexing multiple wavelengths onto a single waveguide provides extensive parallelism. This allows high throughput optical matrix operations by imparting weights at each wavelength. However, existing optical computing architectures face substantial precision and efficiency challenges.
[0004] Many prior optical accelerators rely on analog optics, encoding values in the signal amplitude. However, optical power budgets are highly constrained, limiting distinguishable levels and therefore precision. Supporting multiple analog levels also restricts scalability due to the tradeoff between precision and parallel wavelength channels. Additional challenges arise from the intrinsic losses in nanophotonic components like microring resonators. Propagation loss and weight application reduce optical signal power, while noise floors necessitate minimum power. Moreover, non-linear effects at high intensities cause signal distortion. Combined with limited power, these factors restrict the bit precision, with 4-6 bits typical in prior architectures.
[0005] To reach sufficient precision for an artificial intelligence system, optical accelerators often split digitally encoded values into shorter analog bit slices. Bit serial processing then recombines sliced results. However, throughput substantially declines due to the need for sequential slice-wise operations. Bit slicing also requires area-inefficient duplicate weighting structures. Overall efficiency suffers from the difficulty of precision scaling. While certain accelerators have adopted digital optics, they incur high area overheads rendering them
impractical. So both digital and analog approaches face efficiency obstacles, reflected across optical matrix multiply architectures.
[0006] Changing the representation and processing of data could alleviate these optical computing barriers. One alternative represents values as streams of random bits corresponding to probabilities between zero and one, known as stochastic computing. Logic operations between stochastic bit streams such as AND effectively multiply their probabilities. This permits arithmetic by leveraging simple optical logic like microring weighting. Moreover, encoding data as stochastic bit streams relaxes precision requirements. Since distinct levels denote zero or one bits, precision depends on stream length rather than analog range. Stochastic computing could thus improve achievable precision and efficiency.
[0007] Efficiently implementing reconfigurable optical gates poses challenges however. While optical AND operations suit stochastic multiplication, polymorphic gates enabling different logical functions would provide efficiency through area reuse. Previous optical logic designs suffer from substantially larger footprints than electronic gates though, so size reduction is desired to make implementations practical. Realizing area-efficient stochastic optical computing therefore, requires addressing this optimization across representation, logic, and accumulation.
SUMMARY
[0008] Systems and method directed to stochastic optical computing architecture with reconfigurable logic and accumulation are disclosed. In one embodiment, an optical logic gate includes a single microring resonator having an input port and an output port; one or more input terminals connected to apply logic signals to the microring resonator; and a tuning component to shift a resonance frequency of the microring resonator wherein a transmission between the input port and output port of the microring resonator performs a logic operation on signals applied to the one or more input terminals.
[0009] Some embodiments include a substrate for enabling in-memory stochastic to binary conversion. In one embodiment, the substrate comprises a plurality of bitline groups in a DRAM bank, each bitline group comprising a set of bitlines corresponding to an input stochastic bit vector; sense amplifiers connected to the bitlines; stochastic-to-analog converter circuits connected between the sense amplifiers and analog lines of each bitline group; analog- to-unary converter circuits connected between the analog lines and the bitlines of each bitline group; and unary -to-binary converter circuits connected to the bitlines of each bitline group.
[0010] Some embodiments include an optical logic gate comprising: a single microring resonator; input terminals connected to apply logic signals to the microring resonator; a tuning component to shift a resonance frequency of the microring resonator; and control logic configured to tune the resonance frequency to set a logic function of the optical logic gate by adjusting the tuning component.
[0011] Some embodiments include an accumulation circuit comprising: a photodetector configured to receive a plurality of optical signals over a plurality of wavelengths and convert the optical signals to electrical pulses; an integration circuit connected to integrate the electrical pulses from the photodetector over time to produce an output voltage level proportional to a count of electrical pulses; and an output terminal to provide the output voltage level as an accumulated result of the plurality of optical signals.
[0012] Some embodiments are directed to a computing system comprising: an optical source providing a plurality of optical wavelength channels; an array of optical logic gates, each optical logic gate operable to: receive a subset of the optical wavelength channels and two multi-bit operands, logically combine the multi-bit operands in a bitwise manner by selectively attenuating the subset of optical wavelength channels, and output the selectively attenuated subset as an optical output signal; accumulation circuits each connected to receive optical output signals from a group of the optical logic gates and configured to integrate the optical output signals over time to produce a voltage level representing a sum of the logically combined multi-bit operands; and analog-to-digital conversion circuitry connected to convert the voltage levels into digital results.
[0013] These and additional features provided by the embodiments of the present disclosure will be more fully understood in view of the following detailed description, in conjunction with the drawings.
DESCRIPTION OF THE DRAWINGS
[0014] The appended figures depict certain aspects and are therefore not to be considered limiting of the scope of this disclosure.
[0015] FIG. 1A depicts an optical computing architecture (e.g., aggregation, modulation, modulation (AAM) vector dot product core (VDPC)) for performing vector matrix multiplications in accordance with examples of the present disclosure.
[0016] FIG. IB depicts an optical computing architecture (e.g., modulation, aggregation, modulation (MAM) vector dot product core (VDPC)) for performing vector matrix multiplications in accordance with examples of the present disclosure.
[0017] FIG. 1C depicts details of a summation element in accordance with examples of the present disclosure.
[0018] FIG. 2 depicts details of a VDPC organization, in accordance with examples of the present disclosure.
[0019] FIG. 3A depicts details of an Optical Stochastic Multiplier (OSM), in accordance with examples of the present disclosure.
[0020] FIG. 3B depicts details of an Optical AND Gate (OAG), in accordance with examples of the present disclosure.
[0021] FIG. 3C depicts a cross-section view of an Optical AND Gate (OAG), in accordance with examples of the present disclosure.
[0022] FIGS. 3D-3I depict example transmission spectra for different logic-gate and complementary logic-gate functions corresponding to extracted transmission spectra for drop and through ports of an MRR, in accordance with examples of the present disclosure.
[0023] FIG. 4 depicts additional details of the photo charge accumulator, in accordance with examples of the present disclosure.
[0024] FIG. 5 illustrates a system-level implementation of a Stochastic Computing based Optical Neural Network Accelerator (SCONNA), in accordance with examples of the present disclosure.
[0025] FIGS. 5A-5H depict example timing diagrams, in accordance with examples of the present disclosure.
[0026] FIG. 6 depicts details of a XNOR-Bitcount based Binary Neural Network Accelerator (OXBNN), in accordance with examples of the present disclosure.
[0027] FIG. 6 depicts an example method for performing stochastic-to-binary conversion of input operands, in accordance with examples of the present disclosure.
[0028] FIG. 7 depicts details of an example PCA circuit, in accordance with examples of the present disclosure.
[0029] FIG. 8 depicts a system-level implementation of an OXBNN accelerator, in accordance with examples of the present disclosure.
[0030] To facilitate understanding, identical reference numerals have been used, where possible, to designate identical elements that are common to the drawings. It is contemplated that elements and features of one embodiment may be beneficially incorporated in other embodiments without further recitation.
DETAILED DESCRIPTION
[0031] Aspects of the present disclosure provide apparatuses, methods, processing systems, and computer-readable mediums for stochastic optical computing architecture with reconfigurable logic and accumulation.
[0032] Optical computing has emerged as a promising technology for accelerating workloads like artificial intelligence. However, existing optical computing architectures face substantial precision and efficiency challenges. Many prior optical accelerators rely on analog optics, restricting scalability due to the tradeoff between precision and parallel wavelength channels. This highlights the need for an optimized optical computing architecture tailored to overcome these limitations.
[0033] A technical problem includes that both digital and analog optical computing approaches face efficiency obstacles, reflected across optical matrix multiply architectures. Combined with limited power, these factors restrict the bit precision, with 4-6 bits typical in prior architectures.
[0034] Aspects of the present disclosure provide a stochastic computing based optical architecture with reconfigurable logic gates and specialized accumulation circuits. Such embodiments may take the form of a data center server and/or other computing architecture. The architecture encodes data as streams of random bits corresponding to probabilities between zero and one. Logic operations between stochastic bit streams effectively multiply their probabilities using simple optical logic. Moreover, encoding data as stochastic bit streams relaxes precision requirements. The architecture also includes polymorphic optical logic gates enabling different logical functions for improved efficiency through area reuse. In addition, the accumulation circuits integrate streams of optical bits over time to produce a voltage level result. Specifically, some embodiments may be configured utilizing the optical logic gates and the accumulation circuits to enable bitwise processing layers in an artificial intelligence system. This architecture provides an optimized solution customized for stochastic optical computing to help fully realize its acceleration potential.
[0035] FIG. 1A depicts an optical computing architecture (e.g., aggregation modulation modulation (AAM) vector dot product core (VDPC)) for performing vector matrix multiplications. An array 102 of laser diodes 104 emit light at a plurality of wavelengths i through N, generating N optical wavelength channels. An aggregation block 108 aggregates the generated optical wavelength channels into a single photonic waveguide through dense wavelength division multiplexing (DWDM) (using an Ax 1 multiplexer 106) and then splits the
optical power of these N wavelength channels equally into M separate waveguides (using a l x M splitter 110). A modulation block, also referred to as DIV block 112, employs M arrays of microring resonators (MRRs) 114 (one array per waveguide, with each array having N MRRs; each array referred to as DIV element) to imprint M DIVs of N points each onto the N'M wavelength channels by modulating the analog power amplitudes of the wavelength channels. Another modulation block, referred to as DKV block 116, employs another M arrays of MRRs (one array per waveguide, with each array having N MRRs; each array referred to as DKVelement) to further modulate the Vx M wavelength channels with DKVs, so that the analog power amplitudes of the individual wavelength channels then represent the point-wise products of the utilized DKVs 116 and DIVs 112. A summation block (SB) 120 employs M summation elements (SEs) 122, with each SE 122 having two balanced photodiodes (PDs) upon which the point- wise-product-modulated N wavelength channels are incident to produce an output current that is proportional to the result of the vector dot product (VDP) operation between the corresponding DKV 116 block and DIV 112. The array 102 of laser diodes 104 and SB 120 may be positioned at the two ends of the VDPC, with the aggregation, modulation (DIV block 112), and modulation (DKV block 116) blocks placed in between them. Based on the order in which these intermediate blocks (aggregation, modulation (DIV block 112), modulation (DKV block 116) blocks) are positioned between the array 102 of laser diodes 104 and the SB 120, the MRR-based VDPC organizations may be classified as a MAM (modulation, aggregation, modulation) or AMM (aggregation, modulation, modulation) VDPC organization.
[0036] That is, as depicted in FIG. 1A, the AMM VDPC organization positions the aggregation block 108 after the array 102 of laser diodes 104, and then the DIV block 112 followed by the DKV 116 modulation block. In contrast, the MAM VDPC in FIG. IB, positions the DIV block 124 first after the array 102 of laser diodes 104, and then positions the aggregation block 108 followed by the DKV block 116. Note that the MAM DIV block 124 is structurally different from the AMM DIV block 112. The MAM DIV block 124 employs one MRR 126 per waveguide, and as a result, it can imprint one DIV with N points onto the N wavelength channels. This one DIV (e.g., one array per waveguide, each waveguide having one MRR) is shared among all DKVs 118 in the MAM VDPC 130, whereas each DKV 118 can have a different DIV corresponding to it in the AMM VDPC 100. In examples, M generally is equal N for many MAM and AMM VDPC.
[0037] In both the AMM and MAM VDPC organizations, the combination of a DKV element and the corresponding SE are referred to as VDP element (VDPE). However, the size
and point-wise product precision of MRR-based VDPEs have certain limitations; such limitations can be addressed to improve MRR-based VDPCs, where stochastic computing can be implemented in a meaningful way.
[0038] FIG. 1C depicts details of a summation element 122 in accordance with examples of the present disclosure. The summation element 122 can include a pair of balanced photodiodes 152 and 154 (e.g., balanced photodiode pair) coupled to a trans-impedance amplifier 156. The trans impedance amplifier 156 receives an input current from the balanced photodiodes 152 and 154 and converts the input current to an output voltage.
[0039] FIG. 2 depicts a VDPC organization in accordance with examples of the present disclosure. Similar to the VDPCs of analog optical accelerators, the example VDPC depicted in FIG. 2 also implements multiple VDP operations in parallel. Thus, an array 202 of total V single-wavelength laser diodes (EDs) are used, with each LD sourcing optical power of
amount at a distinct wavelength Xi. The total power from all N EDs (Xi to XN) is multiplexed into a single photonic waveguide through wavelength division multiplexing (WDM) using the wavelength division multiplexor 204. These multiplexed wavelengths are split (using a splitter 206) into M input waveguide arms (IWAs). Each IWA receives N- wavelength optical power and guides it to a VDPE (e.g., 208). There are a total of M IWAs and M VDPEs (207) in the example depicted in FIG. 2. Each VDPE 207 includes at least three components: a cascade of N Optical Stochastic Multipliers (OSMs) 208A-208N; (ii) a bank of fdter MRRs (e.g., 214A- 214N); and (iii) a Photo-Charge Accumulator (PCA) pair. Each OSM (e.g., 208) performs stochastic multiplication between an input bitstream I (e.g., 210) (corresponding to a point in an V-point DIV) and weight bit-stream W (e.g., 212) (corresponding to a point in an V-point DKV). Each OSM 208 receives its bit-streams 7210 and from its corresponding peripherals at a supported bitrate (BR). Bitstream IP 212 provides a weight value along with a sign bit. Bitstream I 210 provides an RELU-activated output value from the previous CNN layer, without a sign bit as RELU has a nonnegative output. The detailed design of OSMs and their peripherals is further explained below. Each OSM 208 performs a bitwise logical AND operation between the I 210 and IT 212 bit-streams to produce a resultant optical bit-stream that represents the stochastic multiplication between the 7210 and W 220 bit-streams. The N optical bit-streams from the cascade of N OSMs 208, with each bit-stream carrying a stochastic multiplication result, reach the bank of fdter MRRs 214. In this bank of fdter MRRs 214, each fdter MRR (e.g., 214A, 214B, 214N) operates on a distinct optical bit-stream Xi. Each fdter MRR (e.g., 214A, 214B, 214N) receives the sign bit from the peripheral W 212 of its
corresponding OSM 208. The sign bit operates the filter to steer the incoming optical bit-stream Xi to the output waveguide arm OWA (if the sign bit is ‘0’) or OWA’ (if the sign bit is ‘ 1’). Thus, the OWA and OWA’ of a VDPE guide the optical bit-streams, carrying the stochastic multiplication results, to PC As 216A/216B. A PC A 216 can include a circuit that collects the optical bit-streams (i.e., stochastic multiplication results) from its corresponding OWA (or OWA’) and generates the accumulation result in a binary format as further described below. As an example, the accumulated result represents an inner product between optical vector inputs. In a VDPE, the OWA coupled PCA (e.g., 216A) combines with the corresponding OWA’ coupled PCA (e.g., 216B) to generate a signed accumulation result.
[0040] As depicted in FIG. 3A, an OSM 300 can include peripherals 302 and an optical logic gate, such as an optical ‘AND’ Gate (OAG) 304. In some embodiments, the optical logic gate maintains operating parameters while reconfigured between logic functions. Similarly, in some embodiments of the optical logic gate, the operating parameters include optical power, bit precision, bit rate, and footprint. The peripherals 302 convert a binary input value 4 (e.g., 306) and binary weight value H'h (308) retrieved from a buffer 310 into unipolar stochastic bitstreams I and W. The OAG 304 performs a multiplication-equivalent bitwise AND operation between the stochastic bit- streams I and W. The peripherals 302 of the OSM 300 can use a lookup table 312 and serializers 314A/314B to generate a combination of unipolar stochastic bit-streams 1314A and W 314B.
[0041] Two unipolar stochastic bitstreams 316A and 316B, for their eventual error-free multiplication using an AND gate, can be generated in combination with each other, so that they are uncorrelated, i.e., the marginal probability of one bit-stream (i.e., / 316 A or W 316B) is equal to its conditional probability given the other bit-stream (i.e., I given W or W given I). For the OSM 300, all combinations of uncorrelated bit-streams I and W can be generated a priori (offline) using a unipolar circuit, and then stored as bit-streams in the bit-vector (bitparallel) format in the lookup table 316. As a result, each entry in the lookup table 316 stores a combination of uncorrelated bit- vectors Iv and Wv. To index into the lookup table 316, the OSM 300 utilizes a unique identifier for each combination of binary values lb 306 and Wb 308 (that are accessed from a buffer 310, such as a scratchpad memory by performing an XOR- based hash function Ib®Wb (e.g.,). Thus, the OSM 300 uses a Ib®Wb value to fetch the desired combination of Iv and Wv from the lookup table 316. Then, the OSM 300 pushes the Iv and Wv through dedicated high-speed serializers 320A/320B, to generate bit-streams 1314A
and W 314B. As one example, if precision B=8-bit for binary lb and Wb, there would be 2xB entries in the lookup table 316, with each entry storing two 2B-bits long bit- vectors.
[0042] The stochastic bit-streams I and W, generated by the peripherals 302 of the OSM 300, are then fed to the OAG 304 via highspeed drivers 322A/322B for stochastic multiplication. Design aspects of the OAG 304 is further described in FIG. 3B. The OAG 304 can be an add-drop microring resonator (MRR) 326 having two operand terminals (realized as embedded PN-junctions 324A/324B) that can take two stochastic bit-streams I 314A and W 314B as inputs at a predefined bitrate (BR).
[0043] The MRR’s temperature can be increased using the integrated microheater 328 to consequently tune its operand-independent resonance from its fabrication-defined initial position y to its programmed position r|, relative to the input optical wavelength position Am. For each bit combination at the operand terminals ((/,JT) = (0,1), (1,0), or (1,1)), the MRR’s 326 resonance passband electro-refractively moves to an operand-driven position. Based on the MRR 326 resonance passband’s programmed position q relative to kin, the drop-port 330 transmission (T(kin)) of the MRR 326 provides bit-wise logical AND operation between the inputs 1314A and W 314B. The passbands of the MRR 326 for different operand inputs and temperature conditions are further depicted in FIGS. 3D-3I.
[0044] FIG. 3B depicts additional details of a microring (MRR) 332 based optical stochastic multiplier (OSM) 304A in accordance with examples of the present disclosure. More specifically, the MRR 332 can accept both input operands electrically, and its drop-port (through-port) optical response can be thermo-optically programmed to make it dynamically follow the truth table of different logic functions, such as AND, OR and XOR (NAND, NOR, and XNOR), at different times. In some embodiments, the control logic configures the optical logic gate to perform XOR logic by controlling the tuning component to shift the resonance frequency to a minimum transmission position relative to an optical input wavelength. Consequently, E-0 circuits built using the MRR 332 can address the above-described shortcomings by providing (1) the ability of all-electrical application of the input operands, (2) compactness through a single-MRR structure of the E-0 gate, and (3) high flexibility through the introduced polymorphism, and consequently, low idle time and improved suitability for use with SIMD/MIMD/SA based architectures. The MRR 332 is similar to an add-drop MRR, but includes four quarter-sized phase-shifting sections embedded therein. In examples, two quarter-sized sections of the MRR 332 are two PN junctions 324A/324B which are operated in the forward bias condition, whereas the remaining two quarter-sized sections integrate micro-
heaters 328A/328B. A cross-section 334 of a PN-junction based section of the MRR 332 is depicted in FIG. 3C and includes a ridge waveguide 336 with an embedded lateral PN junction 324B, fabricated on a buried oxide layer 338. Example dimensions of the P-type regions 340, 342, and 344 and N-type regions, 346, 348, and 350 and corresponding example carrier concentrations are also depicted in FIG. 3C.
[0045] The PN junction based sections (e.g., 324A/324B) of the MRR 332 work at the input terminals where the input logic signals/operand bits are applied. On the other hand, the microheaters 328A/328B integrated sections of the MRR 332 work as programming terminals that are used to program the MRR to perform specific logic-gate functions. Applying a voltage to the microheaters 328A/328B based programming terminals can increase the temperature of the MRR 332, which in turn can shift (red shift) the resonance of the MRR 332 towards a longer wavelength. This is because of the thermo-optic effect in silicon. To program or otherwise configure the MRR 332 to implement a specific logic-gate function, the operand-independent MRR resonance (i.e., the programmed MRR resonance) is adjusted to a specific spectral position with respect to the input optical wavelength, by applying a voltage to the programming terminals (e.g., at 328A/328B). Then, the electrical input logic signals or input operand bits (I 314A and W 314B) are applied to the PN junctions (e.g., 324A/324B) based input terminals of the MRR 332. Upon doing so, the resonance of the MRR 332 shifts (blue shifts) towards a shorter wavelength depending on the combination of the applied input operand bits. This is because of the free-carrier plasma dispersion effect in silicon. Applying the input operand bits to the input terminals makes the through-port 352 and drop-port 354 optical responses of the MRR 332 follow a truth-table of a logic-gate function for which the MRR 332 is programmed. In this manner, the MRR 332 can perform different logic-gate functions at different times. At any given time, the through-port 352 optical response of the MRR 332 follows logical complement of the drop-port 354 optical response. Therefore, AND, OR and XOR functions can be realized (one function at a time) at the drop-port 354 of the MRR 332. Concurrently, the through-port 352 of the MRR 332 can provide complementary logic gate functions such as NAND, NOR and XNOR as discussed below.
[0046] FIGS. 3D-3I depict example transmission spectra for different logic-gate and complementary logic-gate functions corresponding to extracted transmission spectra for drop and through ports of an MRR 332. More specifically, the extracted transmission spectra for different values of the detuning of the operand-independent MRR resonance position K with respect to the input wavelength kin are depicted in FIGS. 3D-3I. As previously described, the
detuning values correspond to different logic-gate functions that the MRR 332 can perform. In addition, transmission spectra for different combinations of the input operand bits are also depicted in FIGS. 3D-3I.
[0047] Transmission spectra corresponding to logic-gate functions AND, OR, and XOR are depicted in FIGS. 3D, 3F, and 3H, respectively. These transmission spectra are drop-port transmission spectra (Torentzian lineshape passbands). Similarly, the transmission spectra corresponding to complementary logic-gate functions NAND, NOR, and XNOR are depicted in in FIGS. 3E, 3G, and 31, respectively. These transmission spectra are through-port transmission spectra (inverse Torentzian lineshape passbands). As depicted in FIG. 3D, the drop port and through port transmission exhibits two clearly distinguishable levels. The full transmission range at the drop port and through port of the MRR 332 is divided into two areas, in which the lower part of the full transmission range is indicated as logic 0, whereas the upper part is indicated as logic 1. If the drop port (DT(Xin)) and through port (TT(Xin)) transmission at Xin falls in the lower part of the full transmission range, then it is referred to as logic ‘0’ transmission. On the other hand, if the drop port and through port transmission at Xm falls in the upper part of the full transmission range, then it is referred to as logic ‘ 1 ’ transmission. However, the vertical spans of the two distinguishable transmission levels differ between the drop port and through port. This is because, similar to the transmission spectra, the spans of transmission levels at the drop port also complement the spans of transmission levels at the through port. The difference between the minimum supported logic ‘ 1 ’ transmission and the maximum supported logic ‘0’ transmission is the sensitivity of optical modulation amplitude (SOMA). SOMA is a property of the photodetector based receiver circuit, and it affects the performance of the MRR 332.
[0048] As an example AND function, as shown in FIG. 3D, to program or configure the MRR 332 to implement AND function, an example 0.9 V voltage (3.52 mW power) is applied to the programming terminals of the MRR 332, shifting the resonance from the initial position, r|, to the programmed position, K, where K has the programmed detuning of about 0.7 nm with respect to Xin. Then, the input operand bits I and W are applied to the input terminals of the device. Doing so induces a blueshift in the MRR resonance, the magnitude of which depends on the specific combination of the applied input operand bits (Zand W, as shown in FIG. 3D). If the applied bit-combination (I, W) is (0,0), the resonance position of the MRR stays at K and the drop port transmission at Xin provides logic ‘0’ level (the bottom dot on the Y-axis). If the applied bit-combination (7, W) is (0,1) or (1,0), the position of the MRR resonance changes, but
the blueshift is the same for both (0,1) and (1,0) bit combinations, and the drop port transmission at Am still remains at logic ‘0’ level (the top dot on the Y-axis). On the other hand, if the applied bit-combination (7, IF) is (1,1), the MRR resonance undergoes a larger blueshift, and the position of the passband with respect to Am changes. As a result, the drop port transmission at kin changes to logic ‘ 1’ level (the dot on the Y-axis). Hence, the drop port transmission at Am for the MRR 332 changes with the applied input operand bits, and follows the truth table of the AND logic function (see the truth table in FIG. 3D). In some embodiments, shifting the resonance frequency includes performing AND logic on the signals applied to the one or more input terminals when the resonance frequency provides a peak transmission at the predefined optical input wavelength. As discussed earlier, since the through-port response provides a logical complement to the drop-port response, this AND function at the drop port of the MRR 332 corresponds to NAND function at the through port of the MRR 332 as illustrated in FIG. 3D).
[0049] Similarly, the MRR 332 can be reconfigured to implement OR (NOR) and XOR (XNOR) gate functions as well, by applying a suitable voltage to the programming terminals of the MRR 332 to set the relative position of K with respect to in as shown in FIGS. 3E, 3G, and 31.
[0050] FIG. 4 depicts additional details of the photo charge accumulator 216 of FIG. 2 in accordance with examples of the present disclosure. In examples, the stochastic multiplication bit-streams generated by OSMs 208 (FIG. 2) are guided to a PCA 216, where they are accumulated to generate a binary output value equivalent to the VDP result. The PCA 216 may include has two stages: (i) a stochastic-to-analog conversion stage; and (ii) an analog-to-binary conversion stage. The stochastic-to-analog stage employs a photodetector 402 and two TIR circuits 404A and 404B, whereon one TIR circuit remains redundant, enabled by the demux 406 and mux 408. The photodetector 402 generates a current pulse for each optical logic ‘ 1’ incident upon it. This current pulse accumulates a certain amount of charge on the capacitor of the active TIR circuit (e.g., the circuit with Cl capacitor); as a result, the capacitor accrues an analog voltage level. Hence, when one or more output optical bit-streams are incident upon the photodetector 402, the total accumulated charge (and thus, the accrued analog voltage level) on the active capacitor (e.g., Cl) is proportional to the total number of 1’s in the incident bitstreams. The number of 1’s that can be accumulated in such manner might be limited, as the charge across the capacitor of TIR circuit 404B can saturate. Once the TIR output saturates, a discharge of the active capacitor (e.g., Cl) is needed to prepare the circuit for the next
accumulation phase. While capacitor Cl is discharging, capacitor C2 of the redundant TIR circuit (e.g., 404A) mitigates the discharge latency by allowing continuation of a concurrent accumulation phase. The output analog voltage computed by the stochastic-to-analog conversion stage represents the unipolar unsealed addition of the stochastic bit-streams. To convert the analog voltage into a binary value, the analog-to-binary stage of the PCA circuit employs an analog-to-digital converter ADC 410. This binary value is the VDP result.
[0051] FIG. 5 illustrates a system-level implementation of a Stochastic Computing based Optical Neural Network Accelerator (SCONNA) 500 in accordance with examples of the present disclosure. More specifically, the SCONNA 500 may include a global memory 502 for storing CNN parameters, and a preprocessing and mapping unit 504 for decomposing the tensors into DIVs/DKVs and mapping them onto VDPEs. The SCONNA 500 can include a mesh of tiles 506A-506N coupled to routers 510, and this mesh network facilitates parameter communication among tiles 506A-506N with network interfaces 508. Each tile 506 can include one or more SCONNA VDPCs 514A-514N interconnected (via H-tree network) with output buffer 522, activation 518, and pooling 520 units. In addition, each tile 506 can include a sum reduction network 516 configured to accumulate a number of partial sum values.
[0052] FIG. 6 depicts details of a XNOR-Bitcount based Binary Neural Network Accelerator (OXBNN) in accordance with examples of the present disclosure. Binary Neural Networks (BNNs) are increasingly preferred over full-precision Convolutional Neural Networks (CNNs) to reduce the memory and computational requirements of inference processing with minimal accuracy drop. BNNs convert CNN model parameters to 1-bit precision, allowing inference of BNNs to be processed with simple XNOR and bitcount operations. This makes BNNs amenable to hardware acceleration. Although existing photonic integrated circuits (PICs) based BNN accelerators provide remarkably higher throughput and energy efficiency than their electronic counterparts, the utilized XNOR and bitcount circuits in these accelerators can be further enhanced to improve area usage, energy efficiency, and throughput. As depicted in FIG. 6, the OXBNN architecture includes an XNOR-Bitcount Processing Core (XPC) 600, which may include an array 602 of total N single-wavelength laser diodes (EDs), with each LD sourcing optical power of P amount at a distinct wavelength
The total power from all N EDs (at wavelength to Aw) is multiplexed into a single photonic waveguide through wavelength division multiplexing (WDM) (e.g., at 604). The optical power containing all these A wavelengths is split (e.g., at splitter 606) into M input waveguides, each
of which connects to an XNOR-Bitcount Processing Element (XPE) 620. An XPC 600 can include a total of A/XPEs 620.
[0053] As depicted in FIG. 6, an XPE 620 in the OXBNN architecture can include: (i) an array 608 of a total of N Optical XNOR Gates (OXGs) (e.g., 616A-616N) that generates an XNOR vector (or an XNOR vector slice) containing N optical bits, and (ii) a Photo-Charge Accumulator (PCA) 612A that performs a bitcount on the generated XNOR vector (or XNOR vector slice). The value A here, which is equal to the number of wavelengths and number of OXGs 616 per XPE 620, is referred to as the size of the XPE 620.
[0054] In an XPE 620, an array of a total of N OXGs 616 couples to an input waveguide (e.g., 601A, 601B, 601N) as depicted in FIG. 6. Each OXG 616 operates upon a unique wavelength traversing the input waveguide (e.g., 601). Each OXG 616 in the array electrically receives two binary operands (i.e., input bit
and weight bit w ) from its corresponding drivers (not shown). The array of OXGs performs a bit-wise logical XNOR between an X-bit input vector slice
= {ij, i. , ... , i ] and an X-bit weight vector slice 14^ = {vtq1, w{, ... , w™} to produce a resultant N-bit XNOR vector slice. Each OXG 616 in the array produces one bit of the resultant XNOR vector slice, and it imprints this bit on its corresponding Ai (by modulating the optical transmission at - to be consequently guided to the bitcount circuit (i.e., PCA 612) via the output waveguide 610A-610N.
[0055] As a result, the PCA 612 receives the N individual optical bits of the X-bit XNOR vector slice concurrently on X distinct wavelengths. The PCA 612 performs bitcount on these optical bits. This processing step, from the bit parallel application of the binary input and weight vector slices at the electrical input terminals of the array of N OXGs 616 to the generation of the bitcount result by the PCA 612, occur with low latency because of the lightspeed operation of the XPE 620. This processing step mapped on an XPE 620 can be referred to as a PASS and the corresponding latency as r. Thus, the XPE 620 can produce one bitcount result for one XNOR vector slice in every single PASS with T latency. Since T can be very low (as low as 20 ps for example), the XPE 620 can achieve very high processing throughput by completing one PASS every T period. For a pass, multiple input and weight vector slices {I- , 12 , ■■■ , I a } and {Wi , VP2 , ... , Wa can be applied to the array of OXGs 616 of an XPE 620
1 in a serial manner at the predefined data rate (DR) of-.
[0056] The Optical XNOR Gate (OXG) 616 depicted in FIG. 6 includes an add-drop microring resonator (MRR), which has two operand terminals (realized as embedded PN-
junctions) that can take two operand bits z and w as inputs for a predefined time-width (usually a little less than the T period). FIG. 3H depicts the passbands of the MRR for different operand inputs and temperature conditions. The MRR’s temperature can be increased using the integrated microheater 328 (e.g., FIG. 3A), to consequently tune its operand-independent resonance from its fabrication-defined initial position operand-independent MRR resonance position r| to its programmed position K relative to the input optical wavelength kin. For each bit combination at the operand terminals ((7,W) = (0,1), (1,0), or (1,1)), the MRR’s resonance passband electro-refractively moves to an operand-driven position. Based on the MRR resonance passband’s programmed position K relative to kin, the through-port transmission (T(kin)) of the MRR provides bit-wise logical XNOR operation between the input bits z and w. [0057] The XNOR vector bits generated by an array of OXGs are guided to a PCA circuit (e.g., PCA circuit 702 of FIG. 7), where a bitcount is performed on the XNOR vector bits to generate an output result. The PCA circuit 702 employs a photodetector 704 and two time integrating receiver (TIR) circuits 706A and 706B, where one of the TIR1 and TIR2 circuits remains redundant, enabled by the demux 708 and mux 710. The photodetector 704 generates a current pulse for each optical logic ‘ 1 ’ incident upon it. The amplitude of a current pulse generated for an optical logic ‘0’ remains under the noise limit; therefore, a logic ‘0’ remains statistically undetected. The current pulse generated by an optical logic ‘ 1’ accumulates a certain statistically significant amount of charge on the capacitor of the active TIR circuit (e.g., the circuit with Cl capacitor); as a result, the TIR circuit 706B outputs a detectable analog voltage level. Hence, when more optical ‘ 1 ’s are incident upon the photodetector 704, the total accumulated charge on the active capacitor (e.g., Cl), and thus, the accrued output analog voltage level, grows proportionally to the total number of optical ‘ 1’s that are incident. This is because current source (a sequence of current pulses) can charge a capacitor linearly following
this equation: 57 = where z is an incident current pulse, 6t is the time-width of the current
pulse, C is the capacitance, and bVis the accrued voltage. The final analog voltage accrued at the TIR output, thus, represents the bitcount result (accumulation result) of the incident optical cl’s.
[0058] However, the number of ‘ l’s that can be accumulated in such a manner might be limited, as the output of the TIR circuit 706B might saturate. Once the output of a TIR circuit 706B saturates, the ongoing accumulation phase ends and the bitcount result (i.e., the final TIR output voltage) is passed through a comparator 712 to generate the activation value for a next
BNN layer. After one accumulation phase, a discharge of the active capacitor (e.g., Cl) is needed to prepare the circuit for the next accumulation phase. While capacitor Cl is discharging, the redundant TIR2 circuit 706 A with capacitor C2 mitigates the discharge latency by allowing a continuation of a concurrent bitcount.
[0059] FIG. 8 depicts a system-level implementation 800 of an OXBNN accelerator. The implementation 800 includes a global memory 802 that stores BNN parameters and a preprocessing and mapping unit 804. The implementation 800 includes a mesh network of tiles 806A-806N. Each tile 806 includes a network interface 808A-N and, as an example, 4 XPCs 814A-814N interconnected (via H-tree) with an output buffer 822 and pooling unit(s) 820.
[0060] Accordingly, embodiments provided herein may include one or more aspects. A first aspect includes an optical logic gate comprising: a microring resonator having an input port and an output port; one or more input terminals connected to apply logic signals to the microring resonator; and a tuning component to shift a resonance frequency of the microring resonator wherein a transmission between the input port and the output port of the microring resonator performs a logic operation on signals applied to the one or more input terminals.
[0061] A second aspect includes the first aspect, wherein shifting the resonance frequency sets a transmission level of the output port relative to a predefined optical input wavelength to the input port.
[0062] A third aspect includes the first aspect and/or the second aspect, wherein shifting the resonance frequency includes performing AND logic on the signals applied to the one or more input terminals when the resonance frequency provides a peak transmission at the predefined optical input wavelength.
[0063] A fourth aspect includes any of the first aspect through the third aspect, wherein shifting the resonance frequency includes performing OR logic on the signals applied to the one or more input terminals when the resonance frequency provides an intermediate transmission level at the predefined optical input wavelength.
[0064] A fifth aspect includes any of the first aspect through the fourth aspect, wherein shifting the resonance frequency includes performing XNOR logic on the signals applied to the one or more input terminals when the resonance frequency provides minimum transmission at the predefined optical input wavelength.
[0065] A sixth aspect includes any of the first aspect through the fifth aspect, wherein the one or more input terminals comprise embedded PN junctions for carrier density modulation to shift the resonance frequency.
[0066] A seventh aspect includes any of the first aspect through the sixth aspect, wherein the tuning component to shift the resonance frequency comprises a microheater disposed adjacent to the microring resonator.
[0067] An eighth aspect includes any of the first aspect through the seventh aspect, wherein the microring resonator, the one or more input terminals, and the tuning component are integrated on a silicon photonic integrated circuit.
[0068] A ninth aspect includes any of the first aspect through the eighth aspect, wherein the one or more input terminals modulate a refractive index of the microring resonator based on applied logic signals.
[0069] A tenth aspect includes any of the first aspect through the ninth aspect, wherein the logic operation on the signals applied to the one or more input terminals comprises AND, OR, XOR, NAND, NOR, or XNOR logic.
[0070] An eleventh aspect includes an optical logic gate comprising: a microring resonator; input terminals connected to apply logic signals to the microring resonator; a tuning component to shift a resonance frequency of the microring resonator; and control logic configured to tune the resonance frequency to set a logic function of the optical logic gate by adjusting the tuning component.
[0071] A twelfth aspect includes the eleventh aspect, wherein the control logic configures the optical logic gate to perform AND logic by controlling the tuning component to shift the resonance frequency to a peak transmission position relative to an optical input wavelength.
[0072] A thirteenth aspect includes the eleventh aspect and/or the twelfth aspect, wherein the control logic configures the optical logic gate to perform OR logic by controlling the tuning component to shift the resonance frequency to an intermediate transmission position relative to an optical input wavelength.
[0073] A fourteenth aspect includes any of the eleventh aspect through the thirteenth aspect, wherein the control logic configures the optical logic gate to perform XOR logic by controlling the tuning component to shift the resonance frequency to a minimum transmission position relative to an optical input wavelength.
[0074] A fifteenth aspect includes any of the eleventh aspect through the fourteenth aspect, wherein the control logic dynamically reconfigures the logic function over time by adjusting the tuning component to shift the resonance frequency of the microring resonator.
[0075] A sixteenth aspect includes any of the eleventh aspect through the fifteenth aspect, wherein dynamic reconfiguration enables time-division multiplexing of logical operations.
[0076] A seventeenth aspect includes any of the eleventh aspect through the sixteenth aspect, wherein the tuning component comprises microheaters to thermally tune a refractive index of the microring resonator.
[0077] An eighteenth aspect includes any of the eleventh aspect through the seventeenth aspect, wherein the control logic controls electrical power supplied to the microheaters to shift the resonance frequency by a predefined detuning amount associated with a target logic function.
[0078] A nineteenth aspect includes any of the eleventh aspect through the eighteenth aspect, wherein the input terminals comprise PN junctions.
[0079] A twentieth aspect includes any of the eleventh aspect through the nineteenth aspect, wherein the optical logic gate maintains operating parameters while reconfigured between logic functions.
[0080] A twenty-first aspect includes any of the eleventh aspect through the twentieth aspect, wherein the operating parameters include optical power, bit precision, bit rate, and footprint.
[0081] A twenty-second aspect includes any of the eleventh aspect through the twenty-first aspect, wherein the control logic receives commands identifying a target logic function and controls the tuning component to configure the optical logic gate with the target logic function. [0082] A twenty -third aspect includes any of the eleventh aspect through the twenty-second aspect, wherein the input terminals modulate a refractive index of the microring resonator based on applied logic signals.
[0083] A twenty-fourth aspect includes any of the eleventh aspect through the twenty -third aspect, wherein the control logic and the tuning component are integrated with the microring resonator on a photonic integrated circuit.
[0084] A twenty-fifth aspect that includes an accumulation circuit comprising: a photodetector configured to receive a plurality of optical signals over a plurality of wavelengths and convert the plurality of optical signals to electrical pulses; an integration circuit connected to integrate the electrical pulses from the photodetector over time to produce an output voltage level proportional to a count of electrical pulses; and an output terminal to provide the output voltage level as an accumulated result of the plurality of optical signals.
[0085] A twenty-sixth aspect includes the twenty-fifth aspect, wherein the output voltage level represents a sum of incident optical pulse amplitudes within the plurality of optical signals.
[0086] A twenty-seventh aspect includes the twenty-fifth aspect and/or the twenty-sixth aspect, wherein the integration circuit comprises one or more capacitors charged by the electrical pulses.
[0087] A twenty-eighth aspect includes any of the twenty-fifth aspect through the twentyseventh aspect, further comprising switching logic to alternate integration across the one or more capacitors.
[0088] A twenty -ninth aspect includes any of the twenty-fifth aspect through the twentyeighth aspect, wherein the photodetector concurrently receives the plurality of optical signals without wavelength filtering.
[0089] A thirtieth aspect includes any of the twenty-fifth aspect through the twenty-ninth aspect, wherein the plurality of optical signals encode probabilistic values as streams of optical bits.
[0090] A thirty-first aspect includes any of the twenty-fifth aspect through the thirtieth aspect, wherein the output voltage level represents a multiplication of the probabilistic values encoded in the plurality of optical signals when computed as a function of optical bit logical products.
[0091] A thirty-second aspect includes any of the twenty-fifth aspect through the thirty- first aspect, wherein the output voltage level saturates upon receiving a threshold number of optical pulses.
[0092] A thirty-third aspect includes any of the twenty-fifth aspect through the thirty- second aspect, further comprising a reset switch connected across integration capacitors to enable discharge upon saturation.
[0093] A thirty-fourth aspect includes any of the twenty-fifth aspect through the thirty- third aspect, wherein the photodetector, the integration circuit and the output terminal are integrated on an integrated circuit.
[0094] A thirty-fifth aspect includes any of the twenty-fifth aspect through the thirty-fourth aspect, wherein the photodetector comprises a balanced photodiode pair connected to the integration circuit.
[0095] A thirty-sixth aspect includes any of the twenty-fifth aspect through the thirty-fifth aspect, further comprising a transimpedance amplifier coupling the photodetector to the integration circuit.
[0096] A thirty-seventh aspect includes any of the twenty-fifth aspect through the thirtysixth aspect, wherein the accumulated result represents an inner product between optical vector inputs.
[0097] A thirty-eighth aspect includes any of the twenty-fifth aspect through the thirtyseventh aspect, wherein the plurality of optical signals are received over at least 50 wavelength channels with at least 0.7 nm channel spacing.
[0098] A thirty-ninth aspect includes a computing system comprising: an optical source providing a plurality of optical wavelength channels; an array of optical logic gates, each optical logic gate operable to: receive a subset of the plurality of optical wavelength channels and two multi-bit operands, logically combine the two multi-bit operands in a bitwise manner by selectively attenuating the subset of the plurality of optical wavelength channels, and output the selectively attenuated subset as an optical output signal; accumulation circuits each connected to receive optical output signals from a group of the array of optical logic gates and configured to integrate the optical output signals over time to produce a voltage level representing a sum of the logically combined multi-bit operands; and analog-to-digital conversion circuitry connected to convert the voltage level into digital results.
[0099] A fortieth aspect includes the thirty -ninth aspect, wherein the array of optical logic gates each include a single microring resonator structure.
[0100] A forty-first aspect includes the thirty -ninth aspect and/or the fortieth aspect, wherein the array of optical logic gates are dynamically reconfigurable to selectively perform AND, OR, XOR, NAND, NOR, or XNOR logical operations on the two multi-bit operands.
[0101] A forty-second aspect includes any of the thirty -ninth aspect through the forty-first aspect, wherein the two multi-bit operands are provided to the array of optical logic gates in a stochastic bit stream format encoding probabilistic values between 0 and 1.
[0102] A forty -third aspect includes any of the thirty-ninth aspect through the forty-second aspect, wherein a length of the stochastic bit streams sets a precision of probabilistic values.
[0103] A forty-fourth aspect includes any of the thirty-ninth aspect through the forty -third aspect, wherein the accumulation circuits are each configured to concurrently receive the optical output signals over multiple wavelengths without wavelength filtering.
[0104] A forty-fifth aspect includes any of the thirty -ninth aspect through the forty-fourth aspect, further comprising control logic connected to reconfigure functions of both the array of optical logic gates and the accumulation circuits over time.
Claims
1. An optical logic gate comprising: a microring resonator having an input port and an output port; one or more input terminals connected to apply logic signals to the microring resonator; and a tuning component to shift a resonance frequency of the microring resonator wherein a transmission between the input port and the output port of the microring resonator performs a logic operation on signals applied to the one or more input terminals.
2. The optical logic gate of claim 1, wherein shifting the resonance frequency sets a transmission level of the output port relative to a predefined optical input wavelength to the input port.
3. The optical logic gate of claim 2, wherein shifting the resonance frequency includes performing AND logic on the signals applied to the one or more input terminals when the resonance frequency provides a peak transmission at the predefined optical input wavelength.
4. The optical logic gate of claim 2, wherein shifting the resonance frequency includes performing OR logic on the signals applied to the one or more input terminals when the resonance frequency provides an intermediate transmission level at the predefined optical input wavelength.
5. The optical logic gate of claim 2, wherein shifting the resonance frequency includes performing XNOR logic on the signals applied to the one or more input terminals when the resonance frequency provides minimum transmission at the predefined optical input wavelength.
6. The optical logic gate of claim 1, wherein the one or more input terminals comprise embedded PN junctions for carrier density modulation to shift the resonance frequency.
7. The optical logic gate of claim 1, wherein the tuning component to shift the resonance frequency comprises a microheater disposed adjacent to the microring resonator.
8. The optical logic gate of claim 1, wherein the microring resonator, the one or more input terminals, and the tuning component are integrated on a silicon photonic integrated circuit.
9. The optical logic gate of claim 1, wherein the one or more input terminals modulate a refractive index of the microring resonator based on applied logic signals.
10. The optical logic gate of claim 1, wherein the logic operation on the signals applied to the one or more input terminals comprises AND, OR, XOR, NAND, NOR, or XNOR logic.
11. An optical logic gate comprising: a microring resonator; input terminals connected to apply logic signals to the microring resonator; a tuning component to shift a resonance frequency of the microring resonator; and control logic configured to tune the resonance frequency to set a logic function of the optical logic gate by adjusting the tuning component.
12. The optical logic gate of claim 11, wherein the control logic configures the optical logic gate to perform AND logic by controlling the tuning component to shift the resonance frequency to a peak transmission position relative to an optical input wavelength.
13. The optical logic gate of claim 11, wherein the control logic configures the optical logic gate to perform OR logic by controlling the tuning component to shift the resonance frequency to an intermediate transmission position relative to an optical input wavelength.
14. The optical logic gate of claim 11, wherein the control logic configures the optical logic gate to perform XOR logic by controlling the tuning component to shift the resonance frequency to a minimum transmission position relative to an optical input wavelength.
15. The optical logic gate of claim 11, wherein the control logic dynamically reconfigures the logic function over time by adjusting the tuning component to shift the resonance frequency of the microring resonator.
16. The optical logic gate of claim 15, wherein dynamic reconfiguration enables time-division multiplexing of logical operations.
17. The optical logic gate of claim 11, wherein the tuning component comprises microheaters to thermally tune a refractive index of the microring resonator.
18. The optical logic gate of claim 17, wherein the control logic controls electrical power supplied to the microheaters to shift the resonance frequency by a predefined detuning amount associated with a target logic function.
19. The optical logic gate of claim 11, wherein the input terminals comprise PN junctions.
20. The optical logic gate of claim 11, wherein the optical logic gate maintains operating parameters while reconfigured between logic functions.
21. The optical logic gate of claim 20, wherein the operating parameters include optical power, bit precision, bit rate, and footprint.
22. The optical logic gate of claim 11, wherein the control logic receives commands identifying a target logic function and controls the tuning component to configure the optical logic gate with the target logic function.
23. The optical logic gate of claim 11, wherein the input terminals modulate a refractive index of the microring resonator based on applied logic signals.
24. The optical logic gate of claim 11, wherein the control logic and the tuning component are integrated with the microring resonator on a photonic integrated circuit.
25. An accumulation circuit comprising: a photodetector configured to receive a plurality of optical signals over a plurality of wavelengths and convert the plurality of optical signals to electrical pulses; an integration circuit connected to integrate the electrical pulses from the photodetector over time to produce an output voltage level proportional to a count of electrical pulses; and an output terminal to provide the output voltage level as an accumulated result of the plurality of optical signals.
26. The accumulation circuit of claim 25, wherein the output voltage level represents a sum of incident optical pulse amplitudes within the plurality of optical signals.
27. The accumulation circuit of claim 25, wherein the integration circuit comprises one or more capacitors charged by the electrical pulses.
28. The accumulation circuit of claim 27, further comprising switching logic to alternate integration across the one or more capacitors.
29. The accumulation circuit of claim 25, wherein the photodetector concurrently receives the plurality of optical signals without wavelength fdtering.
30. The accumulation circuit of claim 25, wherein the plurality of optical signals encode probabilistic values as streams of optical bits.
31. The accumulation circuit of claim 30, wherein the output voltage level represents a multiplication of the probabilistic values encoded in the plurality of optical signals when computed as a function of optical bit logical products.
32. The accumulation circuit of claim 25, wherein the output voltage level saturates upon receiving a threshold number of optical pulses.
33. The accumulation circuit of claim 25, further comprising a reset switch connected across integration capacitors to enable discharge upon saturation.
34. The accumulation circuit of claim 25, wherein the photodetector, the integration circuit and the output terminal are integrated on an integrated circuit.
35. The accumulation circuit of claim 25, wherein the photodetector comprises a balanced photodiode pair connected to the integration circuit.
36. The accumulation circuit of claim 25, further comprising a transimpedance amplifier coupling the photodetector to the integration circuit.
37. The accumulation circuit of claim 25, wherein the accumulated result represents an inner product between optical vector inputs.
38. The accumulation circuit of claim 25, wherein the plurality of optical signals are received over at least 50 wavelength channels with at least 0.7 nm channel spacing.
39. A computing system comprising: an optical source providing a plurality of optical wavelength channels; an array of optical logic gates, each optical logic gate operable to: receive a subset of the plurality of optical wavelength channels and two multi-bit operands, logically combine the two multi-bit operands in a bitwise manner by selectively attenuating the subset of the plurality of optical wavelength channels, and output the selectively attenuated subset as an optical output signal; accumulation circuits each connected to receive optical output signals from a group of the array of optical logic gates and configured to integrate the optical output signals over time to produce a voltage level representing a sum of the logically combined multi-bit operands; and analog-to-digital conversion circuitry connected to convert the voltage level into digital results.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202463560313P | 2024-03-01 | 2024-03-01 | |
| US63/560,313 | 2024-03-01 |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| WO2025188570A1 true WO2025188570A1 (en) | 2025-09-12 |
| WO2025188570A8 WO2025188570A8 (en) | 2025-10-02 |
Family
ID=96991404
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2025/017914 Pending WO2025188570A1 (en) | 2024-03-01 | 2025-02-28 | Stochastic optical computing architecture with reconfigurable logic and accumulation |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025188570A1 (en) |
Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20100182005A1 (en) * | 2009-01-20 | 2010-07-22 | Stephan Biber | Magnetic resonance tomography apparatus with a local coil and method to detect the position of the local coil |
| US20100328744A1 (en) * | 2006-08-24 | 2010-12-30 | Cornell Research Foundation, Inc. | Optical logic device |
| US20160178988A1 (en) * | 2014-12-22 | 2016-06-23 | The University Of Connecticut | Thyristor-Based Optical AND Gate and Thyristor-Based Electrical AND Gate |
| US20210373362A1 (en) * | 2019-02-12 | 2021-12-02 | The Trustees Of Columbia University In The City Of New York | Tunable Optical Frequency Comb Generator In Microresonators |
| US20220051123A1 (en) * | 2019-05-02 | 2022-02-17 | Quantum Machines | Modular and dynamic digital control in a quantum controller |
| US20230118909A1 (en) * | 2021-10-19 | 2023-04-20 | Hewlett Packard Enterprise Development Lp | Optical logic gate decision-making circuit combining non-linear materials on soi |
| US20230350270A1 (en) * | 2022-04-27 | 2023-11-02 | The Trustees Of Boston College | Optical logic circuit devices and methods thereof |
| US20230358952A1 (en) * | 2022-05-09 | 2023-11-09 | Intel Corporation | Reduced bridge structure for a photonic integrated circuit |
-
2025
- 2025-02-28 WO PCT/US2025/017914 patent/WO2025188570A1/en active Pending
Patent Citations (8)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20100328744A1 (en) * | 2006-08-24 | 2010-12-30 | Cornell Research Foundation, Inc. | Optical logic device |
| US20100182005A1 (en) * | 2009-01-20 | 2010-07-22 | Stephan Biber | Magnetic resonance tomography apparatus with a local coil and method to detect the position of the local coil |
| US20160178988A1 (en) * | 2014-12-22 | 2016-06-23 | The University Of Connecticut | Thyristor-Based Optical AND Gate and Thyristor-Based Electrical AND Gate |
| US20210373362A1 (en) * | 2019-02-12 | 2021-12-02 | The Trustees Of Columbia University In The City Of New York | Tunable Optical Frequency Comb Generator In Microresonators |
| US20220051123A1 (en) * | 2019-05-02 | 2022-02-17 | Quantum Machines | Modular and dynamic digital control in a quantum controller |
| US20230118909A1 (en) * | 2021-10-19 | 2023-04-20 | Hewlett Packard Enterprise Development Lp | Optical logic gate decision-making circuit combining non-linear materials on soi |
| US20230350270A1 (en) * | 2022-04-27 | 2023-11-02 | The Trustees Of Boston College | Optical logic circuit devices and methods thereof |
| US20230358952A1 (en) * | 2022-05-09 | 2023-11-09 | Intel Corporation | Reduced bridge structure for a photonic integrated circuit |
Also Published As
| Publication number | Publication date |
|---|---|
| WO2025188570A8 (en) | 2025-10-02 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US8392487B1 (en) | Programmable matrix processor | |
| CN111882052B (en) | Photon convolution neural network system | |
| WO2020176538A1 (en) | Hybrid analog-digital matrix processors | |
| US20240127040A1 (en) | Photonic accelerator for deep neural networks | |
| WO2022160784A1 (en) | Optical computing device and system and convolution computing method | |
| CA3250672A1 (en) | Photonic waveguide networks | |
| CN115906977B (en) | An all-optical convolver based on dual-comb | |
| CN119150939B (en) | Optical convolution computing device | |
| CN116888440A (en) | Optical computing methods and systems based on high-speed time-gated single-photon detector arrays | |
| CN119045123A (en) | Optical chip and optical neural network system based on micro-ring resonator on-chip modulation | |
| WO2023116496A1 (en) | Optical computing device, method and system | |
| CN111630446B (en) | Segmented digital-to-optical phase shift converter | |
| CN102147634A (en) | Optical vector-matrix multiplier based on single-waveguide coupling micro-ring resonant cavity | |
| WO2025188570A1 (en) | Stochastic optical computing architecture with reconfigurable logic and accumulation | |
| CN119312004B (en) | Reconfigurable optical convolution operation device and optical convolution operation method | |
| CN114565091B (en) | Optical neural network device, chip and optical implementation method for neural network calculation | |
| CN115358366A (en) | Acceleration system of convolution operation layer in optical neural network | |
| US5646395A (en) | Differential self-electrooptic effect device | |
| CN115640839A (en) | A storage and calculation integrated circuit, computing system and computing method | |
| CN119312860A (en) | Multi-layer optoelectronic convolutional neural network system and method based on optical parallel method | |
| CN118939077A (en) | A photon parallel matrix multiplication operation chip and its application system | |
| CN116911368A (en) | A photonic convolutional neural network chip that implements BN layer | |
| CN117980920A (en) | Optical computing system, optical computing method and control device | |
| CN117709423B (en) | Deep neural network photon acceleration chip and operation system thereof | |
| US20260087320A1 (en) | Photonic and electronic integrated circuits and methods |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25768288 Country of ref document: EP Kind code of ref document: A1 |