WO2025236263A1 - Multibit analog compute-in-memory using sub-2 radix data format - Google Patents

Multibit analog compute-in-memory using sub-2 radix data format

Info

Publication number
WO2025236263A1
WO2025236263A1 PCT/CN2024/093798 CN2024093798W WO2025236263A1 WO 2025236263 A1 WO2025236263 A1 WO 2025236263A1 CN 2024093798 W CN2024093798 W CN 2024093798W WO 2025236263 A1 WO2025236263 A1 WO 2025236263A1
Authority
WO
WIPO (PCT)
Prior art keywords
data format
binary data
radix
multibit
capacitor
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2024/093798
Other languages
French (fr)
Inventor
Renzhi Liu
Hechen Wang
Richard Dorrance
Brent Carlton
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Intel Corp
Original Assignee
Intel Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Intel Corp filed Critical Intel Corp
Priority to PCT/CN2024/093798 priority Critical patent/WO2025236263A1/en
Publication of WO2025236263A1 publication Critical patent/WO2025236263A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F7/00Methods or arrangements for processing data by operating upon the order or content of the data handled
    • G06F7/38Methods or arrangements for performing computations using exclusively denominational number representation, e.g. using binary, ternary, decimal representation
    • G06F7/48Methods or arrangements for performing computations using exclusively denominational number representation, e.g. using binary, ternary, decimal representation using non-contact-making devices, e.g. tube, solid state device; using unspecified devices
    • G06F7/544Methods or arrangements for performing computations using exclusively denominational number representation, e.g. using binary, ternary, decimal representation using non-contact-making devices, e.g. tube, solid state device; using unspecified devices for evaluating functions by calculation
    • G06F7/5443Sum of products
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F7/00Methods or arrangements for processing data by operating upon the order or content of the data handled
    • G06F7/38Methods or arrangements for performing computations using exclusively denominational number representation, e.g. using binary, ternary, decimal representation
    • G06F7/48Methods or arrangements for performing computations using exclusively denominational number representation, e.g. using binary, ternary, decimal representation using non-contact-making devices, e.g. tube, solid state device; using unspecified devices
    • G06F7/49Computations with a radix, other than binary, 8, 16 or decimal, e.g. ternary, negative or imaginary radices, mixed radix non-linear PCM
    • HELECTRICITY
    • H03ELECTRONIC CIRCUITRY
    • H03MCODING; DECODING; CODE CONVERSION IN GENERAL
    • H03M1/00Analogue/digital conversion; Digital/analogue conversion
    • H03M1/06Continuously compensating for, or preventing, undesired influence of physical parameters
    • H03M1/0617Continuously compensating for, or preventing, undesired influence of physical parameters characterised by the use of methods or means not specific to a particular type of detrimental influence
    • H03M1/0675Continuously compensating for, or preventing, undesired influence of physical parameters characterised by the use of methods or means not specific to a particular type of detrimental influence using redundancy
    • H03M1/069Continuously compensating for, or preventing, undesired influence of physical parameters characterised by the use of methods or means not specific to a particular type of detrimental influence using redundancy by range overlap between successive stages or steps
    • H03M1/0692Continuously compensating for, or preventing, undesired influence of physical parameters characterised by the use of methods or means not specific to a particular type of detrimental influence using redundancy by range overlap between successive stages or steps using a diminished radix representation, e.g. radix 1.95
    • HELECTRICITY
    • H03ELECTRONIC CIRCUITRY
    • H03MCODING; DECODING; CODE CONVERSION IN GENERAL
    • H03M1/00Analogue/digital conversion; Digital/analogue conversion
    • H03M1/66Digital/analogue converters
    • H03M1/662Multiplexed conversion systems

Definitions

  • CiM compute-in-memory
  • CNN convolutional neural network
  • DNN deep neural network
  • CiM does so by directly addressing a “memory wall” in which limited bandwidth between memory and compute hardware in modern architectures causes bottlenecks in the system (e.g., leading to poor computational and energy efficiency) .
  • the development of CiM architectures, however, is challenging to realize as a pure digital system, since the conventional multiply-accumulate (MAC) operation units are too large to fit into high-density Manhattan style memory arrays.
  • MAC multiply-accumulate
  • Analog-based CiM methods may be used to improve efficiency and throughput in-memory MAC computations.
  • charge-domain MAC computation schemes e.g., that use a capacitor array to conduct multiplication and summation operations
  • Analog CiM techniques may still have accuracy disadvantages as compared to digital counterparts.
  • the digital CiM can be considered as “lossless”
  • This condition is especially the case for multibit analog CiM where the multibit representation, e.g. for an N-bit number, would require implementation of precise scaling factors of 1, 2, 2 2 , ..., 2 N-1 in the analog domain (e.g., voltage, current, charge, etc. ) .
  • This requirement is a classic challenge in data-converter design, including digital to analog converter (DAC) and analog to digital converter (ADC) designs, where precision scaling factors are to be implemented to ensure sufficient conversion linearity for preserving data fidelity.
  • DAC digital to analog converter
  • ADC analog to digital converter
  • FIGs. 1A and 1B are schematic diagrams of an example of an analog compute-in-memory (CiM) architecture according to an embodiment
  • FIG. 2 is a comparative plot of an example of a super-2 radix digital to analog conversion curve and a sub-2 radix digital to analog conversion curve according to an embodiment
  • FIG. 3 is a comparative plot of an example of a remapped super-2 radix digital to analog conversion curve and a remapped sub-2 radix digital to analog conversion curve according to an embodiment
  • FIGs. 5 and 6 are flowcharts of examples of methods of operating a performance-enhanced computing system according to an embodiment
  • FIG. 7 is a block diagram of an example of a performance-enhanced computing system according to an embodiment.
  • analog compute-in-memory (CiM) techniques may provide advantages over voltage, current, and time domain methods in terms of linearity and process scalability but may still encounter accuracy challenges.
  • C-2C based weighting schemes for multibit representation in analog CiM alleviates difficulty in physically implementing the exponentially increasing scaling factors by only having unit C-2C designs in a cascaded ladder, this approach does not resolve the multibit weighting precision issue. More particularly, any deviation from the ideal 1: 2 capacitor ratio in the C-2C unit would cause a non-binary scaling in the multibit representation, resulting in missing codes during digital to analog conversion and a lower effective number of bits in analog computation. As a result, the neural network (NN) inference accuracy may degrade significantly.
  • the sub-2 radix scheme described herein adds the flexibility of digital correction through the added radix-2 to sub-2 radix translational layer, to correct systematic multibit weighting error in analog CiM MAC computation.
  • analog in-memory computing can provide superior performance advantages compared to other competing in-memory computing solutions, particularly for edge artificial intelligence (AI) platforms, for achieving both high throughput and high efficiency.
  • AI edge artificial intelligence
  • CMOS complementary metal oxide semiconductor
  • Embodiments can substantially improve the linearity of analog computations without introducing significant hardware or additional calibrations.
  • the sub-2 radix data format described herein for multibit analog CiM macro applies to any type of analog CiM having a nominally radix-2 multibit recombination scheme in the analog domain.
  • a radix-2 to sub-2 radix data format conversion layer 18 is added for the digital input data (D IN ) to be stored in the memory array 12. More particularly, the binary radix-2 data D IN that is to be written in the memory array 12 (e.g., weight data in a weight stationary NN computation scheme) , is processed by the conversion layer 18 to be converted/mapped to the sub-2 radix (e.g., base/radix less than two) data format before the data is stored in the memory array 12.
  • the sub-2 radix format can provide robustness and correct analog multi-bit weighting errors in an analog CiM macro and the conversion layer 18 supports this robustness and error correction.
  • a out C [w 0 b 0 + w 1 b 1 + w 2 b 2 +...w 7 b 7 ] (Eq. 1)
  • Eq. 1 shows how a general 8-bit binary digital data, b 0 to b 7 (least significant bit/LSB to most significant bit/MSB) , and eight corresponding weights w 0 to w 7 , converts to an analog output value A out with a fixed scaling factor C.
  • a out C [m 0 b 0 + m 1 b 1 + m 2 b 2 +...m 7 b 7 ] (Eq. 2)
  • Eq. 2 shows a special case in which the weights are exactly in geometric progression (e.g., radix-m) .
  • radix-m geometric progression
  • the analog MAC computation from an array of stored digital data to analog output activation (OA) is a linear operation for a given set of analog input activation (IA) values. Therefore, the non-ideal weightings from single digital data to analog conversion would directly appear as the same weighting values after multiplication and accumulation in the analog domain (e.g., assuming the weightings are the same for all digital data, which is the case for process variation) . Accordingly, the technology described herein focuses on the linearity issue of a single digital data to analog conversion (e.g., solving the linearity issue means solving the same problem for analog MAC outputs) .
  • capacitor recombination networks are used herein for discussion purposes, other types of devices may be used to conduct multibit recombination operations.
  • resistors and/or transistors may be used in place of capacitors for multibit recombination schemes/circuits.
  • FIG. 2 shows the normalized analog output versus digital inputs for the super-2 radix and sub-2 radix examples.
  • the 8-bit digital binary inputs b 7 b 6 ...b 0 are shown with their respective decimal representation (e.g., a range of 0-255) .
  • a super-2 radix plot 20 demonstrates that the normalized analog output has regions (e.g., “Missing Outputs” ) in the analog domain that cannot be covered by the digital inputs. Thus, there is a significant data conversion accuracy loss, which in turn creates analog CiM MAC computation errors.
  • a sub-2 radix plot 22 demonstrates that the analog outputs would have overlapping output regions. Without correction, this condition still creates data conversion error.
  • the overlapping outputs create redundancy in the data conversion process.
  • the overlapping outputs can be potentially corrected to restore a linear digital to analog conversion curve.
  • this digital correction process of restoring the linearity with radix-m weighting in the analog domain can be conducted in the following operations:
  • Each bit of d 7 , d 6 , ..., d 0 is still binary data, being either 0 or 1.
  • FIG. 3 demonstrates that after applying this correction, with the input radix-2 binary digital code (b 7 b 6 ...b 0 ) 2 remapped to radix-m digital code of (d 7 d 6 ...d 0 ) m and then converted to analog output, a digital to analog conversion super-2 radix curve 30 and a sub-2 radix curve 32 result.
  • the analog output still exhibits missing output values and cannot be corrected to achieve linear digital to analog conversion (e.g., negatively impacting analog CiM computation accuracy) .
  • digital input remapping corrects the sub-2 radix curve 32, which is linear and has no missing analog output values.
  • CMOS process variability results in a high probability that the systematic error can skew the weighting towards super-2 radix (e.g., resulting in irreparable errors in data conversion and hence computation) .
  • a sub-2 radix scheme is adopted for a nominal design point, a much larger margin is created for the overall weighting scheme to remain in sub-2 radix with process variation considered (e.g., systematic error can then be corrected through digital input mapping) .
  • the sub-2 radix data mapping may occur immediately before the digital data is written into the CiM macro for NN weight storage.
  • the actual mapping between the original 8-bit radix-2 input (b 7 b 6 ...b 0 ) 2 , and the remapped 8-bit radix-1.91 digital code (d 7 d 6 ...d 0 ) 1.91 are shown in a curve 40.
  • the proposed input data mapping layer can achieve the mapping results as shown in the curve 40 in at least two different ways:
  • the first approach is to implement the data mapping as a look-up-table (LUT) .
  • LUT look-up-table
  • the LUT size is about 8x256 bits.
  • the second approach is to solve the equation in Eq. 4 via a per bit calculation. Specifically, the solution can be broken down into the following operations for finding d 7 , d 6 , ..., d 0 sequentially:
  • Both approaches can achieve the same functionality for the proposed digital input mapping for the sub-2 radix data format in multi-bit analog CiM.
  • the difference is that the first approach may be “memory heavy” (e.g., LUT storage overhead) , while there is almost no computation involved; and the second approach is “computation heavy” , while the storage is relatively light (e.g., only about 8x8 bits storage is used for sub-2 radix weights in the second approach as opposed to 8x256 bits LUT storage used in the first approach) .
  • both approaches offer flexibility for the sub-2 radix conversion, as the actual sub-2 radix weights or the LUT can be programmed after the chip (e.g., semiconductor package apparatus) is manufactured.
  • the sub-2 radix data format described herein for analog multi-bit CiM is not only limited to the special case where the radix weights are exactly in geometric progression such as, for example, m 0 , m 1 , ...m 7 from Eq. (2) . Indeed, a more general form is presented in Eq. (1) .
  • the only criteria for being a sub-2 radix, rather than a super-2 radix is that the corresponding digital-to-analog conversion curve with original input digital data only results in overlapping analog output values as in the sub-2 radix plot 22 (FIG. 2) without any gaps for missing analog output values as in the super-2 radix plot 20 (FIG. 2) .
  • the loss of effective number of bits is a minor issue when the radix is only slightly smaller than two (e.g., 8-bit of radix-1.91 is equivalent to 7.6 bits of radix-2)
  • the loss of effective number of bits can become major issue if the radix is reduced to a number that is well below two (e.g., 8-bit of radix-1.75 is only about 7 bits in radix-2) .
  • FIG. 5 shows a method 50 of operating a performance-enhanced computing system.
  • the method 50 may be implemented in one or more modules as a set of logic instructions stored in a machine-or computer-readable storage medium such as random access memory (RAM) , read only memory (ROM) , programmable ROM (PROM) , firmware, flash memory, etc., in hardware, or any combination thereof.
  • RAM random access memory
  • ROM read only memory
  • PROM programmable ROM
  • firmware flash memory
  • hardware implementations may include configurable logic, fixed-functionality logic, or any combination thereof.
  • Examples of configurable logic include suitably configured programmable logic arrays (PLAs) , field programmable gate arrays (FPGAs) , complex programmable logic devices (CPLDs) , and general purpose microprocessors.
  • PLAs programmable logic arrays
  • FPGAs field programmable gate arrays
  • CPLDs complex programmable logic devices
  • fixed-functionality logic e.g., fixed-functionality hardware
  • ASICs application specific integrated circuits
  • combinational logic circuits e.g., combinational logic circuits
  • sequential logic circuits e.g., combinational logic circuits
  • CMOS complementary metal oxide semiconductor
  • TTL transistor-transistor logic
  • Illustrated processing block 52 provides for storing multibit weight data to a memory array (e.g., SRAM) in a non-binary data format, wherein the non-binary data format includes a radix that is less than two (e.g., sub-2 radix) .
  • Block 54 conducts, by a capacitor recombination network, MAC operations on first analog signals (e.g., input activations) and the multibit weight data.
  • Block 56 outputs, by the capacitor recombination network, second analog signals (e.g., output activations) based on the MAC operations.
  • the non-binary data format generates overlapping outputs in the second analog signals.
  • the second analog signals are used to draw inferences in an AI application such as, for example, a computer vision application (e.g., automatically classify handwritten digits between 0 and 9) , natural language processing (NLP) application, and so forth.
  • AI application such as, for example, a computer vision application (e.g., automatically classify handwritten digits between 0 and 9) , natural language processing (NLP) application, and so forth.
  • a computer vision application e.g., automatically classify handwritten digits between 0 and 9
  • NLP natural language processing
  • the method 50 therefore enhances performance at least to the extent that the non-binary data format improves the linearity of analog computations without introducing significant hardware and additional calibrations. Accordingly, the method 50 provides robustness in high-precision multibit MAC computations, even in the presence of CMOS process technology manufacturing variability.
  • FIG. 6 shows another method 60 of operating a performance-enhanced computing system.
  • the method 60 may generally be conducted in conjunction with the method 50 (FIG. 5) , already discussed. More particularly, the method 60 may be implemented in one or more modules as a set of logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., in hardware, or any combination thereof.
  • a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc.
  • hardware implementations may include configurable logic, fixed-functionality logic, or any combination thereof.
  • the capacitor recombination network is a C-2C ladder and capacitor pairs in the C-2C ladder include a capacitance ratio (e.g., relative device size) that is greater than two.
  • illustrated processing block 62 provides for determining (e.g., by a conversion layer) , the radix of the non-binary data format based on the capacitance ratio of the capacitor pairs. For example, block 62 might determine that the radix is to be 1.91 when the capacitance ratio is 2.3. Thus, block 62 might involve reading the capacitance ratio from a programmable register and calculating the radix accordingly.
  • Block 64 provides for converting, by a conversion/mapping layer, the multibit weight data from a binary (e.g., radix-2) data format to the non-binary data format.
  • block 64 converts the multibit weight data from the binary data format to the non-binary data format via a lookup table.
  • block 64 converts the multibit weight data from the binary data format to the non-binary data format via a per bit calculation (e.g., Eq. 4) .
  • Block 66 may write, by a control circuit coupled to the conversion layer and the memory array, the multibit weight data to the memory array in the non-binary data format.
  • the system 280 may generally be part of an electronic device/platform having computing functionality (e.g., personal digital assistant/PDA, notebook computer, tablet computer, convertible tablet, server) , communications functionality (e.g., edge networking device/controller, smart phone) , imaging functionality (e.g., camera, camcorder) , media playing functionality (e.g., smart television/TV) , wearable functionality (e.g., watch, eyewear, headwear, footwear, jewelry) , vehicular functionality (e.g., car, truck, motorcycle) , robotic functionality (e.g., autonomous robot) , Internet of Things (IoT) functionality, drone functionality, etc., or any combination thereof.
  • computing functionality e.g., personal digital assistant/PDA, notebook computer, tablet computer, convertible tablet, server
  • communications functionality e.g., edge networking device/controller, smart phone
  • imaging functionality e.g., camera, camcorder
  • media playing functionality e.g., smart television/TV
  • wearable functionality e.g
  • the system 280 includes a host processor 282 (e.g., CPU) having an integrated memory controller (IMC) 284 that is coupled to a system memory 286 (e.g., dual inline memory module/DIMM) .
  • IMC integrated memory controller
  • an IO module 288 is coupled to the host processor 282.
  • the illustrated IO module 288 communicates with, for example, a display 290 (e.g., touch screen, liquid crystal display/LCD, light emitting diode/LED display) , and a network controller 292 (e.g., wired and/or wireless) .
  • the host processor 282 may be combined with the IO module 288, a graphics processor 294, and an AI accelerator 296 into a system on chip (SoC) 298.
  • SoC system on chip
  • the AI accelerator 296 includes logic 300 to perform one or more aspects of the method 50 (FIG. 5) and/or the method 60 (FIG. 6) , already discussed.
  • the logic 300 includes a memory array to store multibit weight data in a non-binary data format, wherein the non-binary data format includes a radix that is less than two (e.g., sub-2 radix) .
  • the logic 300 also includes a capacitor recombination network to conduct MAC operations on first analog signals and the multibit weight data and output second analog signals based on the MAC operations.
  • the capacitor recombination network is a C-2C ladder. In such a case, capacitor pairs in the capacitor recombination network can include a capacitance ratio that is greater than two.
  • the logic 300 is shown within the AI accelerator 296, the logic 300 may reside elsewhere in the computing system 280.
  • the computing system 280 is therefore considered performance-enhanced at least to the extent that the non-binary data format and/or the capacitance ratio improves the linearity of analog computations without introducing significant hardware and additional calibrations. Accordingly, the logic 300 provides robustness in high-precision multibit MAC computations, even in the presence of CMOS process technology manufacturing variability.
  • FIG. 8 shows a semiconductor apparatus 350 (e.g., chip, die, package) .
  • the illustrated apparatus 350 includes one or more substrates 352 (e.g., silicon, sapphire, gallium arsenide) and logic 354 (e.g., transistor array and other integrated circuit/IC components) coupled to the substrate (s) 352.
  • the logic 354 implements one or more aspects of the method 50 (FIG. 5) and/or the method 60 (FIG. 6) , already discussed.
  • the semiconductor apparatus 350 may also be incorporated into the AI accelerator 296 (FIG. 7) .
  • the logic 354 may be implemented at least partly in configurable or fixed-functionality hardware.
  • the logic 354 includes transistor channel regions that are positioned (e.g., embedded) within the substrate (s) 352.
  • the interface between the logic 354 and the substrate (s) 352 may not be an abrupt junction.
  • the logic 354 may also be considered to include an epitaxial layer that is grown on an initial wafer of the substrate (s) 352.
  • Example 1 includes a performance-enhanced computing system comprising a network controller and a processor coupled to the network controller, the processor including logic coupled to one or more substrates, wherein the logic includes a memory array to store multibit weight data in a non-binary data format, wherein the non-binary data format includes a radix that is less than two, and a capacitor recombination network to conduct multiply-accumulate (MAC) operations on first analog signals and the multibit weight data, the capacitor recombination network further to output second analog signals based on the MAC operations.
  • MAC multiply-accumulate
  • Example 2 includes the performance-enhanced computing system of Example 1, wherein the logic further includes a conversion layer to convert the multibit weight data from a binary data format to the non-binary data format.
  • Example 3 includes the performance-enhanced computing system of Example 2, wherein the capacitor recombination network is a C-2C ladder, wherein capacitor pairs in the C-2C ladder include a capacitance ratio that is greater than two, and wherein the conversion layer is to determine the radix of the non-binary data format based on the capacitance ratio of the capacitor pairs.
  • the capacitor recombination network is a C-2C ladder
  • capacitor pairs in the C-2C ladder include a capacitance ratio that is greater than two
  • the conversion layer is to determine the radix of the non-binary data format based on the capacitance ratio of the capacitor pairs.
  • Example 4 includes the performance-enhanced computing system of Example 2, wherein the logic further includes a control circuit coupled to the conversion layer and the memory array, and wherein the control circuit is to write the multibit weight data to the memory array in the non-binary data format.
  • Example 5 includes the performance-enhanced computing system of Example 2, wherein the conversion layer is to convert the multibit weight data from the binary data format to the non-binary data format via a lookup table.
  • Example 6 includes the performance-enhanced computing system of Example 2, wherein the conversion layer is to convert the multibit weight data from the binary data format to the non-binary data format via a per bit calculation.
  • Example 7 includes the performance-enhanced computing system of any one of Examples 1 to 6, wherein the radix of the non-binary data format is to generate overlapping outputs in the second analog signals.
  • Example 8 includes a semiconductor apparatus comprising one or more substrates, and logic coupled to the one or more substrates, wherein the logic is implemented at least partly in one or more of configurable or fixed-functionality hardware, the logic including a memory array to store multibit weight data in a non-binary data format, wherein the non-binary data format includes a radix that is less than two, and a capacitor recombination network to conduct multiply-accumulate (MAC) operations on first analog signals and the multibit weight data, the capacitor recombination network further to output second analog signals based on the MAC operations.
  • MAC multiply-accumulate
  • Example 9 includes the semiconductor apparatus of Example 8, wherein the logic further includes a conversion layer to convert the multibit weight data from a binary data format to the non-binary data format.
  • Example 10 includes the semiconductor apparatus of Example 9, wherein the capacitor recombination network is a C-2C ladder, wherein capacitor pairs in the C-2C ladder include a capacitance ratio that is greater than two, and wherein the conversion layer is to determine the radix of the non-binary data format based on the capacitance ratio of the capacitor pairs.
  • the capacitor recombination network is a C-2C ladder
  • capacitor pairs in the C-2C ladder include a capacitance ratio that is greater than two
  • the conversion layer is to determine the radix of the non-binary data format based on the capacitance ratio of the capacitor pairs.
  • Example 11 includes the semiconductor apparatus of Example 9, wherein the logic further includes a control circuit coupled to the conversion layer and the memory array, and wherein the control circuit is to write the multibit weight data to the memory array in the non-binary data format.
  • Example 12 includes the semiconductor apparatus of Example 9, wherein the conversion layer is to convert the multibit weight data from the binary data format to the non-binary data format via a lookup table.
  • Example 13 includes the semiconductor apparatus of Example 9, wherein the conversion layer is to convert the multibit weight data from the binary data format to the non-binary data format via a per bit calculation.
  • Example 14 includes the semiconductor apparatus of any one of Examples 8 to 13, wherein the radix of the non-binary data format is to generate overlapping outputs in the second analog signals.
  • Example 15 includes the semiconductor apparatus of any one of Examples 8 to 13, wherein the logic coupled to the one or more substrates includes transistor regions that are positioned within the one or more substrates.
  • Example 16 includes a method of operating a performance-enhanced computing system, the method comprising storing multibit weight data to a memory array in a non-binary data format, wherein the non-binary data format includes a radix that is less than two, and conducting, by a capacitor recombination network, multiply-accumulate (MAC) operations on first analog signals and the multibit weight data, and outputting, by the capacitor recombination network, second analog signals based on the MAC operations.
  • MAC multiply-accumulate
  • Example 17 includes the method of Example 16, further including converting, by a conversion layer, the multibit weight data from a binary data format to the non-binary data format.
  • Example 18 includes the method of Example 17, wherein the capacitor recombination network is a C-2C ladder, wherein capacitor pairs in the C-2C ladder include a capacitance ratio that is greater than two, and wherein the method further includes determining, by the conversion layer, the radix of the non-binary data format based on the capacitance ratio of the capacitor pairs.
  • Example 19 includes the method of Example 17, further including writing, by a control circuit coupled to the conversion layer and the memory array, the multibit weight data to the memory array in the non-binary data format, wherein the multibit weight data is converted from the binary data format to the non-binary data format via one or more of a lookup table or a per bit calculation.
  • Example 20 includes the method of any one of Examples 16 to 19, wherein the radix of the non-binary data format generates overlapping outputs in the second analog signals.
  • Example 21 includes an apparatus comprising means for performing the method of any one of Examples 16 to 20.
  • Embodiments are applicable for use with all types of semiconductor integrated circuit ( “IC” ) chips.
  • IC semiconductor integrated circuit
  • Examples of these IC chips include but are not limited to processors, controllers, chipset components, programmable logic arrays (PLAs) , memory chips, network chips, systems on chip (SoCs) , SSD/NAND controller ASICs, and the like.
  • PLAs programmable logic arrays
  • SoCs systems on chip
  • SSD/NAND controller ASICs solid state drive/NAND controller ASICs
  • signal conductor lines are represented with lines. Some may be different, to indicate more constituent signal paths, have a number label, to indicate a number of constituent signal paths, and/or have arrows at one or more ends, to indicate primary information flow direction. This, however, should not be construed in a limiting manner.
  • Any represented signal lines may actually comprise one or more signals that may travel in multiple directions and may be implemented with any suitable type of signal scheme, e.g., digital or analog lines implemented with differential pairs, optical fiber lines, and/or single-ended lines.
  • Example sizes/models/values/ranges may have been given, although embodiments are not limited to the same. As manufacturing techniques (e.g., photolithography) mature over time, it is expected that devices of smaller size could be manufactured.
  • well known power/ground connections to IC chips and other components may or may not be shown within the figures, for simplicity of illustration and discussion, and so as not to obscure certain aspects of the embodiments.
  • arrangements may be shown in block diagram form in order to avoid obscuring embodiments, and also in view of the fact that specifics with respect to implementation of such block diagram arrangements are highly dependent upon the computing system within which the embodiment is to be implemented, i.e., such specifics should be well within purview of one skilled in the art.
  • Coupled may be used herein to refer to any type of relationship, direct or indirect, between the components in question, and may apply to electrical, mechanical, fluid, optical, electromagnetic, electromechanical or other connections.
  • first may be used herein only to facilitate discussion, and carry no particular temporal or chronological significance unless otherwise indicated.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Computational Mathematics (AREA)
  • Computing Systems (AREA)
  • Mathematical Analysis (AREA)
  • Mathematical Optimization (AREA)
  • Pure & Applied Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Nonlinear Science (AREA)
  • Analogue/Digital Conversion (AREA)

Abstract

Systems, apparatuses and methods may provide for technology that includes a memory array to store multibit weight data in a non-binary data format, wherein the non-binary data format includes a radix that is less than two, and a capacitor recombination network to conduct multiply-accumulate (MAC) operations on first analog signals and the multibit weight data, the capacitor recombination network further to output second analog signals based on the MAC operations.

Description

MULTIBIT ANALOG COMPUTE-IN-MEMORY USING SUB-2 RADIX DATA FORMAT BACKGROUND
In recent years, compute-in-memory (CiM) architectures have become one of the leading hardware candidates for accelerating the execution and training of convolutional neural network (CNN) and deep neural network (DNN) applications. CiM does so by directly addressing a “memory wall” in which limited bandwidth between memory and compute hardware in modern architectures causes bottlenecks in the system (e.g., leading to poor computational and energy efficiency) . The development of CiM architectures, however, is challenging to realize as a pure digital system, since the conventional multiply-accumulate (MAC) operation units are too large to fit into high-density Manhattan style memory arrays.
Analog-based CiM methods may be used to improve efficiency and throughput in-memory MAC computations. Among these analog CiM solutions, charge-domain MAC computation schemes (e.g., that use a capacitor array to conduct multiplication and summation operations) have been shown to provide significant advantages over voltage, current, and time-domain methods in terms of linearity and process scalability.
Analog CiM techniques, however, may still have accuracy disadvantages as compared to digital counterparts. Whereas the digital CiM can be considered as “lossless” , there is computation accuracy loss that is intrinsic to analog CiM due to non-ideality of the analog circuits. This condition is especially the case for multibit analog CiM where the multibit representation, e.g. for an N-bit number, would require implementation of precise scaling factors of 1, 2, 22, …, 2N-1 in the analog domain (e.g., voltage, current, charge, etc. ) . This requirement is a classic challenge in data-converter design, including digital to analog converter (DAC) and analog to digital converter (ADC) designs, where precision scaling factors are to be implemented to ensure sufficient conversion linearity for preserving data fidelity.
BRIEF DESCRIPTION OF THE DRAWINGS
The various advantages of the embodiments will become apparent to one skilled in the art by reading the following specification and appended claims, and by referencing the following drawings, in which:
FIGs. 1A and 1B are schematic diagrams of an example of an analog compute-in-memory (CiM) architecture according to an embodiment;
FIG. 2 is a comparative plot of an example of a super-2 radix digital to analog conversion curve and a sub-2 radix digital to analog conversion curve according to an embodiment;
FIG. 3 is a comparative plot of an example of a remapped super-2 radix digital to analog conversion curve and a remapped sub-2 radix digital to analog conversion curve according to an embodiment;
FIG. 4 is a plot of an example of an original 8-bit digital input versus a remapped digital input for a radix-1.91 data format according to an embodiment;
FIGs. 5 and 6 are flowcharts of examples of methods of operating a performance-enhanced computing system according to an embodiment;
FIG. 7 is a block diagram of an example of a performance-enhanced computing system according to an embodiment; and
FIG. 8 is an illustration of an example of a semiconductor package apparatus according to an embodiment.
DETAILED DESCRIPTION
As already noted, analog compute-in-memory (CiM) techniques may provide advantages over voltage, current, and time domain methods in terms of linearity and process scalability but may still encounter accuracy challenges. Although previously developed C-2C based weighting schemes for multibit representation in analog CiM alleviates difficulty in physically implementing the exponentially increasing scaling factors by only having unit C-2C designs in a cascaded ladder, this approach does not resolve the multibit weighting precision issue. More particularly, any deviation from the ideal 1: 2 capacitor ratio in the C-2C unit would cause a non-binary scaling in the multibit representation, resulting in missing codes during digital to analog conversion and a lower effective number of bits in analog computation. As a result, the neural network (NN) inference accuracy may degrade significantly.
Indeed, binary weighting precision is a classic problem in digital to analog converter (DAC) and analog to digital converter (ADC) designs. To overcome this problem, traditional methods may conduct calibration on the analog components through small calibration units for fine adjustment post silicon fabrication. This approach may be suitable for DAC and ADC designs where the analog components are large such that overhead on the calibration units is acceptable. In compact analog CiM macros, however, this approach adds significant overhead  to the macro, thus degrading the MAC computation density for correcting the multi-bit conversion accuracy.
As will be discussed in greater detail, the technology described herein provides a robust multibit representation scheme for analog CiM macros that achieves high computation accuracy. More particularly, it has been determined that a sub-2 radix (e.g., base of less than two) weighting scheme for the analog components in analog CiM computation and data storage hardware may generate redundancies that can be exploited to allow errors that can be exploited to allow systematic errors and global mismatches on the weightings that occur during chip fabrication due to manufacturing variability. Through a reconfigurable sub-2 radix translation in the digital domain, the digital codes are then converted to correct the weighting errors from the analog domain (e.g., therefore achieving robust and high-accuracy multi-bit analog MAC operations) .
Embodiments therefore adopt a sub-2 radix weighting scheme for multibit recombination in analog CiM MAC computation. For example, an ideal C-2C ladder for binary radix-2 recombination is now skewed intentionally in design to achieve a sub-2 radix weighting. By doing so, the stored weight data (e.g., in a weight stationary scheme) in an analog CiM macro also has a corresponding sub-2 radix bit representation. Accordingly, a translational mapping layer from the radix-2 (e.g., binary) data format to the sub-2 radix data format is added for input weight data (e.g., in the 8-bit integer/INT8 data format before the data is written into the CiM macro) . By contrast, no special treatment is applied to the analog output of the CiM MAC computation to restore linearity.
The sub-2 radix scheme described herein adds the flexibility of digital correction through the added radix-2 to sub-2 radix translational layer, to correct systematic multibit weighting error in analog CiM MAC computation. Moreover, analog in-memory computing can provide superior performance advantages compared to other competing in-memory computing solutions, particularly for edge artificial intelligence (AI) platforms, for achieving both high throughput and high efficiency.
The technology described herein addresses a significant disadvantage of analog CiM when compared to existing digital accelerators –robustness in high-precision multibit MAC computation, which is exacerbated with increased complementary metal oxide semiconductor (CMOS) process technology manufacturing variability. Embodiments can substantially improve the linearity of analog computations without introducing significant hardware or additional calibrations. Moreover, the sub-2 radix data format described herein for multibit  analog CiM macro applies to any type of analog CiM having a nominally radix-2 multibit recombination scheme in the analog domain.
Turning now to FIGs. 1A and 1B, an analog CiM architecture 10 is shown in which a sub-2 radix data format scheme is added to a memory array 12 (e.g., static random access memory/SRAM organized into sub-banks) and a capacitor recombination network 14. As best seen in FIG. 1B, the ideal 1: 2 ratio of capacitor pairs (e.g., relative device sizes) in the capacitor recombination network 14 has been changed to reflect an overall sub-2 radix weighting scheme. Originally, to achieve ideal radix-2 (e.g., base two, binary) weighting, the Ceq capacitance, which is the equivalent capacitance on each single-bit partial output activation (OA) line 16 seen from the capacitor recombination network 14, and the C2C capacitance, which is the series connecting capacitor in the C-2C ladder, meets a requirement/constraint of C2C =2·Ceq. For the proposed sub-2 radix weighting scheme, the C2C is now sized such that C2C =n·Ceq and n>2. Albeit counter-intuitive, a capacitance ratio greater than two in the capacitor recombination network 14 results in a sub-2 radix weighting.
Additionally, a radix-2 to sub-2 radix data format conversion layer 18 is added for the digital input data (DIN) to be stored in the memory array 12. More particularly, the binary radix-2 data DIN that is to be written in the memory array 12 (e.g., weight data in a weight stationary NN computation scheme) , is processed by the conversion layer 18 to be converted/mapped to the sub-2 radix (e.g., base/radix less than two) data format before the data is stored in the memory array 12. As will be discussed in greater detail, the sub-2 radix format can provide robustness and correct analog multi-bit weighting errors in an analog CiM macro and the conversion layer 18 supports this robustness and error correction.
Benefits of sub-2 radix data format for multi-bit analog CiM macro
Without losing generality, an 8-bit digital to analog conversion can be used as an example to explain the importance of sub-2 radix data format.
Aout = C [w0b0+ w1b1+ w2b2+…w7b7]    (Eq. 1)
Eq. 1 shows how a general 8-bit binary digital data, b0 to b7 (least significant bit/LSB to most significant bit/MSB) , and eight corresponding weights w0 to w7, converts to an analog output value Aout with a fixed scaling factor C.
Aout = C [m0b0+ m1b1+ m2b2+…m7b7]    (Eq. 2)
Eq. 2 shows a special case in which the weights are exactly in geometric progression (e.g., radix-m) . Here, when m=2, an ideal radix-2 conversion results as in a conventional digital  to analog converter. For sub-2 radix digital to analog conversion, m<2; whereas for super-2 radix, m>2.
The analog MAC computation from an array of stored digital data to analog output activation (OA) is a linear operation for a given set of analog input activation (IA) values. Therefore, the non-ideal weightings from single digital data to analog conversion would directly appear as the same weighting values after multiplication and accumulation in the analog domain (e.g., assuming the weightings are the same for all digital data, which is the case for process variation) . Accordingly, the technology described herein focuses on the linearity issue of a single digital data to analog conversion (e.g., solving the linearity issue means solving the same problem for analog MAC outputs) .
In the capacitor recombination network 14 based multibit recombination scheme, two different radix examples are further used for illustration. When C2C =2.3·Ceq, it can be shown mathematically that the radix is 1.91 (m=1.91, e.g., sub-2 radix) . Meanwhile, when C2C =1.7·Ceq, the radix is 2.12 (m=2.12, e.g., super-2 radix) .
Although capacitor recombination networks are used herein for discussion purposes, other types of devices may be used to conduct multibit recombination operations. For example, resistors and/or transistors may be used in place of capacitors for multibit recombination schemes/circuits.
FIG. 2 shows the normalized analog output versus digital inputs for the super-2 radix and sub-2 radix examples. For ease of reading, the 8-bit digital binary inputs b7b6…b0 are shown with their respective decimal representation (e.g., a range of 0-255) . A super-2 radix plot 20 demonstrates that the normalized analog output has regions (e.g., “Missing Outputs” ) in the analog domain that cannot be covered by the digital inputs. Thus, there is a significant data conversion accuracy loss, which in turn creates analog CiM MAC computation errors. By contrast, a sub-2 radix plot 22 demonstrates that the analog outputs would have overlapping output regions. Without correction, this condition still creates data conversion error. As opposed to the super-2 radix plot 20, however, where there is information loss due to the missing outputs, in the sub-2 radix plot 22, the overlapping outputs create redundancy in the data conversion process. The overlapping outputs can be potentially corrected to restore a linear digital to analog conversion curve.
Specifically, this digital correction process of restoring the linearity with radix-m weighting in the analog domain can be conducted in the following operations:
- For the original 8-bit digital input (b7b6…b02, the value is scaled by S= (11111111) m/ (11111111) 2. For example, when correcting for an 8-bit radix-1.91 data format, the result would scale full-scale input data from 255 to about 195, and thus S=0.76.
- For the scaled digital input, which is S· (b7b6…b02, find the 8-bit radix-m representation (d7d6…d0m, such that (d7d6…d0m = S· (b7b6…b02 (Eq. 3) . Since finding the exact matching condition may not be practical, this operation can essentially minimize the error for Eq. 3. In other words, for each 8-bit input (b7b6…b02, this operation finds 8-bits data of d7, d6, …, d0 that is the closest to satisfy Eq. (4) as shown below. Each bit of d7, d6, …, d0 is still binary data, being either 0 or 1.
m0·d0+ m1·d1+…+m7·d7 = S· [20b0+ 21b1+ 22b2+…27b7 = (m0+ m1+…+m7) / (20+ 21+…
+2) · [20b0+ 21b1+ 22b2+…27b7]    (Eq. 4)
FIG. 3 demonstrates that after applying this correction, with the input radix-2 binary digital code (b7b6…b02 remapped to radix-m digital code of (d7d6…d0m and then converted to analog output, a digital to analog conversion super-2 radix curve 30 and a sub-2 radix curve 32 result. Using the same radix-2.12 as an example for the super-2 radix case, even after proposed digital input remapping, the analog output still exhibits missing output values and cannot be corrected to achieve linear digital to analog conversion (e.g., negatively impacting analog CiM computation accuracy) . For the sub-2 radix example of radix-1.91, however, digital input remapping corrects the sub-2 radix curve 32, which is linear and has no missing analog output values.
If the original design point is ideal radix-2 weighting in the analog domain, then CMOS process variability results in a high probability that the systematic error can skew the weighting towards super-2 radix (e.g., resulting in irreparable errors in data conversion and hence computation) . Meanwhile, if a sub-2 radix scheme is adopted for a nominal design point, a much larger margin is created for the overall weighting scheme to remain in sub-2 radix with process variation considered (e.g., systematic error can then be corrected through digital input mapping) .
Implementation of conversion layer for sub-2 radix data format in analog CiM
The aforementioned examples demonstrate that by adopting sub-2 radix weighting, the systematic weighting error in the multi-bit representation for an analog CiM can be corrected through digital input mapping. Implementation of this mapping layer can be as follows.
Turning now to FIG. 4 and continuing to use the same sub-2 radix example (e.g., radix-1.91) in a C-2C ladder based multibit analog CiM, the sub-2 radix data mapping may occur  immediately before the digital data is written into the CiM macro for NN weight storage. The actual mapping between the original 8-bit radix-2 input (b7b6…b02, and the remapped 8-bit radix-1.91 digital code (d7d6…d01.91 are shown in a curve 40. Even though the illustrated example is a sub-2 radix representation for (d7d6…d01.91, both digital input and remapped digital inputs are shown after decimal conversion with the radix-2 format for discussion purposes (e.g., decimal converted (b7b6…b02 versus decimal converted (d7d6…d02) .
The proposed input data mapping layer can achieve the mapping results as shown in the curve 40 in at least two different ways:
- The first approach is to implement the data mapping as a look-up-table (LUT) . In the case of 8-bit input data remapping as shown in the plot 40, the LUT size is about 8x256 bits.
- The second approach is to solve the equation in Eq. 4 via a per bit calculation. Specifically, the solution can be broken down into the following operations for finding d7, d6, …, d0 sequentially:
- If S· [20b0+ 21b1+ 22b2+…27b7] > m7, then d7 = 1; otherwise d7 = 0;
- If S· [20b0+ 21b1+ 22b2+…27b7] -m7·d7 > m6, then d6 = 1; otherwise d6 = 0;
- In general, once d7, d6, …, di+1 is found, then do the following:
if S· [20b0+ 21b1+ 22b2+…27b7] -m7·d7 -m6·d6 -…-mi+1·di+1 > mi, then di = 1; otherwise di = 0;
- Repeat the above operations until d0 is found.
Both approaches can achieve the same functionality for the proposed digital input mapping for the sub-2 radix data format in multi-bit analog CiM. The difference is that the first approach may be “memory heavy” (e.g., LUT storage overhead) , while there is almost no computation involved; and the second approach is “computation heavy” , while the storage is relatively light (e.g., only about 8x8 bits storage is used for sub-2 radix weights in the second approach as opposed to 8x256 bits LUT storage used in the first approach) . Additionally, both approaches offer flexibility for the sub-2 radix conversion, as the actual sub-2 radix weights or the LUT can be programmed after the chip (e.g., semiconductor package apparatus) is manufactured.
Of particular note is that the sub-2 radix data format described herein for analog multi-bit CiM is not only limited to the special case where the radix weights are exactly in geometric progression such as, for example, m0, m1, …m7 from Eq. (2) . Indeed, a more general form is presented in Eq. (1) . In this case, the only criteria for being a sub-2 radix, rather than a super-2 radix, is that the corresponding digital-to-analog conversion curve with original input digital  data only results in overlapping analog output values as in the sub-2 radix plot 22 (FIG. 2) without any gaps for missing analog output values as in the super-2 radix plot 20 (FIG. 2) . Furthermore, although a C-2C ladder based analog CiM architecture has been used as an example for discussion purposes, the technology described herein applies to other analog CiM implementations as long as the multi-bit recombination occurs in the analog domain through device sizing (e.g., capacitance ratios) . Therefore, the sub-2 radix weighting described herein can be applied to the analog values of the device sizes, while the conversion layer (e.g., sub-2 radix data mapping layer) remains the same as in the analog CiM architecture 10 (FIGs. 1A and 1B) .
Thus, for super-2 radix cases, there may be significant inference accuracy drops. This condition can be understood as the aforementioned missing output issue leading to uncorrectable errors (e.g., thus resulting in a linearity issue and computation error in analog MAC, and therefore degraded inferenced accuracy) . By contrast, the inference accuracies for sub-2 radix cases are comparable to the ideal radix-2 case, and sometimes even better than radix-2, showing the resilience and robustness against analog non-ideality by adopting sub-2 radix. When the radix is substantially smaller than two, then the inference accuracy may drop again because even though the linearity can be corrected for sub-2 radix by the technology described herein, the effective number of bits are still reduced for sub-2 radix format. Although the loss of effective number of bits is a minor issue when the radix is only slightly smaller than two (e.g., 8-bit of radix-1.91 is equivalent to 7.6 bits of radix-2) , the loss of effective number of bits can become major issue if the radix is reduced to a number that is well below two (e.g., 8-bit of radix-1.75 is only about 7 bits in radix-2) .
FIG. 5 shows a method 50 of operating a performance-enhanced computing system. The method 50 may be implemented in one or more modules as a set of logic instructions stored in a machine-or computer-readable storage medium such as random access memory (RAM) , read only memory (ROM) , programmable ROM (PROM) , firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations may include configurable logic, fixed-functionality logic, or any combination thereof. Examples of configurable logic (e.g., configurable hardware) include suitably configured programmable logic arrays (PLAs) , field programmable gate arrays (FPGAs) , complex programmable logic devices (CPLDs) , and general purpose microprocessors. Examples of fixed-functionality logic (e.g., fixed-functionality hardware) include suitably configured application specific integrated circuits (ASICs) , combinational logic circuits, and sequential logic circuits. The configurable  or fixed-functionality logic can be implemented with complementary metal oxide semiconductor (CMOS) logic circuits, transistor-transistor logic (TTL) logic circuits, or other circuits.
Illustrated processing block 52 provides for storing multibit weight data to a memory array (e.g., SRAM) in a non-binary data format, wherein the non-binary data format includes a radix that is less than two (e.g., sub-2 radix) . Block 54 conducts, by a capacitor recombination network, MAC operations on first analog signals (e.g., input activations) and the multibit weight data. Block 56 outputs, by the capacitor recombination network, second analog signals (e.g., output activations) based on the MAC operations. In one example, the non-binary data format generates overlapping outputs in the second analog signals. In an embodiment, the second analog signals are used to draw inferences in an AI application such as, for example, a computer vision application (e.g., automatically classify handwritten digits between 0 and 9) , natural language processing (NLP) application, and so forth.
The method 50 therefore enhances performance at least to the extent that the non-binary data format improves the linearity of analog computations without introducing significant hardware and additional calibrations. Accordingly, the method 50 provides robustness in high-precision multibit MAC computations, even in the presence of CMOS process technology manufacturing variability.
FIG. 6 shows another method 60 of operating a performance-enhanced computing system. The method 60 may generally be conducted in conjunction with the method 50 (FIG. 5) , already discussed. More particularly, the method 60 may be implemented in one or more modules as a set of logic instructions stored in a machine-or computer-readable storage medium such as RAM, ROM, PROM, firmware, flash memory, etc., in hardware, or any combination thereof. For example, hardware implementations may include configurable logic, fixed-functionality logic, or any combination thereof.
In one example, the capacitor recombination network is a C-2C ladder and capacitor pairs in the C-2C ladder include a capacitance ratio (e.g., relative device size) that is greater than two. In such a case, illustrated processing block 62 provides for determining (e.g., by a conversion layer) , the radix of the non-binary data format based on the capacitance ratio of the capacitor pairs. For example, block 62 might determine that the radix is to be 1.91 when the capacitance ratio is 2.3. Thus, block 62 might involve reading the capacitance ratio from a programmable register and calculating the radix accordingly. Block 64 provides for converting, by a conversion/mapping layer, the multibit weight data from a binary (e.g., radix-2) data format  to the non-binary data format. In one example, block 64 converts the multibit weight data from the binary data format to the non-binary data format via a lookup table. In another example, block 64 converts the multibit weight data from the binary data format to the non-binary data format via a per bit calculation (e.g., Eq. 4) . Block 66 may write, by a control circuit coupled to the conversion layer and the memory array, the multibit weight data to the memory array in the non-binary data format.
Turning now to FIG. 7, a performance-enhanced computing system 280 is shown. The system 280 may generally be part of an electronic device/platform having computing functionality (e.g., personal digital assistant/PDA, notebook computer, tablet computer, convertible tablet, server) , communications functionality (e.g., edge networking device/controller, smart phone) , imaging functionality (e.g., camera, camcorder) , media playing functionality (e.g., smart television/TV) , wearable functionality (e.g., watch, eyewear, headwear, footwear, jewelry) , vehicular functionality (e.g., car, truck, motorcycle) , robotic functionality (e.g., autonomous robot) , Internet of Things (IoT) functionality, drone functionality, etc., or any combination thereof.
In the illustrated example, the system 280 includes a host processor 282 (e.g., CPU) having an integrated memory controller (IMC) 284 that is coupled to a system memory 286 (e.g., dual inline memory module/DIMM) . In an embodiment, an IO module 288 is coupled to the host processor 282. The illustrated IO module 288 communicates with, for example, a display 290 (e.g., touch screen, liquid crystal display/LCD, light emitting diode/LED display) , and a network controller 292 (e.g., wired and/or wireless) . The host processor 282 may be combined with the IO module 288, a graphics processor 294, and an AI accelerator 296 into a system on chip (SoC) 298.
In an embodiment, the AI accelerator 296 includes logic 300 to perform one or more aspects of the method 50 (FIG. 5) and/or the method 60 (FIG. 6) , already discussed. Thus, the logic 300 includes a memory array to store multibit weight data in a non-binary data format, wherein the non-binary data format includes a radix that is less than two (e.g., sub-2 radix) . The logic 300 also includes a capacitor recombination network to conduct MAC operations on first analog signals and the multibit weight data and output second analog signals based on the MAC operations. As already noted, the capacitor recombination network is a C-2C ladder. In such a case, capacitor pairs in the capacitor recombination network can include a capacitance ratio that is greater than two. Although the logic 300 is shown within the AI accelerator 296, the logic 300 may reside elsewhere in the computing system 280.
The computing system 280 is therefore considered performance-enhanced at least to the extent that the non-binary data format and/or the capacitance ratio improves the linearity of analog computations without introducing significant hardware and additional calibrations. Accordingly, the logic 300 provides robustness in high-precision multibit MAC computations, even in the presence of CMOS process technology manufacturing variability.
FIG. 8 shows a semiconductor apparatus 350 (e.g., chip, die, package) . The illustrated apparatus 350 includes one or more substrates 352 (e.g., silicon, sapphire, gallium arsenide) and logic 354 (e.g., transistor array and other integrated circuit/IC components) coupled to the substrate (s) 352. In an embodiment, the logic 354 implements one or more aspects of the method 50 (FIG. 5) and/or the method 60 (FIG. 6) , already discussed. The semiconductor apparatus 350 may also be incorporated into the AI accelerator 296 (FIG. 7) .
The logic 354 may be implemented at least partly in configurable or fixed-functionality hardware. In one example, the logic 354 includes transistor channel regions that are positioned (e.g., embedded) within the substrate (s) 352. Thus, the interface between the logic 354 and the substrate (s) 352 may not be an abrupt junction. The logic 354 may also be considered to include an epitaxial layer that is grown on an initial wafer of the substrate (s) 352.
Additional Notes and Examples:
Example 1 includes a performance-enhanced computing system comprising a network controller and a processor coupled to the network controller, the processor including logic coupled to one or more substrates, wherein the logic includes a memory array to store multibit weight data in a non-binary data format, wherein the non-binary data format includes a radix that is less than two, and a capacitor recombination network to conduct multiply-accumulate (MAC) operations on first analog signals and the multibit weight data, the capacitor recombination network further to output second analog signals based on the MAC operations.
Example 2 includes the performance-enhanced computing system of Example 1, wherein the logic further includes a conversion layer to convert the multibit weight data from a binary data format to the non-binary data format.
Example 3 includes the performance-enhanced computing system of Example 2, wherein the capacitor recombination network is a C-2C ladder, wherein capacitor pairs in the C-2C ladder include a capacitance ratio that is greater than two, and wherein the conversion layer is to determine the radix of the non-binary data format based on the capacitance ratio of the capacitor pairs.
Example 4 includes the performance-enhanced computing system of Example 2, wherein the logic further includes a control circuit coupled to the conversion layer and the memory array, and wherein the control circuit is to write the multibit weight data to the memory array in the non-binary data format.
Example 5 includes the performance-enhanced computing system of Example 2, wherein the conversion layer is to convert the multibit weight data from the binary data format to the non-binary data format via a lookup table.
Example 6 includes the performance-enhanced computing system of Example 2, wherein the conversion layer is to convert the multibit weight data from the binary data format to the non-binary data format via a per bit calculation.
Example 7 includes the performance-enhanced computing system of any one of Examples 1 to 6, wherein the radix of the non-binary data format is to generate overlapping outputs in the second analog signals.
Example 8 includes a semiconductor apparatus comprising one or more substrates, and logic coupled to the one or more substrates, wherein the logic is implemented at least partly in one or more of configurable or fixed-functionality hardware, the logic including a memory array to store multibit weight data in a non-binary data format, wherein the non-binary data format includes a radix that is less than two, and a capacitor recombination network to conduct multiply-accumulate (MAC) operations on first analog signals and the multibit weight data, the capacitor recombination network further to output second analog signals based on the MAC operations.
Example 9 includes the semiconductor apparatus of Example 8, wherein the logic further includes a conversion layer to convert the multibit weight data from a binary data format to the non-binary data format.
Example 10 includes the semiconductor apparatus of Example 9, wherein the capacitor recombination network is a C-2C ladder, wherein capacitor pairs in the C-2C ladder include a capacitance ratio that is greater than two, and wherein the conversion layer is to determine the radix of the non-binary data format based on the capacitance ratio of the capacitor pairs.
Example 11 includes the semiconductor apparatus of Example 9, wherein the logic further includes a control circuit coupled to the conversion layer and the memory array, and wherein the control circuit is to write the multibit weight data to the memory array in the non-binary data format.
Example 12 includes the semiconductor apparatus of Example 9, wherein the conversion layer is to convert the multibit weight data from the binary data format to the non-binary data format via a lookup table.
Example 13 includes the semiconductor apparatus of Example 9, wherein the conversion layer is to convert the multibit weight data from the binary data format to the non-binary data format via a per bit calculation.
Example 14 includes the semiconductor apparatus of any one of Examples 8 to 13, wherein the radix of the non-binary data format is to generate overlapping outputs in the second analog signals.
Example 15 includes the semiconductor apparatus of any one of Examples 8 to 13, wherein the logic coupled to the one or more substrates includes transistor regions that are positioned within the one or more substrates.
Example 16 includes a method of operating a performance-enhanced computing system, the method comprising storing multibit weight data to a memory array in a non-binary data format, wherein the non-binary data format includes a radix that is less than two, and conducting, by a capacitor recombination network, multiply-accumulate (MAC) operations on first analog signals and the multibit weight data, and outputting, by the capacitor recombination network, second analog signals based on the MAC operations.
Example 17 includes the method of Example 16, further including converting, by a conversion layer, the multibit weight data from a binary data format to the non-binary data format.
Example 18 includes the method of Example 17, wherein the capacitor recombination network is a C-2C ladder, wherein capacitor pairs in the C-2C ladder include a capacitance ratio that is greater than two, and wherein the method further includes determining, by the conversion layer, the radix of the non-binary data format based on the capacitance ratio of the capacitor pairs.
Example 19 includes the method of Example 17, further including writing, by a control circuit coupled to the conversion layer and the memory array, the multibit weight data to the memory array in the non-binary data format, wherein the multibit weight data is converted from the binary data format to the non-binary data format via one or more of a lookup table or a per bit calculation.
Example 20 includes the method of any one of Examples 16 to 19, wherein the radix of the non-binary data format generates overlapping outputs in the second analog signals.
Example 21 includes an apparatus comprising means for performing the method of any one of Examples 16 to 20.
Embodiments are applicable for use with all types of semiconductor integrated circuit ( “IC” ) chips. Examples of these IC chips include but are not limited to processors, controllers, chipset components, programmable logic arrays (PLAs) , memory chips, network chips, systems on chip (SoCs) , SSD/NAND controller ASICs, and the like. In addition, in some of the drawings, signal conductor lines are represented with lines. Some may be different, to indicate more constituent signal paths, have a number label, to indicate a number of constituent signal paths, and/or have arrows at one or more ends, to indicate primary information flow direction. This, however, should not be construed in a limiting manner. Rather, such added detail may be used in connection with one or more exemplary embodiments to facilitate easier understanding of a circuit. Any represented signal lines, whether or not having additional information, may actually comprise one or more signals that may travel in multiple directions and may be implemented with any suitable type of signal scheme, e.g., digital or analog lines implemented with differential pairs, optical fiber lines, and/or single-ended lines.
Example sizes/models/values/ranges may have been given, although embodiments are not limited to the same. As manufacturing techniques (e.g., photolithography) mature over time, it is expected that devices of smaller size could be manufactured. In addition, well known power/ground connections to IC chips and other components may or may not be shown within the figures, for simplicity of illustration and discussion, and so as not to obscure certain aspects of the embodiments. Further, arrangements may be shown in block diagram form in order to avoid obscuring embodiments, and also in view of the fact that specifics with respect to implementation of such block diagram arrangements are highly dependent upon the computing system within which the embodiment is to be implemented, i.e., such specifics should be well within purview of one skilled in the art. Where specific details (e.g., circuits) are set forth in order to describe example embodiments, it should be apparent to one skilled in the art that embodiments can be practiced without, or with variation of, these specific details. The description is thus to be regarded as illustrative instead of limiting.
The term “coupled” may be used herein to refer to any type of relationship, direct or indirect, between the components in question, and may apply to electrical, mechanical, fluid, optical, electromagnetic, electromechanical or other connections. In addition, the terms “first” , “second” , etc. may be used herein only to facilitate discussion, and carry no particular temporal or chronological significance unless otherwise indicated.
As used in this application and in the claims, a list of items joined by the term “one or more of” may mean any combination of the listed terms. For example, the phrases “one or more of A, B or C” may mean A; B; C; A and B; A and C; B and C; or A, B and C.
Those skilled in the art will appreciate from the foregoing description that the broad techniques of the embodiments can be implemented in a variety of forms. Therefore, while the embodiments have been described in connection with particular examples thereof, the true scope of the embodiments should not be so limited since other modifications will become apparent to the skilled practitioner upon a study of the drawings, specification, and following claims.

Claims (20)

  1. A performance-enhanced computing system comprising:
    a network controller; and
    a processor coupled to the network controller, the processor including logic coupled to one or more substrates, wherein the logic includes:
    a memory array to store multibit weight data in a non-binary data format, wherein the non-binary data format includes a radix that is less than two, and
    a capacitor recombination network to conduct multiply-accumulate (MAC) operations on first analog signals and the multibit weight data, the capacitor recombination network further to output second analog signals based on the MAC operations.
  2. The performance-enhanced computing system of claim 1, wherein the logic further includes a conversion layer to convert the multibit weight data from a binary data format to the non-binary data format.
  3. The performance-enhanced computing system of claim 2, wherein the capacitor recombination network is a C-2C ladder, wherein capacitor pairs in the C-2C ladder include a capacitance ratio that is greater than two, and wherein the conversion layer is to determine the radix of the non-binary data format based on the capacitance ratio of the capacitor pairs.
  4. The performance-enhanced computing system of claim 2, wherein the logic further includes a control circuit coupled to the conversion layer and the memory array, and wherein the control circuit is to write the multibit weight data to the memory array in the non-binary data format.
  5. The performance-enhanced computing system of claim 2, wherein the conversion layer is to convert the multibit weight data from the binary data format to the non-binary data format via a lookup table.
  6. The performance-enhanced computing system of claim 2, wherein the conversion layer is to convert the multibit weight data from the binary data format to the non-binary data format via a per bit calculation.
  7. The performance-enhanced computing system of any one of claims 1 to 6, wherein the radix of the non-binary data format is to generate overlapping outputs in the second analog signals.
  8. A semiconductor apparatus comprising:
    one or more substrates; and
    logic coupled to the one or more substrates, wherein the logic is implemented at least partly in one or more of configurable or fixed-functionality hardware, the logic including:
    a memory array to store multibit weight data in a non-binary data format, wherein the non-binary data format includes a radix that is less than two; and
    a capacitor recombination network to conduct multiply-accumulate (MAC) operations on first analog signals and the multibit weight data, the capacitor recombination network further to output second analog signals based on the MAC operations.
  9. The semiconductor apparatus of claim 8, wherein the logic further includes a conversion layer to convert the multibit weight data from a binary data format to the non-binary data format.
  10. The semiconductor apparatus of claim 9, wherein the capacitor recombination network is a C-2C ladder, wherein capacitor pairs in the C-2C ladder include a capacitance ratio that is greater than two, and wherein the conversion layer is to determine the radix of the non-binary data format based on the capacitance ratio of the capacitor pairs.
  11. The semiconductor apparatus of claim 9, wherein the logic further includes a control circuit coupled to the conversion layer and the memory array, and wherein the control circuit is to write the multibit weight data to the memory array in the non-binary data format.
  12. The semiconductor apparatus of claim 9, wherein the conversion layer is to convert the multibit weight data from the binary data format to the non-binary data format via a lookup table.
  13. The semiconductor apparatus of claim 9, wherein the conversion layer is to convert the multibit weight data from the binary data format to the non-binary data format via a per bit calculation.
  14. The semiconductor apparatus of any one of claims 8 to 13, wherein the radix of the non-binary data format is to generate overlapping outputs in the second analog signals.
  15. The semiconductor apparatus of any one of claims 8 to 13, wherein the logic coupled to the one or more substrates includes transistor regions that are positioned within the one or more substrates.
  16. A method of operating a performance-enhanced computing system, the method comprising:
    storing multibit weight data to a memory array in a non-binary data format, wherein the non-binary data format includes a radix that is less than two; and
    conducting, by a capacitor recombination network, multiply-accumulate (MAC) operations on first analog signals and the multibit weight data; and
    outputting, by the capacitor recombination network, second analog signals based on the MAC operations.
  17. The method of claim 16, further including converting, by a conversion layer, the multibit weight data from a binary data format to the non-binary data format.
  18. The method of claim 17, wherein the capacitor recombination network is a C-2C ladder, wherein capacitor pairs in the C-2C ladder include a capacitance ratio that is greater than two, and wherein the method further includes determining, by the conversion layer, the radix of the non-binary data format based on the capacitance ratio of the capacitor pairs.
  19. The method of claim 17, further including writing, by a control circuit coupled to the conversion layer and the memory array, the multibit weight data to the memory array in the non-binary data format, wherein the multibit weight data is converted from the binary data format to the non-binary data format via one or more of a lookup table or a per bit calculation.
  20. The method of any one of claims 16 to 19, wherein the radix of the non-binary data format generates overlapping outputs in the second analog signals.
PCT/CN2024/093798 2024-05-17 2024-05-17 Multibit analog compute-in-memory using sub-2 radix data format Pending WO2025236263A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/CN2024/093798 WO2025236263A1 (en) 2024-05-17 2024-05-17 Multibit analog compute-in-memory using sub-2 radix data format

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2024/093798 WO2025236263A1 (en) 2024-05-17 2024-05-17 Multibit analog compute-in-memory using sub-2 radix data format

Publications (1)

Publication Number Publication Date
WO2025236263A1 true WO2025236263A1 (en) 2025-11-20

Family

ID=97719177

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/093798 Pending WO2025236263A1 (en) 2024-05-17 2024-05-17 Multibit analog compute-in-memory using sub-2 radix data format

Country Status (1)

Country Link
WO (1) WO2025236263A1 (en)

Citations (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5638071A (en) * 1994-09-23 1997-06-10 National Semiconductor Corporation Efficient architecture for correcting component mismatches and circuit nonlinearities in A/D converters
US20130021181A1 (en) * 2011-07-22 2013-01-24 Texas Instuments Incorporated Non-binary successive approximation analog to digital converter
US20140247177A1 (en) * 2013-03-01 2014-09-04 Infineon Technologies Ag Data conversion with redundant split-capacitor arrangement
US9059734B1 (en) * 2013-09-10 2015-06-16 Maxim Integrated Products, Inc. Systems and methods for capacitive digital to analog converters
US9531400B1 (en) * 2015-11-04 2016-12-27 Avnera Corporation Digitally calibrated successive approximation register analog-to-digital converter
US20180034479A1 (en) * 2016-07-29 2018-02-01 Western Digital Technologies, Inc. Non-binary encoding for non-volatile memory
US20200242474A1 (en) * 2019-01-24 2020-07-30 Microsoft Technology Licensing, Llc Neural network activation compression with non-uniform mantissas
CN115904311A (en) * 2021-09-24 2023-04-04 英特尔公司 Analog multiply-accumulate unit for computations in multi-bit memory cells
US20230289066A1 (en) * 2023-03-22 2023-09-14 Intel Corporation Reconfigurable multibit analog in-memory computing with compact computation

Patent Citations (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5638071A (en) * 1994-09-23 1997-06-10 National Semiconductor Corporation Efficient architecture for correcting component mismatches and circuit nonlinearities in A/D converters
US20130021181A1 (en) * 2011-07-22 2013-01-24 Texas Instuments Incorporated Non-binary successive approximation analog to digital converter
US20140247177A1 (en) * 2013-03-01 2014-09-04 Infineon Technologies Ag Data conversion with redundant split-capacitor arrangement
US9059734B1 (en) * 2013-09-10 2015-06-16 Maxim Integrated Products, Inc. Systems and methods for capacitive digital to analog converters
US9531400B1 (en) * 2015-11-04 2016-12-27 Avnera Corporation Digitally calibrated successive approximation register analog-to-digital converter
US20180034479A1 (en) * 2016-07-29 2018-02-01 Western Digital Technologies, Inc. Non-binary encoding for non-volatile memory
US20200242474A1 (en) * 2019-01-24 2020-07-30 Microsoft Technology Licensing, Llc Neural network activation compression with non-uniform mantissas
CN115904311A (en) * 2021-09-24 2023-04-04 英特尔公司 Analog multiply-accumulate unit for computations in multi-bit memory cells
US20230289066A1 (en) * 2023-03-22 2023-09-14 Intel Corporation Reconfigurable multibit analog in-memory computing with compact computation

Similar Documents

Publication Publication Date Title
US9742424B2 (en) Analog-to-digital converter
WO2020139895A1 (en) Circuits and methods for in-memory computing
US11762700B2 (en) High-energy-efficiency binary neural network accelerator applicable to artificial intelligence internet of things
US10979065B1 (en) Signal processing circuit, in-memory computing device and control method thereof
CN111934688A (en) Successive approximation type analog-to-digital converter and method
WO2021056980A1 (en) Convolutional neural network oriented two-phase coefficient adjustable analog multiplication calculation circuit
US20240396568A1 (en) Adaptive analog partial sum accumulation technology for energy-efficient compute-in-memory
US20200043557A1 (en) Nand flash memory with reconfigurable neighbor assisted llr correction with downsampling and pipelining
US20230251943A1 (en) Row repair and accuracy improvements in analog compute-in-memory architectures
US20220262426A1 (en) Memory System Capable of Performing a Bit Partitioning Process and an Internal Computation Process
WO2024103480A1 (en) Computing-in-memory circuit and chip, and electronic device
WO2020020092A1 (en) Digital to analog converter
US12504721B2 (en) Energy efficient digital to time converter (DTC) for edge computing
CN115664422B (en) A distributed successive approximation analog-to-digital converter and its operation method
CN118645135A (en) Storage and calculation circuit, control method thereof, and chip
Lin et al. An 11t1c bit-level-sparsity-aware computing-in-memory macro with adaptive conversion time and computation voltage
US20230289066A1 (en) Reconfigurable multibit analog in-memory computing with compact computation
US6747588B1 (en) Method for improving successive approximation analog-to-digital converter
CN110855295B (en) Digital-to-analog converter and control method
Jeong et al. HYTEC: Compact and Energy-Efficient Analog-Digital Hybrid CIM With Transpose Ternary eDRAM
CN114330694B (en) Circuit and method for implementing convolution operation
US11764801B2 (en) Computing-in-memory circuit
US20240113725A1 (en) Embedded sar-adc with least significant bit skipping based relu activation function
US10778240B1 (en) Device and method for digital to analog conversion
US20240340019A1 (en) Stable Low-Power Analog-to-Digital Converter (ADC) Reference Voltage

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24938287

Country of ref document: EP

Kind code of ref document: A1