WO2025189339A1 - Reshaping convolution based on configuration of deep neural network accelerator - Google Patents

Reshaping convolution based on configuration of deep neural network accelerator

Info

Publication number
WO2025189339A1
WO2025189339A1 PCT/CN2024/081131 CN2024081131W WO2025189339A1 WO 2025189339 A1 WO2025189339 A1 WO 2025189339A1 CN 2024081131 W CN2024081131 W CN 2024081131W WO 2025189339 A1 WO2025189339 A1 WO 2025189339A1
Authority
WO
WIPO (PCT)
Prior art keywords
tensor
activation
convolution
mac
reshaped
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2024/081131
Other languages
French (fr)
Inventor
Haiyun HONG
Yuanyuan Li
Xu QIAN
Peiqing Jiang
Yinghao Li
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Intel Corp
Original Assignee
Intel Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Intel Corp filed Critical Intel Corp
Priority to PCT/CN2024/081131 priority Critical patent/WO2025189339A1/en
Publication of WO2025189339A1 publication Critical patent/WO2025189339A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0464Convolutional networks [CNN, ConvNet]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0495Quantised networks; Sparse networks; Compressed networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/06Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons
    • G06N3/063Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons using electronic means
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/084Backpropagation, e.g. using gradient descent

Definitions

  • This disclosure relates generally to neural networks (also referred to as “deep neural networks” or “DNN” ) , and more specifically, reshaping convolutions based on configurations of DNN accelerators.
  • DNN deep neural networks
  • DNNs are used extensively for a variety of artificial intelligence applications ranging from computer vision to speech recognition and natural language processing due to their ability to achieve high accuracy.
  • the high accuracy comes at the expense of significant computation cost.
  • DNNs have extremely high computing demands as there can be hundreds of millions of MAC (multiply-accumulate) operations as well as a large amount of data to read and write. Therefore, techniques to improve efficiency of DNNs are needed.
  • FIG. 1 illustrates an example DNN, in accordance with various embodiments.
  • FIG. 2 illustrates an example convolution, in accordance with various embodiments.
  • FIG. 3 is a block diagram of a DNN system, in accordance with various embodiments.
  • FIG. 4 illustrates an example data processing cell, in accordance with various embodiments.
  • FIG. 5 illustrates an example data processing unit, in accordance with various embodiments.
  • FIG. 6 is a block diagram of a DNN module, in accordance with various embodiments.
  • FIG. 7 illustrates a convolution before reshaping, in accordance with various embodiments.
  • FIG. 8 illustrates a reshaped convolution, in accordance with various embodiments.
  • FIG. 9 illustrates reshaping of an activation tensor in a channel, in accordance with various embodiments.
  • FIG. 10 illustrates computing a channel of an output tensor of a convolution without reshaping, in accordance with various embodiments.
  • FIG. 11 illustrates computing a channel of an output tensor of a reshaped convolution, in accordance with various embodiments.
  • FIG. 12 illustrates an example process of improving convolution compute efficiency, in accordance with various embodiments.
  • FIG. 13 is a flowchart showing a method of executing a convolution, in accordance with various embodiments.
  • FIG. 14 is a block diagram of an example computing device, in accordance with various embodiments.
  • DNNs are widely used in the domains of computer vision, speech recognition, image, and video processing mainly due to their ability to achieve beyond human-level accuracy.
  • AI artificial intelligence
  • a DNN layer may include one or more deep learning operations (also referred to as “neural network operations” ) , such as matrix multiplication, convolution, pooling, elementwise operation, linear operation, nonlinear operation, and so on.
  • Input or output data of deep learning operations may be arranged in data structures called tensors.
  • a tensor is a data structure having multiple elements across one or more dimensions. Examples of tensors include vector (which is one-dimensional (1D) tensor) , matrix (which is two-dimensional (2D) tensor) , three-dimensional (3D) tensors, four-dimensional (4D) tensors, and even higher dimensional tensors.
  • a dimension of a tensor may correspond to an axis, e.g., an axis in a coordinate system. A dimension may be measured by the number of data points along the axis. The dimensions of a tensor may define the shape of the tensor.
  • a DNN layer may receive one or more input tensors and compute an output tensor from the one or more input tensors.
  • the input tensors include an activation tensor (also referred to as “input feature map (IFM) ” ) including one or more activations (also referred to as “input elements” ) and a weight tensor.
  • the weight tensor may be a kernel (a 2D weight tensor) , a filter (a 3D weight tensor) , or a group of filters (a 4D weight tensor) .
  • the workload of performing a deep learning operation may be visualized as a nested loop structure having a 4D spatial shape. This structure spans across the output channel, spatial height, spatial width, and input channel, which are the four dimensions of the spatial shape.
  • the innermost loop may represent a hardware-unrolled stencil computation.
  • the spatial height and spatial width may be the height and width, respectively, of the output tensor. For instance, the spatial height and spatial width may be the height and width of a 2D matrix corresponding to a single channel in the output tensor.
  • DNN accelerators are typically processors designed to expedite computations in DNNs.
  • AI accelerator or “AI processor”
  • the size of DNN models e.g., the number of internal parameters
  • the diversity of computation workloads causes challenges in the design of DNN accelerators, such as efficiently running a variety of models on a single hardware system, accommodating a broad range of model sizes, managing operational intensity (measured in operations per byte) , meeting latency and power requirements, and so on.
  • hardware needs to support flexible and various stencil settings, while software can choose the optimal setting based on cost metric or workload’s pattern.
  • Currently available DNN accelerators can support various fixed stencil settings.
  • Stencil refers to the fundamental computational unit within a single computational cycle.
  • Such fixed stencil settings can offer certain level of flexibility, allowing for optimal performance across various workload shapes.
  • the mapping of these settings is usually managed by a deep learning compiler or software stack.
  • the stencil settings supported by these DNN accelerators are often fixed and limited. This is usually due to the complexities of hardware implementation, the need for efficient operation, and the extensive validation efforts required. Even though stencil settings can provide some level of adaptability, their range is constrained by these practical considerations.
  • a DNN accelerator may include one or more MAC arrays that perform deep learning operations, such as convolutions, matrix multiplications, and so on.
  • a MAC array is typically organized into a two-dimensional array of MAC units that perform MAC operations.
  • the fixed stencil settings of the DNN accelerator cause the stencils processed by the MAC units to adhere to certain patterns, such as 4x4x16x8, 8x2x16x8, and so on.
  • a MAC array may process a 4x4x16x8 output subtensor within one computational cycle, meaning the MAC unit can compute an output subtensor that spans across 16 output channels, 4 spatial heights, 4 spatial widths, and 8 input channels in a single cycle.
  • a workload having a size of 16x16x32x8 would take 32 (4x4x2x1) cycles, and the utilization of the MAC units would be 100%.
  • Some workload may have spatial shapes that do not align with any of the stencil settings of the DNN accelerator. For instance, a dimension of the spatial shape of a workload may not be a multiple of the corresponding dimension of any stencil setting.
  • Such “stencil-unfriendly” workloads can suffer from low utilization of the MAC units.
  • a workload with a size of 1x16x32x8 would also take 32 cycles, and the utilization of the MAC units would be reduced to 25%.
  • Many DNNs models may have stencil-unfriendly workloads.
  • workloads for executing Transformer-based models may feature a kernel size of [1, 1] or unbalanced spatial shapes like [1, n] or [n, 1] . It can be difficult to design one chip with fixed stencil to run diverse workloads efficiently.
  • Embodiments of the present disclosure may improve on at least some of the challenges and issues described above by reshaping deep learning operations in DNNs based on configuration of MAC arrays in DNN accelerators.
  • the deep learning operations may include convulsions, matrix multiplications, and so on.
  • an input tensor of a deep learning operation may be reshaped based on the configuration of the DNN accelerator running the deep learning operation to align the spatial shape of the workload for performing the deep learning operation with a stencil setting of the DNN accelerator. That way, despite that the stencil settings of DNN accelerators are fixed and limited, the computational efficiency of the DNN accelerators can reach a desirable level for workloads with various spatial shapes, including stencil-unfriendly workloads.
  • a DNN accelerator may include one or more tiles.
  • a tile may be referred to as a compute block, which may include a data processing unit (aka neural processing unit) .
  • a data processing unit may include one or more MAC arrays, each of which have MAC units arranged in one or more columns and one or more rows.
  • the number of MAC unit (s) in a column of a MAC array may indicate the height of the MAC array.
  • the number of MAC unit (s) in a row of the array may indicate the width of the MAC array.
  • a deep learning operation may be performed by at least one MAC array and may be reshaped based on the height or width of a MAC array to increase utilization of the MAC units in the process of performing the deep learning operation.
  • a convolution may have an activation tensor and a weight tensor, which are inputs to the convolution.
  • the output of the convolution may be an output tensor, which may not be aligned with the configuration (e.g., the stencil setting) of the DNN accelerator.
  • the height or width of the output tensor may not be a multiple of the height or width of the MAC array, which can result in undesirable utilization of MAC units in the MAC array.
  • the convolution, before being performed, may be transformed into a convolution aligned with the configuration of the DNN accelerator by reshaping the activation tensor.
  • the activation tensor may have a 3D shape defined by a height measured by the number of activations in a column, a width measured by the number of weights in a row, and a depth measured by the number of input channels of the convolution.
  • one or more dimensions of the activation tensor may be modified.
  • the height or width of the activation tensor of the convolution may be modified so that the height or width of the output tensor would be a multiple of the height or width of the MAC array.
  • the height or width of the activation tensor may be changed to become a multiple of the height or width of the MAC array.
  • the MAC array may compute the output tensor of the reshaped convolution.
  • the computational efficiency for performing the reshaped convolution can be better than the computational efficiency for performing the original condition as the utilization of the MAC units would be better.
  • the output tensor of the reshaped convolution may then be reshaped to generate the output tensor of the original convolution.
  • the present disclosure provides a spatial reshape mechanism that can optimize utilization of MAC units in DNN accelerators with fixed stencil settings.
  • the reshaping may introduce little or even no data movement overhead or computation overhead.
  • many stencil-unfriendly workloads such as workloads in Transformer-based models and 1D convolutions in audio applications, can achieve similar or even the same computational efficiency as convolution-based models. Therefore, desirable hardware utilization and efficiency can be achieved even for stencil-unfriendly workloads.
  • the phrase “A or B” or the phrase “A and/or B” means (A) , (B) , or (A and B) .
  • the phrase “A, B, or C” or the phrase “A, B, and/or C” means (A) , (B) , (C) , (A and B) , (A and C) , (B and C) , or (A, B, and C) .
  • the terms “comprise, ” “comprising, ” “include, ” “including, ” “have, ” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion.
  • a method, process, device, or DNN accelerator that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such method, process, device, or DNN accelerators.
  • the term “or” refers to an inclusive “or” and not to an exclusive “or. ”
  • FIG. 1 illustrates an example DNN 100, in accordance with various embodiments.
  • the DNN 100 may be executed by a DNN accelerator, e.g., the DNN accelerator 302 in FIG. 3.
  • the DNN 100 may be a convolution-based DNN.
  • the DNN 100 may be other types of DNNs.
  • the DNN 100 includes a sequence of layers comprising a plurality of convolutional layers 110 (individually referred to as “convolutional layer 110” ) , a plurality of pooling layers 120 (individually referred to as “pooling layer 120” ) , and a plurality of fully-connected layers 130 (individually referred to as “fully-connected layer 130” ) .
  • the DNN 100 may include fewer, more, or different layers.
  • the layers of the DNN 100 execute tensor computation that includes many tensor operations, such as matrix multiplications, convolutions (e.g., multiply-accumulate (MAC) operations, etc. ) , pooling operations, elementwise operations (e.g., elementwise addition, elementwise multiplication, etc. ) , other types of tensor operations, or some combination thereof.
  • tensor computation that includes many tensor operations, such as matrix multiplications, convolutions (e.g., multiply-accumulate (MAC) operations, etc. ) , pooling operations, elementwise operations (e.g., elementwise addition, elementwise multiplication, etc. ) , other types of tensor operations, or some combination thereof.
  • MAC multiply-accumulate
  • the convolutional layers 110 summarize the presence of features in inputs to the DNN 100.
  • the convolutional layers 110 function as feature extractors.
  • the first layer of the DNN 100 is a convolutional layer 110.
  • a convolutional layer 110 performs a convolution on an input tensor 140 (also referred to as IFM 140) and a filter 150.
  • IFM 140 is represented by a 7 ⁇ 7 ⁇ 3 three-dimensional (3D) matrix.
  • the IFM 140 includes 3 input channels, each of which is represented by a 7 ⁇ 7 two-dimensional (2D) matrix.
  • the 7 ⁇ 7 2D matrix includes 7 input elements (also referred to as input points) in each row and 7 input elements in each column.
  • the filter 150 is represented by a 3 ⁇ 3 ⁇ 3 3D matrix.
  • the filter 150 includes 3 kernels, each of which may correspond to a different input channel of the IFM 140.
  • a kernel is a 2D matrix of weights, where the weights are arranged in columns and rows.
  • a kernel can be smaller than the IFM.
  • each kernel is represented by a 3 ⁇ 3 2D matrix.
  • the 3 ⁇ 3 kernel includes 3 weights in each row and 3 weights in each column. Weights can be initialized and updated by backpropagation using gradient descent. The magnitudes of the weights can indicate importance of the filter 150 in extracting features from the IFM 140.
  • the convolution includes MAC operations with the input elements in the IFM 140 and the weights in the filter 150.
  • the convolution may be a standard convolution 163 or a depthwise convolution 183. In the standard convolution 163, the whole filter 150 slides across the IFM 140. All the input channels are combined to produce an output tensor 160 (also referred to as OFM 160) .
  • the OFM 160 is represented by a 5 ⁇ 5 2D matrix.
  • the 5 ⁇ 5 2D matrix includes 5 output elements (also referred to as output points) in each row and 5 output elements in each column.
  • the standard convolution includes one filter in the embodiments of FIG. 1. In embodiments where there are multiple filters, the standard convolution may produce multiple output channels in the OFM 160.
  • the multiplication applied between a kernel-sized patch of the IFM 140 and a kernel may be a dot product.
  • a dot product is the elementwise multiplication between the kernel-sized patch of the IFM 140 and the corresponding kernel, which is then summed, always resulting in a single value. Because it results in a single value, the operation is often referred to as the “scalar product. ”
  • Using a kernel smaller than the IFM 140 is intentional as it allows the same kernel (set of weights) to be multiplied by the IFM 140 multiple times at different points on the IFM 140. Specifically, the kernel is applied systematically to each overlapping part or kernel-sized patch of the IFM 140, left to right, top to bottom.
  • the result from multiplying the kernel with the IFM 140 one time is a single value.
  • the multiplication result is a 2D matrix of output elements.
  • the 2D output matrix (i.e., the OFM 160) from the standard convolution 163 is referred to as an OFM.
  • the depthwise convolution 183 produces a depthwise output tensor 180.
  • the depthwise output tensor 180 is represented by a 5 ⁇ 5 ⁇ 3 3D matrix.
  • the depthwise output tensor 180 includes 3 output channels, each of which is represented by a 5 ⁇ 5 2D matrix.
  • the 5 ⁇ 5 2D matrix includes 5 output elements in each row and 5 output elements in each column.
  • Each output channel is a result of MAC operations of an input channel of the IFM 140 and a kernel of the filter 150.
  • the first output channel (patterned with dots) is a result of MAC operations of the first input channel (patterned with dots) and the first kernel (patterned with dots)
  • the second output channel (patterned with horizontal strips) is a result of MAC operations of the second input channel (patterned with horizontal strips) and the second kernel (patterned with horizontal strips)
  • the third output channel (patterned with diagonal stripes) is a result of MAC operations of the third input channel (patterned with diagonal stripes) and the third kernel (patterned with diagonal stripes) .
  • the number of input channels equals the number of output channels, and each output channel corresponds to a different input channel.
  • the input channels and output channels are referred to collectively as depthwise channels.
  • a pointwise convolution 193 is then performed on the depthwise output tensor 180 and a 1 ⁇ 1 ⁇ 3 tensor 190 to produce the OFM 160.
  • the OFM 160 is then passed to the next layer in the sequence.
  • the OFM 160 is passed through an activation function.
  • An example activation function is rectified linear unit (ReLU) .
  • ReLU is a calculation that returns the value provided as input directly, or the value zero if the input is zero or less.
  • the convolutional layer 110 may receive several images as input and calculate the convolution of each of them with each of the kernels. This process can be repeated several times.
  • the OFM 160 is passed to the subsequent convolutional layer 110 (i.e., the convolutional layer 110 following the convolutional layer 110 generating the OFM 160 in the sequence) .
  • the subsequent convolutional layers 110 perform a convolution on the OFM 160 with new kernels and generate a new feature map.
  • the new feature map may also be normalized and resized.
  • the new feature map can be kernelled again by a further subsequent convolutional layer 110, and so on.
  • a convolutional layer 110 has four hyperparameters: the number of kernels, the size F kernels (e.g., a kernel is of dimensions F ⁇ F ⁇ D pixels) , the S step with which the window corresponding to the kernel is dragged on the image (e.g., a step of one means moving the window one pixel at a time) , and the zero-padding P (e.g., adding a black contour of P pixels thickness to the input image of the convolutional layer 110) .
  • the convolutional layers 110 may perform various types of convolutions, such as 2-dimensional convolution, dilated or atrous convolution, spatial separable convolution, depthwise separable convolution, transposed convolution, and so on.
  • the DNN 100 includes 16 convolutional layers 110. In other embodiments, the DNN 100 may include a different number of convolutional layers.
  • the pooling layers 120 down-sample feature maps generated by the convolutional layers, e.g., by summarizing the presence of features in the patches of the feature maps.
  • a pooling layer 120 is placed between two convolution layers 110: a preceding convolutional layer 110 (the convolution layer 110 preceding the pooling layer 120 in the sequence of layers) and a subsequent convolutional layer 110 (the convolution layer 110 subsequent to the pooling layer 120 in the sequence of layers) .
  • a pooling layer 120 is added after a convolutional layer 110, e.g., after an activation function (e.g., ReLU, etc. ) has been applied to the OFM 160.
  • an activation function e.g., ReLU, etc.
  • a pooling layer 120 receives feature maps generated by the preceding convolution layer 110 and applies a pooling operation to the feature maps.
  • the pooling operation reduces the size of the feature maps while preserving their important characteristics. Accordingly, the pooling operation improves the efficiency of the DNN and avoids over-learning.
  • the pooling layers 120 may perform the pooling operation through average pooling (calculating the average value for each patch on the feature map) , max pooling (calculating the maximum value for each patch of the feature map) , or a combination of both.
  • the size of the pooling operation is smaller than the size of the feature maps.
  • the pooling operation is 2 ⁇ 2 pixels applied with a stride of two pixels, so that the pooling operation reduces the size of a feature map by a factor of 2, e.g., the number of pixels or values in the feature map is reduced to one quarter the size.
  • a pooling layer 120 applied to a feature map of 6 ⁇ 6 results in an output pooled feature map of 3 ⁇ 3.
  • the output of the pooling layer 120 is inputted into the subsequent convolution layer 110 for further feature extraction.
  • the pooling layer 120 operates upon each feature map separately to create a new set of the same number of pooled feature maps.
  • the fully-connected layers 130 are the last layers of the DNN.
  • the fully-connected layers 130 may be convolutional or not.
  • the fully-connected layers 130 receive an input operand.
  • the input operand defines the output of the convolutional layers 110 and pooling layers 120 and includes the values of the last feature map generated by the last pooling layer 120 in the sequence.
  • the fully-connected layers 130 apply a linear combination and an activation function to the input operand and generate a vector.
  • the vector may contain as many elements as there are classes: element i represents the probability that the image belongs to class i. Each element is therefore between 0 and 1, and the sum of all is worth one.
  • probabilities are calculated by the last fully-connected layer 130 by using a logistic function (binary classification) or a SoftMax function (multi-class classification) as an activation function.
  • FIG. 2 illustrates an example convolution, in accordance with various embodiments.
  • the convolution may be a deep learning operation in a convolutional layer of a DNN, e.g., a convolutional layer 110 in FIG. 1.
  • the convolution can be executed on an activation tensor 210 and filters 220 (individually referred to as “filter 220” ) .
  • the filters may constitute a weight tensor of the convolution.
  • the result of the convolution is an output tensor 230.
  • the convolution is performed by a DNN accelerator.
  • An example of the DNN accelerator may be the DNN accelerator 302 in FIG. 3.
  • the convolution may be performed by the sparse cell array 370 in the DNN accelerator 302.
  • the activation tensor 210 includes activations (also referred to as “input activations, ” “elements, ” or “input elements” ) arranged in a 3D matrix.
  • the activation tensor 210 may also be referred to as an input tensor of the convolution.
  • An input element is a data point in the activation tensor 210.
  • the activation tensor 210 has a spatial size of 7 ⁇ 7 ⁇ 3, i.e., the activation tensor 210 includes three input channels and each input channel has a 7 ⁇ 7 2D matrix.
  • Each input element in the activation tensor 210 may be represented by a (X, Y, Z) coordinate. In other embodiments, the height, width, or depth of the activation tensor 210 may be different.
  • each filter 220 in FIG. 2 has a spatial size of 2 ⁇ 3 ⁇ 3, i.e., the filter 220 includes 2 convolutional kernels with a spatial size of 2 ⁇ 3.
  • the height, width, or depth of the filter 220 may be different.
  • the spatial size of the convolutional kernels is smaller than the spatial size of the 2D matrix of each input channel in the activation tensor 210.
  • An activation or weight may take one or more bytes in a memory.
  • the number of bytes for an activation or weight may depend on the data format. For example, when the activation or weight has an INT8 format, the activation takes one byte. When the activation or weight has a FP16 format, the activation or weight takes two bytes. Other data formats may be used for activations or weights.
  • each filter 220 slides across the activation tensor 210 and generates a 2D matrix for an output channel in the output tensor 230.
  • the 2D matrix has a spatial size of 5 ⁇ 5.
  • the output tensor 230 includes activations (also referred to as “output activations, ” “elements, ” or “output element” ) arranged in a 3D matrix.
  • An output activation is a data point in the output tensor 230.
  • the output tensor 230 has a spatial size H out ⁇ W out ⁇ C out , where H out is the height of the 3D matrix (i.e., the length along the Y axis, which indicates the number of output activations in a column in the 2D matrix of each output channel) , W out is the width of the 3D matrix (i.e., the length along the X axis, which indicates the number of output activations in a row in the 2D matrix of each output channel) , and C out is the depth of the 3D matrix (i.e., the length along the Z axis, which indicates the number of output channels) .
  • C out may equal the number of filters 220 in the convolution.
  • H out and W out may depend on the heights and weights of the activation tensor 210 and each filter 220. In an example where the kernel size is 1 ⁇ 1, H out and W out may equal to H in and W in , respectively.
  • an output activation may include 8 bits, e.g., one byte.
  • an output activation may include more than one byte. For instance, an output element may include two bytes.
  • a vector 235 is produced.
  • the vector 235 is highlighted with slashes in FIG. 2.
  • the vector 235 includes a sequence of output activations, which are arranged along the Z axis.
  • the output activations in the vector 235 have the same (X, Y) coordinate, but the output activations correspond to different output channels and have different Z coordinates.
  • the dimension of the vector 235 along the Z axis may equal the total number of output channels in the output tensor 230.
  • the MAC operations on a 2 ⁇ 3 ⁇ 3 subtensor (e.g., the subtensor 215) and a filter 220 may be performed by a plurality of MAC units.
  • One or more MAC units may receive an input operand (e.g., an activation operand 217 shown in FIG. 2) and a weight operand (e.g., the weight operand 227 shown in FIG. 2) .
  • the activation operand 217 includes a sequence of activations having the same (x, y) coordinate but different z coordinates.
  • the activation operand 217 includes an activation from each of the input channels in the activation tensor 210.
  • the weight operand 227 includes a sequence of weights having the same (x, y) coordinate but different z coordinates.
  • the weight operand 227 includes a weight from each of the channels in the filter 220.
  • Activations in the activation operand 217 and weights in the weight operand 227 may be sequentially fed into a MAC unit.
  • the MAC unit may receive an activation and a weight ( “an activation-weight pair” ) at a time and multiple the activation and the weight.
  • the position of the activation in the activation operand 217 may match the position of the weight in the weight operand 227.
  • the activation and weight may correspond to the same channel.
  • Activations or weights may be floating-point numbers.
  • Floating-point numbers may have various data formats, such as FP32, FP16, BF16, and so on.
  • a floating-point number may be a positive or negative number with a decimal point.
  • a floating-point number may be represented by a sequence of bits that includes one or more bits representing the sign of the floating-point number (e.g., positive or negative) , bits representing an exponent of the floating-point number, and bits representing a mantissa of the floating-point number.
  • the mantissa is the part of a floating-point number that represents the significant digits of that number.
  • the mantissa is multiplied by the base raised to the exponent to give the actual value of the floating-point number.
  • the output activations in the output tensor 230 may be further processed based on one or more activation functions before they are stored or inputted into the next layer of the DNN.
  • the processing based on the one or more activation functions may be at least part of the post processing of the convolution.
  • the post processing may include one or more other computations, such as offset computation, bias computation, and so on.
  • the results of the post processing may be stored in a local memory of the compute block and be used as input to the next DNN layer.
  • the input activations in the activation tensor 210 may be results of post processing of the previous DNN layer.
  • FIG. 3 is a block diagram of a DNN system 300, in accordance with various embodiments.
  • the whole DNN system 300 or a part of the DNN system 300 may be implemented in one or more computing devices, such as the computing device 1400 in FIG. 14.
  • the DNN system 300 can generate and execute DNNs, such as Transformer-based models, convolution-based models, and so on.
  • the DNN system 300 includes a DNN module 301 and a DNN accelerator 302.
  • the DNN system 300 may include multiple DNN modules or multiple DNN accelerators.
  • the DNN module 301 and DNN accelerator 302 may include different types of processing units.
  • the DNN module 301 may be implemented by one or more central processing units (CPUs) .
  • the DNN accelerator 302 may also be referred to as an AI accelerator or an AI processor.
  • the DNN module 301 and DNN accelerator 302 may be implemented in the same chip or separate chips.
  • the DNN module 301 facilitates generation and deployment of DNNs.
  • the DNN module 301 may generate and train DNNs.
  • the DNN module 301 can define the layered architecture of a DNN.
  • the DNN module 301 can also determine the internal parameters of the DNN through a DNN training process.
  • the DNN module 301 may also determine one or more hyperparameters that define how the DNN is trained.
  • An example hyperparameter is a sparsity ratio that defines the sparsity level of one or more deep learning tensors for the DNN.
  • the DNN module 301 may also compress DNNs, e.g., during or after training.
  • the DNN module 301 may prune weights in one or more layers of a DNN by changing nonzero valued weight to zeros.
  • the DNN module 301 may prune weights based on a target weight sparsity ratio.
  • a weight sparsity ratio may be the ratio of the number of zero-valued weights to the total number of weights.
  • the DNN module 301 may prune weight of a layer to achieve a target sparsity ratio after one or more epochs.
  • the DNN module 301 may prevent the pruned weights from changing values during the rest of the training process.
  • the DNN module 301 may allow the pruned weights to change values so that a pruned, zero-valued weight may have a nonzero value after further training.
  • the DNN module 301 may prune weights of the layer again after one or more additional epochs.
  • the DNN module 301 may deploy trained, compressed, or validated DNNs for use in deep learning applications.
  • the DNN module 301 may distribute trained, compressed, or validated DNNs to devices or systems which may use the DNNs to perform tasks (e.g., image classification, motion planning, etc. ) for which the DNNs were trained.
  • the DNN module 301 may facilitate deployment of the DNNs using the DNN accelerator 302. For instance, the DNN module 301 may receive data from a device or system coupled with the DNN system 300 and input the received data (or data generated by the DNN module 301, e.g., based on the received data) into a DNN.
  • the DNN module 301 may generate instructions (e.g., configuration files) that control the operation of the DNN accelerator 302 during the DNN execution.
  • the DNN module 301 may receive an output of the DNN from the DNN accelerator 302.
  • the DNN module 301 may transmit the output of the DNN (or a result of processing the output of the DNN by the DNN module 301) to the device or system.
  • the DNN module 301 may control execution processes of trained, compressed, or validated DNNs.
  • the DNN module 301 may function as a deep learning complier for DNNs executed by the DNN accelerator 302.
  • the DNN module 301 facilitates reshaping deep learning operations to be executed by the DNN accelerator 302 so that the spatial shape of the deep learning operations may be aligned with the configurations (e.g., stencil settings) of the DNN accelerator 302.
  • the DNN module 301 may change the shape of activation tensors of convolutions so that the spatial heights or widths of output tensors may be multiples of the height or width of one or more MAC arrays in the DNN accelerator 302.
  • the DNN module 301 may cause the DNN accelerator 302 to execute reshaped convolutions.
  • the DNN module 301 may receive output tensors of reshaped convolution from the DNN accelerator and reshape the output tensors to generate output tensors of the original convolutions. Certain aspects of the DNN module 301 are provided below in conjunction with FIG. 6.
  • the DNN accelerator 302 executes DNNs provided by the DNN module 301.
  • the DNN accelerator 302 can execute a DNN by running deep learning operations in the DNN.
  • the process of carrying out a deep learning operation is also referred to as a process of executing the deep learning operation or performing the deep learning operation.
  • the execution of the DNN may be for training the DNN or for using the DNN to perform AI tasks.
  • the DNN accelerator 302 includes components designed for optimal efficiency in running convolution-based DNNs.
  • the DNN accelerator 302 includes a memory 310, a DMA (direct memory access) engine 320, and compute blocks 330 (individually referred to as “compute block 330” ) .
  • a compute block 330 may also referred to as a tile of the DNN accelerator 302.
  • each compute block 330 includes a local memory 340, a sparsity mode module 350, a load module 360, a sparse cell array 370 (also referred to as a data processing unit) , and a drain module 380.
  • Some or all the components of the compute block 330 can be implemented on the same chip. In other embodiments, alternative configurations, different or additional components may be included in the compute block 330. Further, functionality attributed to a component of the compute block 330 may be accomplished by a different component included in the compute block 330, a different compute block 330, another component of the DNN accelerator 302, or a different system.
  • a component of the compute block 330 may be implemented in hardware, software, firmware, or some combination thereof.
  • the process of converting a dense tensor to a sparse tensor may be referred to as sparsity encoding.
  • Sparsity encoding may also generate a sparsity tensor.
  • Each element in the sparsity tensor may correspond to a different element in the dense tensor and indicate whether the element in the dense tensor is zero or not.
  • the sparsity tensor may indicate positions of elements of the sparse tensor in the dense tensor.
  • the sparsity tensor may be a sparsity bitmap, each element of which is a bit.
  • a sparse tensor may be converted to a dense tensor through a densifying process, in which one or more zeros may be added to the sparse tensor based on the sparsity tensor.
  • a storage unit may store a single byte, and data larger than a single byte may be stored in storage units with consecutive memory addresses, i.e., adjacent storage units.
  • a storage unit can store an integer number in the INT8 format, versus two storage units may be needed to store a number in the FP16 or BF16 format, which has 16 bits.
  • 16 bits can be transferred from the local memory 340 in a single read cycle. In other embodiments, 16 bits can be transferred from the local memory 340 in multiple read cycles, such as two cycles.
  • the sparsity mode module 350 determines sparsity modes in which the compute block 330 operates to execute DNN layers. For instance, the sparsity mode module 350 may determine whether to accelerate a layer based on weight sparsity, activation sparsity, or both.
  • the sparsity mode module 350 select the sparsity mode for a layer from a group of sparsity modes that includes, for example, combined sparsity mode in which the layer is accelerated based on both weight sparsity and activation sparsity, activation sparsity mode in which the layer is accelerated based on activation sparsity but not based on weight sparsity, weight sparsity mode in which the layer is accelerated based on weight sparsity but not based on activation sparsity, and a dense mode in which the layer is not accelerated based on sparsity.
  • the sparsity mode module 350 may determine the sparsity mode for all the compute blocks 330 that executes the layer.
  • the sparsity mode module 350 may receive configuration parameters from the DNN module 301.
  • a configuration parameter may correspond to a layer and indicate whether to accelerate the layer based on weight sparsity.
  • the sparsity mode module 350 may determine the sparsity mode of the layer based on the configuration parameter.
  • the load module 360 loads data from the local memory 340 to the sparse cell array 370.
  • the load module 360 may read tensors from the local memory 340.
  • the tensors may include sparse activation tensors, sparse weight tensors, activation sparsity tensors, weight sparsity tensors, and so on.
  • the load module 360 may load data based on the sparsity mode determined by the sparsity mode module 350.
  • the load module 360 may select different data to transmit to the sparse cell array 370 in different sparsity modes.
  • the load module 360 may transmit an activation sparsity tensor and a weight sparsity tensor of a layer to the sparse cell array 370 in the combined sparsity mode, while transmit the activation sparsity tensor but not the weight sparsity tensor to the sparse cell array 370 in the activation sparsity mode and transmit the weight sparsity tensor but not the activation sparsity tensor to the sparse cell array 370 in the weight sparsity mode.
  • the load module 360 does not transmit either the activation sparsity tensor or the weight sparsity tensor to the sparse cell array 370.
  • the load module 360 may process (e.g., densify) data stored in the local memory 340 before providing the data to the sparse cell array 370.
  • the load module 360 while operating in the weight sparsity mode, may densify sparse activation tensors to generate dense activation tensors based on corresponding activation sparsity tensors.
  • the load module 360 may add one or more zeros into a sparse activation tensor based on an activation sparsity tensor associated with the sparse activation tensor to generate the dense activation tensor.
  • the dense activation tensor includes one or more elements than the sparse activation tensor.
  • the additional element (s) are zero-valued.
  • the load module 360 may identify one or more elements in the activation sparsity tensor that correspond to the zero-valued element (s) , determine the position of each of the zero-valued element (s) in the dense activation tensor, and insert the zero-valued element (s) into the sparse activation tensor based on the determined positions. After the densification, the load module 360 may transmit the dense activation tensors to the sparse cell array 370. The load module 360 may also transmit corresponding sparse weight tensors and weight sparsity tensors to the sparse cell array 370. Activation sparsity tensor of the dense activation tensors may not be loaded to the sparse cell array 370.
  • the load module 360 while operating in the activation sparsity mode, may densify sparse weight tensors to generate dense weight tensors based on corresponding weight sparsity tensors by inserting zeros into sparse weight tensors.
  • the densification of sparse weight tensors may be similar to the densification of sparse activation tensors described above.
  • the load module 360 may transmit the dense weight tensors to the sparse cell array 370.
  • the load module 360 may also transmit corresponding sparse activation tensors and activation sparsity tensors to the sparse cell array 370. Weight sparsity tensor of the dense weight tensors may not be loaded to the sparse cell array 370.
  • the load module 360 while operating in the dense mode, may densify both sparse weight tensors and sparse activation tensors.
  • the load module 360 may generate the input tensor and weight tensor of the layer and transmit the tensors to the sparse cell array 370 for executing the layer without sparsity acceleration.
  • the sparse cell array 370 may include one or more data processing cells. Each data processing cell may include one or more MAC units that can perform MAC operations. The MAC units in a data processing cell may be arranged in an array that includes rows and columns. The data processing cells may be arranged in one or more rows and one or more columns in the sparse cell array 370. All the MAC units in the sparse cell array 370 may constitute a bigger array that includes more rows and columns. In some embodiments (e.g., embodiments where the compute block 330 executes a convolutional layer) , a computation in an MAC unit may be an MAC operation on an activation operand and a weight operand.
  • the activation operand may be an activation tensor that may include one or more activations in the input tensor of the convolution. Different activations may be in different input channels.
  • the weight operand may be a weight tensor that may include one or more weights in the filter of the convolution. The values of the weights are determined through training the DNN. The weights in the weight operand may be in different input channels.
  • an MAC unit includes one or more multipliers for performing multiplications.
  • An MAC unit may also include one or more accumulators ( “adders” ) for performing accumulations.
  • a column of MAC units is referred to as an MAC column.
  • An MAC column may be associated with one or more MAC lanes.
  • An MAC lane is a path for loading data e.g., by the load module 360, into an MAC column.
  • An MAC lane may be also referred to as a data transmission lane or data loading lane.
  • An MAC column may have multiple MAC lanes.
  • the loading bandwidth of the MAC column is an aggregation of the loading bandwidths of all the MAC lanes associated with the MAC column.
  • MAC lanes With a certain number of MAC lanes, data can be fed into the same number of independent MAC units simultaneously.
  • an MAC column has four MAC lanes for feeding activations or weights into the MAC column and each MAC lane may have a bandwidth of 16 bytes
  • the four MAC lanes can have a total loading bandwidth of 64 bytes.
  • the sparse cell array 370 may be capable of depthwise convolution, standard convolution, or both.
  • an MAC unit may perform an MAC operation that includes a sequence of multiplications for an input operand and a weight operand.
  • Each multiplication in the sequence (also referred to as a cycle) is a multiplication of a different activation in the input operand with a different weight in the weight operand.
  • the activation and weight in the same cycle may correspond to the same channel.
  • the sequence of multiplication produces a product operand that includes a sequence of products.
  • the MAC operation may also include accumulations in which multiple product operands are accumulated to produce an output operand of the MAC unit.
  • the sparse cell array 370 may output multiple output operands at a time, each of which is generated by a different MAC unit.
  • MAC operations may include accumulations across the channels. For instance, as opposed to generating an output operand, a MAC unit may accumulate products across different channels to generate a single output point.
  • the sparse cell array 370 may include sparsity acceleration logic for facilitating sparsity acceleration.
  • each data processing cell in the sparse cell array 370 may include one or more sparsity modules.
  • each MAC column or each MAC row may have a corresponding sparsity module that accelerates MAC operations in the MAC column or MAC row.
  • a sparsity module accelerates computations in the sparse cell array 370 based on sparsity in activations, sparsity in weights, or both.
  • the sparsity module may include a storage unit that stores a sparsity tensor, which may be loaded to the storage unit by the load module 360.
  • the sparsity tensor may be an activation sparsity tensor, a weight sparsity tensor, or a combined sparsity tensor.
  • An activation sparsity tensor may be the sparsity tensor of an activation tensor and has the same number of elements as the activation tensor.
  • An element in the activation sparsity tensor may indicate whether the corresponding element in the activation tensor is zero or not. For instance, a zero-valued in the activation sparsity tensor may indicate that the corresponding element in the activation tensor is zero.
  • a one-valued in the activation sparsity tensor may indicate that the corresponding element in the activation tensor is nonzero.
  • a weight sparsity tensor may be the sparsity tensor of a weight tensor and has the same number of elements as the weight tensor.
  • An element in the weight sparsity tensor may indicate whether the corresponding element in the weight tensor is zero or not. For instance, a zero-valued in the weight sparsity tensor may indicate that the corresponding element in the weight tensor is zero.
  • a one-valued in the weight sparsity tensor may indicate that the corresponding element in the weight tensor is nonzero.
  • the sparsity module may generate a combined sparsity tensor using an activation sparsity tensor and a weight sparsity tensor.
  • the sparsity module may multiply an element of the activation sparsity tensor with a corresponding element of the weight sparsity tensor to compute an element of the combined sparsity tensor.
  • the positions of the three elements in their corresponding sparsity tensors may match.
  • each element in a sparsity tensor may be a bit, and the sparsity tensor may be referred to as a sparsity bitmap.
  • the sparsity module may use the sparsity tensor to identify activations and weights to be used in MAC operations by the MAC units.
  • the sparsity module may identify activations and weights that correspond to nonzero valued elements of a combined sparsity tensor.
  • the sparsity module may identify activations and weights that correspond to nonzero valued elements of an activation sparsity tensor.
  • the sparsity module may identify activations and weights that correspond to nonzero valued elements of a weight sparsity tensor.
  • the sparsity module may be bypassed in the dense mode as no sparsity acceleration would be conducted.
  • the drain module 380 drains data from the sparse cell array 370 and writes the data to the local memory 340.
  • the data may be outputs of MAC operations performed by MAC units in the sparse cell array 370.
  • the drain module 380 may drain data on a cell level. For each data processing cell, the drain module 380 may drain outputs of MAC units in the data processing cell based on a row index or column index of each MAC unit. For instance, the drain module 380 may use a sequence of cycles to drain data from a data processing cell.
  • the drain module 380 may drain the output of some of the MAC units in each cycle. The sequence of the cycles may be configured based on a configuration parameter indicating the operation mode of the load module 360.
  • the drain module 380 may determine whether to drain the output of an MAC unit based on the column index of the MAC unit when the load module operates in the activation sparsity mode versus based on the row index of the MAC unit when the load module operates in the weight sparsity mode. For instance, for MAC operations where the load module 360 operates in the activation sparsity mode, the drain module 380 may drain the output of a different MAC column in each cycle. The sequence of cycles may start with the first MAC column (e.g., the MAC column on the left side of the data processing cell) and end with the last MAC column (e.g., the MAC column on the right side of the data processing cell) .
  • the first MAC column e.g., the MAC column on the left side of the data processing cell
  • the last MAC column e.g., the MAC column on the right side of the data processing cell
  • the drain module 380 may drain the output of a different MAC row in each cycle.
  • the sequence of cycles may start with the first MAC row (e.g., the MAC row at the top of the data processing cell) and end with the last MAC row (e.g., the MAC column at the bottom of the data processing cell) .
  • the drain module 380 may determine whether to drain the output of an MAC unit based on the row index of the MAC unit when the load module operates in the activation sparsity mode versus based on the column index of the MAC unit when the load module operates in the weight sparsity mode.
  • the drain module 380 may also include sparsity encoding logic that can convert outputs of the sparse cell array 370 from a dense format to a sparse format.
  • the drain module 380 may be implemented with one or more sparsity encoders.
  • a sparsity encoder converts dense data to compressed data based on sparsity in the dense data.
  • the sparsity encoder may remove zeros in an activation tensor computed by the sparse cell array 370 to convert the activation tensor to a compressed activation tensor.
  • the sparsity encoder may also generate sparsity tensors, including activation sparsity tensors.
  • the data drained from the sparse cell array 370 may be at least part of an output tensor (e.g., the output tensor 230 in FIG. 2) of a deep learning operation.
  • the sparsity encoder may generate a compressed version of the output tensor.
  • the sparsity encoder may identify every zero-valued activation in the output tensor and remove these activations from the output tensor to generate a compressed activation tensor (aka “sparse activation tensor” ) .
  • the sparsity encoder may also generate one or more sparsity tensors for the output tensor.
  • a sparsity tensor may correspond to a portion of the output tensor (e.g., the vector 235 in FIG. 2) .
  • the sparsity tensor may include sparsity elements (e.g., bits) , each of which corresponds to a different activation in the vector and indicates whether the corresponding activation is zeroed or not.
  • the weight register files 420 store weights to be processed in MAC operations.
  • four weight register files 420 are grouped into a storage set that stores data to be used by a column of MAC units 410.
  • a weight register file 420 may correspond to a MAC unit 410 and store data to be processed by the MAC unit.
  • all the 16 weight register files 420 constitute a weight storage unit.
  • the activation register files 430 stores activations to be processed in MAC operations.
  • four activation register files 430 are grouped into a storage set that stores data to be used by a row of MAC units 410.
  • an activation register file 430 may correspond to a MAC unit 410 and store data to be processed by the MAC unit.
  • all the 16 activation register files 430 constitute an activation storage unit.
  • the row buffers 440 store outputs of the MAC units 410. Each row buffer 440 may drain outputs of a single row of MAC units 410.
  • each sparsity module 460 facilitates dynamic sparsity-based acceleration in the data processing cell 400.
  • each sparsity module 460 includes a sparsity tensor storage unit 465 and a control logic 467.
  • the sparsity tensor storage unit 465 stores combined sparsity tensors.
  • a combined sparsity tensor stored in the sparsity tensor storage unit 465 may correspond to an activation tensor and a weight tensor.
  • a nonzero element in the combined sparsity tensor may correspond to a nonzero activation-weight pair that includes a nonzero activation and a nonzero weight.
  • the position of the nonzero activation in the activation tensor may match the position of the nonzero weight in the weight tensor.
  • the product of the nonzero activation and nonzero weight would be nonzero.
  • the control logic 467 may control transmission of activations and weights stored from the weight register files 420 and the activation register files 430 to the MAC units 410 based on sparsity tensors. For instance, the control logic 467 may select a subset of the weights stored in the weight register files 420 and select a subset of activations stored in the activation register files 430 based on a combined sparsity tensor. The selected weights and activations constitute nonzero activation-weight pairs. The control logic 467 may transmit the selected weights and activations to the MAC units 410 for performing MAC operations. The other weights stored in the weight register files 420 and the other activations stored in the activation register files 430 are skipped from computation. In the embodiments of FIG.
  • each sparsity module 460 controls sparsity acceleration in a respective MAC unit 410.
  • the sparsity acceleration is either based on both weight sparsity and activation sparsity
  • 16 sparsity modules 460 are used for acceleration computations in the 16 MAC units 410.
  • the data processing cell 400 is associated with multiplexers (MUXs) 403, 404, 405, and 406.
  • MUXs multiplexers
  • the MUX 403 facilitates loading weights, e.g., from the local memory 340, into the weight register files 420.
  • the MUX 404 facilitates loading activations, e.g., from the local memory 340, into the activation register files 430.
  • the MUX 405 facilitates loading sparsity tensors into the sparsity tensor storage unit 465.
  • the MUX 406 may be a drain MUX that can facilitate draining outputs of the MAC units 410, e.g., to the local memory 340.
  • the data processing cell 400 may also execute matrix multiplications converted from Fourier transform operations.
  • the MAC units 410 may perform MAC operations in the two sequences of matrix multiplications converted from the Fourier transform operation.
  • the weight register files 420 may be used to store data points in transformation tensor of the Fourier transform operation.
  • the activation register file 430 may be used to store data points in the input tensor of the Fourier transform operation.
  • the row buffers 440 may store data points in the output tensor of the Fourier transform operation.
  • FIG. 5 illustrates a data processing unit 500, in accordance with various embodiments.
  • the data processing unit 500 may be an example of the sparse cell array 370 in FIG. 3.
  • the data processing unit 500 includes data processing cells 510 (individually referred to as “data processing cell 510” ) arranged in four columns and four rows, an activation memory 520, and a weight memory 530.
  • the data processing unit 500 may also be referred to as a data processing unit.
  • the data processing unit 500 may include fewer, more, or different components. For instance, the data processing unit 500 may include a different number of columns, rows, or data processing cells 510.
  • Each data processing cell 510 may perform sparsity accelerated MAC operations.
  • the data processing cells 510 may facilitate dynamic sparsity mode. For instance, the sparsity modes of a data processing cell 510 may be dynamically changed between a combined sparsity mode, an activation sparsity mode, a weight sparsity mode, and a dense mode.
  • An embodiment of a data processing cell 510 may be the data processing cell 400 in FIG. 4.
  • the activation memory 520 stores activations, such as activations in input tensors of deep learning operations. Activations may be loaded from the activation memory 520 to data processing cells 510.
  • the weight memory 530 stores weights, such as weights in filters of deep learning operations. Weights may be loaded from the weight memory 530 to data processing cells 510.
  • the activation memory 520 or weight memory 530 may be a buffer.
  • the data processing unit 500 may include a dense data memory and a sparse data memory in lieu of the activation memory 520 and weight memory 530.
  • the dense data memory may store dense tensors, e.g., dense tensors generated by the load module 360.
  • the sparse data memory may store sparse tensors.
  • the data processing unit 500 may also execute matrix multiplications in Fourier transform operations.
  • the activation memory 520 may be used to store input tensors of the Fourier transform operations.
  • the weight memory 530 may be used to store transformation matrices of the Fourier transform operations.
  • FIG. 6 is a block diagram of a DNN module 600, in accordance with various embodiments.
  • the DNN module 600 facilitates transformation of matrix multiplications to convolutions.
  • the DNN module 600 may be an embodiment of the DNN module 301 in FIG. 3.
  • the DNN module 600 includes an interface module 610, a training module 620, a compressing module 630, a validating module 640, a reshaping module 650, and a datastore 660.
  • the DNN module 600 includes an interface module 610, a training module 620, a compressing module 630, a validating module 640, a reshaping module 650, and a datastore 660.
  • different or additional components may be included in the DNN module 600.
  • functionality attributed to a component of the DNN module 600 may be accomplished by a different component included in the DNN module 600 or a different module or system.
  • the interface module 610 facilitates communications of the DNN module 600 with other modules or systems. For example, the interface module 610 establishes communications between the DNN module 600 with an external database to receive data that can be used to train DNNs or input into DNNs to perform tasks. As another example, the interface module 610 supports the DNN module 600 to distribute DNNs to other systems, e.g., computing devices configured to apply DNNs to perform tasks.
  • the training module 620 trains DNNs by using a training dataset.
  • the training module 620 forms the training dataset.
  • the training dataset includes training images and training labels.
  • the training labels describe ground-truth classifications of objects in the training images.
  • each label in the training dataset corresponds to an object in a training image.
  • a part of the training dataset may be used to initially train the DNN, and the rest of the training dataset may be held back as a validation subset used by the validating module 640 to validate performance of a trained DNN.
  • the portion of the training dataset not including the tuning subset and the validation subset may be used to train the DNN.
  • the training module 620 also determines hyperparameters for training the DNN.
  • Hyperparameters are variables specifying the DNN training process. Hyperparameters are different from parameters inside the DNN (e.g., weights of filters) .
  • hyperparameters include variables determining the architecture of the DNN, such as number of hidden layers, etc.
  • Hyperparameters also include variables which determine how the DNN is trained, such as batch size, number of epochs, etc.
  • a batch size defines the number of training samples to work through before updating the parameters of the DNN. The batch size is the same as or smaller than the number of samples in the training dataset.
  • the training dataset can be divided into one or more batches.
  • the number of epochs defines how many times the entire training dataset is passed forward and backwards through the entire network.
  • the number of epochs defines the number of times that the deep learning algorithm works through the entire training dataset.
  • One epoch means that each training sample in the training dataset has had an opportunity to update the parameters inside the DNN.
  • An epoch may include one or more batches.
  • the number of epochs may be 1, 5, 10, 50, 100, 500, 1000, or even larger.
  • the training module 620 defines the architecture of the DNN, e.g., based on some of the hyperparameters.
  • the architecture of the DNN includes an input layer, an output layer, and a plurality of hidden layers.
  • the input layer of an DNN may include tensors (e.g., a multidimensional array) specifying attributes of the input image, such as the height of the input image, the width of the input image, and the depth of the input image (e.g., the number of bits specifying the color of a pixel in the input image) .
  • the output layer includes labels of objects in the input layer.
  • the hidden layers are layers between the input layer and output layer.
  • the hidden layers include one or more convolutional layers and one or more other types of layers, such as pooling layers, fully-connected layers, normalization layers, SoftMax or logistic layers, and so on.
  • the convolutional layers of the DNN abstract the input image to a feature map that is represented by a tensor specifying the feature map height, the feature map width, and the feature map channels (e.g., red, green, blue images include 3 channels) .
  • a pooling layer is used to reduce the spatial volume of input image after convolution. It is used between two convolution layers.
  • a fully-connected layer involves weights, biases, and neurons. It connects neurons in one layer to neurons in another layer. It is used to classify images between different categories by training.
  • the training module 620 also adds an activation function to a hidden layer or the output layer.
  • An activation function of a layer transforms the weighted sum of the input of the layer to an output of the layer.
  • the activation function may be, for example, a ReLU activation function, a tangent activation function, or other types of activation functions.
  • the training module 620 inputs a training dataset into the DNN.
  • the training dataset includes a plurality of training samples.
  • An example of a training sample includes an object in an image and a ground-truth label of the object.
  • the training module 620 modifies the parameters inside the DNN ( “internal parameters of the DNN” ) to minimize the error between labels of the training objects that are generated by the DNN and the ground-truth labels of the objects.
  • the internal parameters include weights of filters in the convolutional layers of the DNN.
  • the training module 620 uses a cost function to minimize the error.
  • the training module 620 may train the DNN for a predetermined number of epochs.
  • the number of epochs is a hyperparameter that defines the number of times that the deep learning algorithm will work through the entire training dataset.
  • One epoch means that each sample in the training dataset has had an opportunity to update internal parameters of the DNN.
  • the training module 620 may stop updating the parameters in the DNN.
  • the DNN having the updated parameters is referred to as a trained DNN.
  • the compressing module 630 compresses DNNs. For instance, the compressing module 630 may add pruning operations to DNN layers to reduce computational complexity or memory usage. A pruning operation may prune weight tensors of a DNN layer by changing one or more nonzero valued weights of the layer to zeros. The modification may be done before, during, or after training. Weights may be pruned during training, during inference, or a combination of both.
  • the compressing module 630 may determine a sparsity ratio for a DNN layer. The sparsity ratio may be a ratio of the number of zero-valued weight to the total number of weights in the layer. The compressing module 630 may perform the pruning operation till the sparsity ratio of the DNN layer meets a target sparsity ration, such as 10%, 20%, 30%, 60%, 50%, and so on.
  • a target sparsity ration such as 10%, 20%, 30%, 60%, 50%, and so on.
  • the compressing module 630 may select one or more layers in a DNN and modify each selected layer with a pruning operation. For instance, the compressing module 630 may select computationally complex layers, such as layers with large filters. For a pruning operation of a layer or of a type of layer, the compressing module 630 may determine a weight threshold that would not cause a loss of the accuracy of the DNN to exceed an accuracy loss constraint. A pruning operation may modify weights having absolute values above the weight threshold to zeros and leave the other weights unchanged. The weight pruning can reduce memory storage as zero-valued weights may not be stored. Also, the number of operations in the layer can be reduced as computations on zero-valued weights can be skipped without impacting the output of the layer. In some embodiments, the compressing module 630 may also measure energy saving, final DNN accuracy, or layer-wise sparsity caused by pruning operations.
  • the compressing module 630 may fine tune the DNN, e.g., through a retraining process.
  • the compressing module 630 may fine tunes DNNs after weights are pruned.
  • the fine-tuning process is a retraining or further training process. For instance, after weights in a DNN are pruned, the compressing module 630 may further train the DNN by inputting a training dataset into the DNN.
  • the values of the unpruned weights in the DNN may be modified based on outputs of the DNN and ground-truth labels of the training samples in the training dataset.
  • the values of the pruned weights i.e., zero) are not changed during the fine-tuning process.
  • the compressing module 630 may place a mask over a pruned weight block and the mask can prevent values in the pruned weight blocks from being changed during the fine-tuning process.
  • the values of all weights, including the pruned weights may be changed during the fine-tuning process.
  • the compressing module 630 may perform a new pruning process, e.g., by selecting weight blocks and pruning the selected weight blocks.
  • the weight pruning process may be repeated multiple times before the fine-tuning process is done.
  • the number of epochs in the fine-tuning process may be different from the number of epochs in the training process in which the pre-pruning values of the weights are determined.
  • the fine-tuning process may have less epochs than the training process.
  • the number of epochs in the fine-tuning process may be relatively small, such as 2, 3, 6, 5, and so on.
  • the validating module 640 verifies accuracy of trained or compressed DNNs.
  • the validating module 640 inputs samples in a validation dataset into a trained DNN and uses the outputs of the DNN to determine the model accuracy.
  • a validation dataset may be formed of some or all the samples in the training dataset. Additionally or alternatively, the validation dataset includes additional samples, other than those in the training sets.
  • the validating module 640 may determine an accuracy score measuring the precision, recall, or a combination of precision and recall of the DNN.
  • the reshaping module 650 may perform the padding by adding one or more zeros to one or more boundaries of the activation tensor.
  • the reshaping module 650 may add one or more columns (or rows) of zeros into the activation tensor to increase the width (or height) of the activation tensor to make the product of the height and the weight of the output tensor to a multiple of the height or width of the MAC array.
  • the reshaping module 650 may make no change to the number of input channels in the activation tensor. Even though the activation tensor is reshaped in various examples, the reshaping module 650 may reshape the weight tensor in some embodiments.
  • the reshaping module 650 may reshape the activation tensor to change the spatial size of the output tensor in each channel to 16 ⁇ 16 based on a determination that the MAC array is an 8 ⁇ 8 array.
  • the reshaping module 650 may determine the computational efficiency for each of the options and select the option that can achieve the best computational efficiency. In an example where the MAC array is a N ⁇ N grid, the reshaping module 650 may determine the compute efficiency using the following algorithm:
  • ceil denotes the ceil function that returns the smallest integer value that is greater than or equal to the number input into the ceil function
  • W out is the width of the output tensor
  • H out is the height of the output tensor.
  • the reshaping module 650 may determine that the spatial size 16 ⁇ 16 would result in the best computational efficiency and may select to changing the 128 ⁇ 2 spatial size to 16 ⁇ 16, as opposed to 32 ⁇ 8 or 8 ⁇ 32.
  • the reshaping module 650 may also generate one or more configuration parameters that indicate how the activation tensor is reshaped. Such configuration parameters may be referred to as reshaping parameters, which the reshaping module 650 may use to generate the output tensor of the original convolution.
  • the reshaping module 650 may also determine whether to change the stride of the convolution to ensure that the output tensor of the reshaped convolution would have the same elements (even though different arrangements of the elements) as the output tensor of the original convolution.
  • the reshaping module 650 may determine the stride of the reshaped convolution based on the shape of the activation tensor and the shape of the weight tensor.
  • the reshaping module 650 may facilitate the sparse cell array 370 to read activations in the activation tensor in accordance with the new shape of the activation tensor.
  • the sparse cell array 370 may read an activation operand at a time and send the activation operand to a MAC unit for processing in a single computational cycle.
  • the activation operand may also be referred to as a context.
  • An example of the activation operand may be a vector in the activation tensor, e.g., at least part of a row or a column in a channel of the activation tensor.
  • the reshaping module 650 may generate configuration parameters that configure the sparse cell array 370 to read the activation operands in the reshaped activation tensor.
  • the reshaping module 650 may configure storage pointers associated with the memory 310 or the local memory 340.
  • the storage pointers may indicate the boundary of the activation operands, e.g., the memory address where the first activation in every activation operand is stored, the memory address where the last activation in every activation operand is stored, and so on.
  • the sparse cell array 370 may read activation operands based on the storage pointers. The layout of the activations in the memory 310 or the local memory 340 may remain the same.
  • the reshaping module 650 may receive the output tensor of the reshaped convolution from the DNN accelerator 302. The reshaping module 650 may further transform the output tensor of the reshaped convolution to generate the output tensor of the original convolution. For instance, the reshaping module 650 may modify the shape of the output tensor of the reshaped convolution based on the reshaping parameter (s) . In some embodiments, the reshaping module 650 may modify the shape of the output tensor of the reshaped convolution in each output channel. The number of output channels of the reshaped convolution may be the same as the total number of output channels of the original convolution.
  • the datastore 660 stores data received, generated, used, or otherwise associated with the DNN module 600.
  • the datastore 660 stores the datasets used by the training module 620 and validating module 640.
  • the datastore 660 may also store data generated by the training module 620 and validating module 640, such as the hyperparameters for training DNNs, internal parameters of trained DNNs (e.g., weights, etc. ) , data for sparsity acceleration (e.g., sparsity bitmap, etc. ) , and so on.
  • the datastore 660 may store tensors (e.g., transposed tensors, reshaped tensors, etc.
  • the datastore 660 is a component of the DNN module 600. In other embodiments, the datastore 660 may be external to the DNN module 600 and communicate with the DNN module 600 through a network.
  • FIG. 7 illustrates a convolution before reshaping, in accordance with various embodiments.
  • the convolution has an activation tensor 710, a weight tensor 720, and an output tensor 730.
  • the activation tensor 710 has a spatial size H in ⁇ W in ⁇ C in , where H in is the height of the 3D matrix (i.e., the length along the Y axis, which indicates the number of activations in a column in the 3D matrix of each input channel) , W in is the width of the 3D matrix (i.e., the length along the X axis, which indicates the number of activations in a row in the 2D matrix of each input channel) , and C in is the depth of the 3D matrix (i.e., the length along the Z axis, which indicates the number of input channels) .
  • the weight tensor 720 includes multiple filters 725. Each filter 725 includes weights arranged in a 3D matrix. The values of the weights may be determined through training the DNN.
  • a filter 725 has a spatial size H f ⁇ W f ⁇ C f , where H f is the height of the filter (i.e., the length along the Y axis, which indicates the number of weights in a column in each kernel) , W f is the width of the filter (i.e., the length along the X axis, which indicates the number of weights in a row in each kernel) , and C f is the depth of the filter (i.e., the length along the Z axis, which indicates the number of channels) . In some embodiments, C f equals C in .
  • the convolution may be performed by the DNN accelerator 302 in FIG. 3, which processes the activation tensor 710 and weight tensor 720 and computes the output tensor 730.
  • the output tensor 730 has a spatial size H out ⁇ W out ⁇ C out , where H out is the height of the 3D matrix (i.e., the length along the Y axis, which indicates the number of output activations in a column in the 2D matrix of each output channel) , W out is the width of the 3D matrix (i.e., the length along the X axis, which indicates the number of output activations in a row in the 2D matrix of each output channel) , and C out is the depth of the 3D matrix (i.e., the length along the Z axis, which indicates the number of output channels) .
  • C out may equal the number of filters 725 in the weight tensor 720.
  • H out and W out may depend on the heights and weights of the activation
  • the workload of performing the convolution may be represented by a 4D tensor C out ⁇ H out ⁇ W out ⁇ C in .
  • the DNN accelerator performing the convolution may have a N ⁇ N MAC array, meaning the MAC array has N columns and N rows.
  • the compute efficiency of the MAC array may be denoted as:
  • ceil denotes the ceil function that returns the smallest integer value that is greater than or equal to the number input into the ceil function.
  • N is a fixed number
  • the computational efficiency can change as H out or W out changes. For instance, the computational efficiency for a convolution workload with H out and W out each being a multiple of N may be higher than the computational efficiency for a convolution workload with H out or W out not being a multiple of N for computing the same number of output elements. For some convolution workloads, the computational efficiency can be improved by changing H out and W out .
  • FIG. 8 illustrates a reshaped convolution, in accordance with various embodiments.
  • the reshaped convolution may be converted from the convolution in FIG. 7.
  • the convolution in FIG. 7 is referred to as the original convolution.
  • the reshaped convolution has an activation tensor 810, a weight tensor 820, and an output tensor 830.
  • the activation tensor 810 has a spatial size H′ in ⁇ W′ in ⁇ C′ in .
  • the weight tensor 820 includes multiple filters 825. Each filter 825 includes weights arranged in a 3D matrix. The values of the weights may be determined through training the DNN.
  • a filter 825 has a spatial size H′ f ⁇ W′ f ⁇ C′ f .
  • the output tensor 830 has a spatial size H′ out ⁇ W′ out ⁇ C′ out .
  • the activation tensor 810 may be generated by reshaping the activation tensor 710.
  • C′ in may remain the same, but H in and W in are changed.
  • H in ⁇ W in may be the same as H′ in ⁇ W′ in .
  • the weight tensor 820 may be the same as the weight tensor 720, meaning the number of filters 825 in the weight tensor 820 may equal the number of filters 725 in the weight tensor 720 and the shape of each filter 825 may be the same as the shape of each filter 725.
  • the shape of the output tensor 830 may be different from the shape of the output tensor 730.
  • H out is different from H′ out
  • W out may be different from W′ out
  • H′ out ⁇ W′ out may be the same as H out ⁇ W out
  • the workload of performing the convolution is changed from C out ⁇ H out ⁇ W out ⁇ C in to C out ⁇ H′ out ⁇ W′ out ⁇ C in .
  • the reshaping may be done based on the configuration of the MAC array, e.g., based on the height or width of the MAC array.
  • C out ⁇ H out ⁇ W out ⁇ C in is 256x2048x1x48 and the height or width of the MAC array is 4
  • C out ⁇ H′ ou ⁇ W′ tut ⁇ C in is set to 256x128x16x48 to increase utilization of MAC units in the MAC array and increase computation.
  • the reshaped convolution workload would take less computational cycles than the original convolution workload as the utilization of the MAC units in the MAC array would be higher.
  • FIG. 9 illustrates reshaping of an activation tensor in a channel, in accordance with various embodiments.
  • FIG. 9 shows a matrix 910, which is a 2D matrix in a single channel of the activation tensor.
  • the matrix 910 has a spatial shape of 2048 ⁇ 1.2048 may be the spatial height of the activation tensor. 1 may be the spatial width of the activation tensor.
  • the matrix 910 is reshaped to generate a matrix 920.
  • the matrix 920 has a spatial shape of 128 ⁇ 16.
  • the total number of activations in the matrix 920 is the same as the total number of activations in the matrix 910.
  • the arrangement of the activations is changed by the reshaping.
  • the reshaping may not require data movement in the memory (e.g., the memory 310 or the local memory 340) where the activations are stored.
  • the layout of the activations in the memory may remain the same, but the MAC array may read the activations from the memory in a different pattern that matches the shape of the matrix 920.
  • the MAC array reads activations from the memory at a context level.
  • a context may be an operand to be processed by a MAC unit in a computational cycle.
  • a context may have one activation.
  • a context may have 16 activations.
  • Each context may be associated with a storage pointer that indicates a memory address of the context (e.g., the memory address where the first activation in the context is stored) . The reshaping may be performed without making any data movement in the memory. Rather, the storage pointers for reading the contexts may be changed so that the MAC array may read 16 activations per channel to compute a stencil as opposed to reading one activation per channel to compute a stencil.
  • FIG. 10 illustrates computing a channel of an output tensor of a convolution without reshaping, in accordance with various embodiments.
  • the convolution may be an example of the convolution in FIG. 7.
  • the output tensor in the embodiment of FIG. 10 may be an example of the output tensor 730 in FIG. 7.
  • the convolution in FIG. 10 has an activation tensor 1010 with a shape of 2048 ⁇ 1 ⁇ 48 and kernels 1020 having a shape of 1 ⁇ 1.
  • Each channel of the output tensor is a tenor 1030 having a shape of 2048 ⁇ 1.
  • FIG. 11 illustrates computing a channel of an output tensor of a reshaped convolution, in accordance with various embodiments.
  • the reshaped convolution may be an example of the reshaped convolution in FIG. 8.
  • the output tensor in the embodiment of FIG. 10 may be an example of the output tensor 830 in FIG. 8.
  • the reshaped convolution in FIG. 11 has an activation tensor 1110 with a shape of 128 ⁇ 16 ⁇ 48 and kernels 1120 having a shape of 1 ⁇ 1.
  • the kernels 1120 may be the same as the kernel 1020 in FIG. 10.
  • Each channel of the output tensor is a tenor 1130 having a shape of 128 ⁇ 16.
  • the reshaped convolution in FIG. 11 is computationally identical to the convolution in FIG. 10, meaning each input channel convolves with its corresponding filter, with the result being the sum of all these convolutions. Despite the reshaping, the computation remains the same so that the computational accuracy is not impacted while the computational efficiency is improved.
  • FIG. 12 illustrates an example process 1200 of improving computational efficiency of deep learning operations, in accordance with various embodiments.
  • the process 1200 is a process of improving computational efficiency of a convolution in a DNN.
  • a reshaping module 1210 receives an activation tensor 1201 of the convolution and generates a reshaped activation tensor 1202 from the activation tensor 1201.
  • a convolution operator 1220 performs a reshaped convolution on the reshaped activation tensor 1202 and a weight tensor 1203 and computes an output tensor 1204.
  • the reshaping module 1210 receives the output tensor 1204 and generates a reshaped output tensor 1205 from the output tensor 1204.
  • the reshaped output tensor 1205 may be the output of the original convolution, which may be processed in the next deep learning operation of the DNN.
  • An example of the reshaping module 1210 may be the reshaping module 650 in FIG. 6.
  • An example of the convolution operator 1220 may include the data processing cell 900 in FIG. 9 or the data processing unit 1000 in FIG. 10.
  • the DNN system 300 computes 1330, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution.
  • a total number of channels in the activation tensor is equal to the total number of channels in the reshaped activation tensor.
  • a stride of the reshaped convolution is different from a stride of the convolution.
  • the DNN system 300 determines the stride of the reshaped convolution based on the shape of the activation tensor and a shape of the weight tensor.
  • the DNN system 300 stores the activations in a memory based on the shape of the activation tensor.
  • the DNN system 300 reads, by the MAC array, the activations from the memory based on a shape of the reshaped activation tensor.
  • the DNN system 300 performs, by the MAC array, the reshaped convolution on the reshaped activation tensor and a weight tensor of the convolution.
  • the DNN system 300 generates 1340 an output tensor of the convolution by modifying a shape of the output tensor of the reshaped convolution. In some embodiments, the DNN system 300 modifies the shape of the output tensor of the reshaped convolution by modifying the height or width of a matrix in the output tensor of the reshaped convolution.
  • FIG. 14 is a block diagram of an example computing device 1400, in accordance with various embodiments.
  • the computing device 1400 can be used as at least part of the DNN system 300.
  • a number of components are illustrated in FIG. 14 as included in the computing device 1400, but any one or more of these components may be omitted or duplicated, as suitable for the application.
  • some or all of the components included in the computing device 1400 may be attached to one or more motherboards. In some embodiments, some or all of these components are fabricated onto a single system on a chip (SoC) die. Additionally, in various embodiments, the computing device 1400 may not include one or more of the components illustrated in FIG.
  • SoC system on a chip
  • the computing device 1400 may include interface circuitry for coupling to the one or more components.
  • the computing device 1400 may not include a display device 1406, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which a display device 1406 may be coupled.
  • the computing device 1400 may not include an audio input device 1418 or an audio output device 1408 but may include audio input or output device interface circuitry (e.g., connectors and supporting circuitry) to which an audio input device 1418 or audio output device 1408 may be coupled.
  • the computing device 1400 may include a processing device 1402 (e.g., one or more processing devices) .
  • the processing device 1402 processes electronic data from registers and/or memory to transform that electronic data into other electronic data that may be stored in registers and/or memory.
  • the computing device 1400 may include a memory 1404, which may itself include one or more memory devices such as volatile memory (e.g., DRAM) , nonvolatile memory (e.g., read-only memory (ROM) ) , high bandwidth memory (HBM) , flash memory, solid state memory, and/or a hard drive.
  • the memory 1404 may include memory that shares a die with the processing device 1402.
  • the memory 1404 includes one or more non-transitory computer-readable media storing instructions executable to perform operations for executing convolutions (e.g., the method 1300 described in conjunction with FIG. 13) or some operations performed by the DNN system 300.
  • the instructions stored in the one or more non-transitory computer-readable media may be executed by the processing device 1402.
  • the computing device 1400 may include a communication chip 1412 (e.g., one or more communication chips) .
  • the communication chip 1412 may be configured for managing wireless communications for the transfer of data to and from the computing device 1400.
  • wireless and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data through the use of modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not.
  • the communication chip 1412 may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.10 family) , IEEE 802.16 standards (e.g., IEEE 802.16-2005 Amendment) , Long-Term Evolution (LTE) project along with any amendments, updates, and/or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as "3GPP2" ) , etc. ) .
  • IEEE Institute for Electrical and Electronic Engineers
  • Wi-Fi IEEE 802.10 family
  • IEEE 802.16 standards e.g., IEEE 802.16-2005 Amendment
  • LTE Long-Term Evolution
  • LTE Long-Term Evolution
  • UMB ultramobile broadband
  • WiMAX Broadband Wireless Access
  • the communication chip 1412 may operate in accordance with a Global System for Mobile Communication (GSM) , General Packet Radio Service (GPRS) , Universal Mobile Telecommunications System (UMTS) , High Speed Packet Access (HSPA) , Evolved HSPA (E-HSPA) , or LTE network.
  • GSM Global System for Mobile Communication
  • GPRS General Packet Radio Service
  • UMTS Universal Mobile Telecommunications System
  • HSPA High Speed Packet Access
  • E-HSPA Evolved HSPA
  • the communication chip 1412 may operate in accordance with Enhanced Data for GSM Evolution (EDGE) , GSM EDGE Radio Access Network (GERAN) , Universal Terrestrial Radio Access Network (UTRAN) , or Evolved UTRAN (E-UTRAN) .
  • the communication chip 1412 may operate in accordance with Code-division Multiple Access (CDMA) , Time Division Multiple Access (TDMA) , Digital Enhanced Cordless Telecommunications (DECT) , Evolution-Data Optimized (EV-DO) , and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond.
  • the communication chip 1412 may operate in accordance with other wireless protocols in other embodiments.
  • the computing device 1400 may include an antenna 1422 to facilitate wireless communications and/or to receive other wireless communications (such as AM or FM radio transmissions) .
  • the communication chip 1412 may manage wired communications, such as electrical, optical, or any other suitable communication protocols (e.g., the Ethernet) .
  • the communication chip 1412 may include multiple communication chips. For instance, a first communication chip 1412 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second communication chip 1412 may be dedicated to longer-range wireless communications such as global positioning system (GPS) , EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others.
  • GPS global positioning system
  • a first communication chip 1412 may be dedicated to wireless communications
  • a second communication chip 1412 may be dedicated to wired communications.
  • the computing device 1400 may include battery/power circuitry 1414.
  • the battery/power circuitry 1414 may include one or more energy storage devices (e.g., batteries or capacitors) and/or circuitry for coupling components of the computing device 1400 to an energy source separate from the computing device 1400 (e.g., AC line power) .
  • the computing device 1400 may include a display device 1406 (or corresponding interface circuitry, as discussed above) .
  • the display device 1406 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD) , a light-emitting diode display, or a flat panel display, for example.
  • LCD liquid crystal display
  • the computing device 1400 may include an audio output device 1408 (or corresponding interface circuitry, as discussed above) .
  • the audio output device 1408 may include any device that generates an audible indicator, such as speakers, headsets, or earbuds, for example.
  • the computing device 1400 may include an audio input device 1418 (or corresponding interface circuitry, as discussed above) .
  • the audio input device 1418 may include any device that generates a signal representative of a sound, such as microphones, microphone arrays, or digital instruments (e.g., instruments having a musical instrument digital interface (MIDI) output) .
  • MIDI musical instrument digital interface
  • the computing device 1400 may include a GPS device 1416 (or corresponding interface circuitry, as discussed above) .
  • the GPS device 1416 may be in communication with a satellite-based system and may receive a location of the computing device 1400, as known in the art.
  • the computing device 1400 may include another output device 1410 (or corresponding interface circuitry, as discussed above) .
  • Examples of the other output device 1410 may include an audio codec, a video codec, a printer, a wired or wireless transmitter for providing information to other devices, or an additional storage device.
  • the computing device 1400 may include another input device 1420 (or corresponding interface circuitry, as discussed above) .
  • Examples of the other input device 1420 may include an accelerometer, a gyroscope, a compass, an image capture device, a keyboard, a cursor control device such as a mouse, a stylus, a touchpad, a bar code reader, a Quick Response (QR) code reader, any sensor, or a radio frequency identification (RFID) reader.
  • the computing device 1400 may have any desired form factor, such as a handheld or mobile computer system (e.g., a cell phone, a smart phone, a mobile internet device, a music player, a tablet computer, a laptop computer, a netbook computer, an ultrabook computer, a personal digital assistant (PDA) , an ultramobile personal computer, etc. ) , a desktop computer system, a server or other networked computing component, a printer, a scanner, a monitor, a set-top box, an entertainment control unit, a vehicle control unit, a digital camera, a digital video recorder, or a wearable computer system.
  • the computing device 1400 may be any other electronic device that processes data.
  • Example 1 provides a method, including receiving an activation tensor of a convolution, the activation tensor having one or more dimensions and a shape defined by the one or more dimensions; generating a reshaped activation tensor by modifying the shape of the activation tensor based on a structure of a MAC array; computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution; and generating an output tensor of the convolution by modifying a shape of the output tensor of the reshaped convolution.
  • Example 2 provides the method of example 1, in which modifying the shape of the activation tensor includes changing a total number of activations in a dimension of the activation tensor so that a total number of elements in a dimension of the output tensor of the reshaped convolution is a multiple of a total number of MAC units arranged in a row or column of the MAC array.
  • Example 3 provides the method of example 1 or 2, in which modifying the shape of the activation tensor includes adding one or more zero-valued activations into the activation tensor.
  • Example 4 provides the method of example 3, in which modifying the shape of the activation tensor further includes determining a spatial size of the activation tensor in a channel of the convolution; and determining whether the spatial size is a multiple of a total number of MAC units arranged in a row or column of the MAC array, in which the one or more zero-valued activations are added into the activation tensor in response to a determination that the spatial size is a multiple of a total number of MAC units arranged in a row or column of the MAC array.
  • Example 5 provides the method of any one of examples 1-4, in which a stride of the reshaped convolution is different from a stride of the convolution.
  • Example 6 provides the method of example 5, further including determining the stride of the reshaped convolution based on a shape of the activation tensor.
  • Example 11 provides one or more non-transitory computer-readable media storing instructions executable to perform operations, the operations including receiving an activation tensor of a convolution, the activation tensor having one or more dimensions and a shape defined by the one or more dimensions; generating a reshaped activation tensor by modifying the shape of the activation tensor based on a structure of a MAC array; computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution; and generating an output tensor of the convolution by modifying a shape of the output tensor of the reshaped convolution.
  • Example 13 provides the one or more non-transitory computer-readable media of example 11 or 12, in which modifying the shape of the activation tensor includes adding one or more zero-valued activations into the activation tensor.
  • Example 14 provides the one or more non-transitory computer-readable media of any one of examples 11-13, in which a stride of the reshaped convolution is different from a stride of the convolution.
  • Example 15 provides the one or more non-transitory computer-readable media of any one of examples 11-14, in which computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution includes storing activations of the activation tensor in a memory based on the shape of the activation tensor; and reading, by the MAC array, the activations from the memory based on a shape of the reshaped activation tensor.
  • Example 16 provides the one or more non-transitory computer-readable media of any one of examples 11-15, in which computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution includes performing, by the MAC array, the reshaped convolution on the reshaped activation tensor and a weight tensor of the convolution.
  • Example 17 provides the one or more non-transitory computer-readable media of any one of examples 11-16, in which the activation tensor includes a plurality of input channels, an input channel corresponds to a matrix including activations arranged in one or more rows and one or more columns, and modifying the shape of the activation tensor includes modifying a height or width of the matrix.
  • Example 18 provides an apparatus, including a computer processor for executing computer program instructions; and a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations including receiving an activation tensor of a convolution, the activation tensor having one or more dimensions and a shape defined by the one or more dimensions, generating a reshaped activation tensor by modifying the shape of the activation tensor based on a structure of a MAC array, computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution, and generating an output tensor of the convolution by modifying a shape of the output tensor of the reshaped convolution.
  • Example 19 provides the apparatus of example 18, in which modifying the shape of the activation tensor includes changing a total number of activations in a dimension of the activation tensor so that a total number of elements in a dimension of the output tensor of the reshaped convolution is a multiple of a total number of MAC units arranged in a row or column of the MAC array.
  • Example 20 provides the apparatus of example 18 or 19, in which modifying the shape of the activation tensor includes adding one or more zero-valued activations into the activation tensor.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Biomedical Technology (AREA)
  • Biophysics (AREA)
  • General Health & Medical Sciences (AREA)
  • Data Mining & Analysis (AREA)
  • Evolutionary Computation (AREA)
  • Computational Linguistics (AREA)
  • Molecular Biology (AREA)
  • Computing Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Mathematical Physics (AREA)
  • Software Systems (AREA)
  • Artificial Intelligence (AREA)
  • Neurology (AREA)
  • Complex Calculations (AREA)

Abstract

A convolution in a deep neural network (DNN) may have an activation tensor, where activations are arranged in one or more dimensions, and a weight tensor. The convolution is to be performed by at least one multiply-accumulate (MAC) array including MAC units arranged in one or more rows and one or more columns. The MAC array may have a height measured by the number of MAC units in a single column and a width measured by the number of MAC units in a single row. The convolution, before being performed, may be reshaped by changing one or more dimensions of the activation tensor so that a dimension of the output tensor would become a multiple of the height or width of the MAC array. The MAC array may compute an output tensor of the reshaped convolution, which may then be reshaped to generate the output tensor of the original convolution.

Description

RESHAPING CONVOLUTION BASED ON CONFIGURATION OF DEEP NEURAL NETWORK ACCELERATOR Technical Field
This disclosure relates generally to neural networks (also referred to as “deep neural networks” or “DNN” ) , and more specifically, reshaping convolutions based on configurations of DNN accelerators.
Background
DNNs are used extensively for a variety of artificial intelligence applications ranging from computer vision to speech recognition and natural language processing due to their ability to achieve high accuracy. However, the high accuracy comes at the expense of significant computation cost. DNNs have extremely high computing demands as there can be hundreds of millions of MAC (multiply-accumulate) operations as well as a large amount of data to read and write. Therefore, techniques to improve efficiency of DNNs are needed.
Brief Description of the Drawings
Embodiments will be readily understood by the following detailed description in conjunction with the accompanying drawings. To facilitate this description, like reference numerals designate like structural elements. Embodiments are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings.
FIG. 1 illustrates an example DNN, in accordance with various embodiments.
FIG. 2 illustrates an example convolution, in accordance with various embodiments.
FIG. 3 is a block diagram of a DNN system, in accordance with various embodiments.
FIG. 4 illustrates an example data processing cell, in accordance with various embodiments.
FIG. 5 illustrates an example data processing unit, in accordance with various embodiments.
FIG. 6 is a block diagram of a DNN module, in accordance with various embodiments.
FIG. 7 illustrates a convolution before reshaping, in accordance with various embodiments.
FIG. 8 illustrates a reshaped convolution, in accordance with various embodiments.
FIG. 9 illustrates reshaping of an activation tensor in a channel, in accordance with various embodiments.
FIG. 10 illustrates computing a channel of an output tensor of a convolution without reshaping, in accordance with various embodiments.
FIG. 11 illustrates computing a channel of an output tensor of a reshaped convolution, in accordance with various embodiments.
FIG. 12 illustrates an example process of improving convolution compute efficiency, in accordance with various embodiments.
FIG. 13 is a flowchart showing a method of executing a convolution, in accordance with various embodiments.
FIG. 14 is a block diagram of an example computing device, in accordance with various embodiments.
Detailed Description
Overview
The last decade has witnessed a rapid rise in artificial intelligence (AI) based data processing, particularly based on DNNs. DNNs are widely used in the domains of computer vision, speech recognition, image, and video processing mainly due to their ability to achieve beyond human-level accuracy. The significant improvements in DNN model size and accuracy coupled with the rapid increase in computing power of execution platforms have led to the adoption of DNN applications even within resource constrained mobile and edge devices that have limited energy availability.
A DNN layer may include one or more deep learning operations (also referred to as “neural network operations” ) , such as matrix multiplication, convolution, pooling, elementwise operation, linear operation, nonlinear operation, and so on. Input or output data of deep learning operations may be arranged in data structures called tensors. A tensor is a data structure having multiple elements across one or more dimensions. Examples of tensors include vector (which is one-dimensional (1D) tensor) , matrix (which is two-dimensional (2D) tensor) , three-dimensional (3D) tensors, four-dimensional (4D) tensors, and even higher dimensional tensors. A dimension of a tensor may correspond to an axis,  e.g., an axis in a coordinate system. A dimension may be measured by the number of data points along the axis. The dimensions of a tensor may define the shape of the tensor. A DNN layer may receive one or more input tensors and compute an output tensor from the one or more input tensors. For a convolution layer, the input tensors include an activation tensor (also referred to as “input feature map (IFM) ” ) including one or more activations (also referred to as “input elements” ) and a weight tensor. The weight tensor may be a kernel (a 2D weight tensor) , a filter (a 3D weight tensor) , or a group of filters (a 4D weight tensor) .
The workload of performing a deep learning operation (e.g., convolution, matrix multiplication operation, etc. ) may be visualized as a nested loop structure having a 4D spatial shape. This structure spans across the output channel, spatial height, spatial width, and input channel, which are the four dimensions of the spatial shape. The innermost loop may represent a hardware-unrolled stencil computation. The spatial height and spatial width may be the height and width, respectively, of the output tensor. For instance, the spatial height and spatial width may be the height and width of a 2D matrix corresponding to a single channel in the output tensor.
DNN accelerators (also referred to as “AI accelerator” or “AI processor” ) are typically processors designed to expedite computations in DNNs. As the size of DNN models (e.g., the number of internal parameters) are having a substantial growth, the diversity of computation workloads causes challenges in the design of DNN accelerators, such as efficiently running a variety of models on a single hardware system, accommodating a broad range of model sizes, managing operational intensity (measured in operations per byte) , meeting latency and power requirements, and so on. To efficiently run diverse models, hardware needs to support flexible and various stencil settings, while software can choose the optimal setting based on cost metric or workload’s pattern. Currently available DNN accelerators can support various fixed stencil settings. Stencil refers to the fundamental computational unit within a single computational cycle. Such fixed stencil settings can offer certain level of flexibility, allowing for optimal performance across various workload shapes. The mapping of these settings is usually managed by a deep learning compiler or software stack. However, the stencil settings supported by these DNN accelerators are often fixed and limited. This is usually due to the complexities of hardware implementation, the need for efficient operation, and the extensive validation efforts required. Even though stencil  settings can provide some level of adaptability, their range is constrained by these practical considerations. A DNN accelerator may include one or more MAC arrays that perform deep learning operations, such as convolutions, matrix multiplications, and so on. A MAC array is typically organized into a two-dimensional array of MAC units that perform MAC operations. The fixed stencil settings of the DNN accelerator cause the stencils processed by the MAC units to adhere to certain patterns, such as 4x4x16x8, 8x2x16x8, and so on.
In an example where the stencil setting is 4x4x16x8, a MAC array may process a 4x4x16x8 output subtensor within one computational cycle, meaning the MAC unit can compute an output subtensor that spans across 16 output channels, 4 spatial heights, 4 spatial widths, and 8 input channels in a single cycle. A workload having a size of 16x16x32x8 would take 32 (4x4x2x1) cycles, and the utilization of the MAC units would be 100%. However, not all workloads can get desirable hardware utilization and computational efficiency. Some workload may have spatial shapes that do not align with any of the stencil settings of the DNN accelerator. For instance, a dimension of the spatial shape of a workload may not be a multiple of the corresponding dimension of any stencil setting. Such “stencil-unfriendly” workloads can suffer from low utilization of the MAC units. In an example, a workload with a size of 1x16x32x8 would also take 32 cycles, and the utilization of the MAC units would be reduced to 25%. Many DNNs models may have stencil-unfriendly workloads. For instance, workloads for executing Transformer-based models may feature a kernel size of [1, 1] or unbalanced spatial shapes like [1, n] or [n, 1] . It can be difficult to design one chip with fixed stencil to run diverse workloads efficiently.
Embodiments of the present disclosure may improve on at least some of the challenges and issues described above by reshaping deep learning operations in DNNs based on configuration of MAC arrays in DNN accelerators. The deep learning operations may include convulsions, matrix multiplications, and so on. In an example, an input tensor of a deep learning operation may be reshaped based on the configuration of the DNN accelerator running the deep learning operation to align the spatial shape of the workload for performing the deep learning operation with a stencil setting of the DNN accelerator. That way, despite that the stencil settings of DNN accelerators are fixed and limited, the computational efficiency of the DNN accelerators can reach a desirable level for workloads with various spatial shapes, including stencil-unfriendly workloads.
In various embodiments of the present disclosure, a DNN accelerator may include one or more tiles. A tile may be referred to as a compute block, which may include a data processing unit (aka neural processing unit) . A data processing unit may include one or more MAC arrays, each of which have MAC units arranged in one or more columns and one or more rows. The number of MAC unit (s) in a column of a MAC array may indicate the height of the MAC array. The number of MAC unit (s) in a row of the array may indicate the width of the MAC array. A deep learning operation may be performed by at least one MAC array and may be reshaped based on the height or width of a MAC array to increase utilization of the MAC units in the process of performing the deep learning operation.
Taking convolutions as example, a convolution may have an activation tensor and a weight tensor, which are inputs to the convolution. The output of the convolution may be an output tensor, which may not be aligned with the configuration (e.g., the stencil setting) of the DNN accelerator. For instance, the height or width of the output tensor may not be a multiple of the height or width of the MAC array, which can result in undesirable utilization of MAC units in the MAC array. The convolution, before being performed, may be transformed into a convolution aligned with the configuration of the DNN accelerator by reshaping the activation tensor. The activation tensor may have a 3D shape defined by a height measured by the number of activations in a column, a width measured by the number of weights in a row, and a depth measured by the number of input channels of the convolution. To reshape the activation tensor, one or more dimensions of the activation tensor may be modified. For instance, the height or width of the activation tensor of the convolution may be modified so that the height or width of the output tensor would be a multiple of the height or width of the MAC array. In some embodiments, the height or width of the activation tensor may be changed to become a multiple of the height or width of the MAC array. The MAC array may compute the output tensor of the reshaped convolution. The computational efficiency for performing the reshaped convolution can be better than the computational efficiency for performing the original condition as the utilization of the MAC units would be better. The output tensor of the reshaped convolution may then be reshaped to generate the output tensor of the original convolution.
The present disclosure provides a spatial reshape mechanism that can optimize utilization of MAC units in DNN accelerators with fixed stencil settings. The reshaping may  introduce little or even no data movement overhead or computation overhead. With the reshaping, many stencil-unfriendly workloads, such as workloads in Transformer-based models and 1D convolutions in audio applications, can achieve similar or even the same computational efficiency as convolution-based models. Therefore, desirable hardware utilization and efficiency can be achieved even for stencil-unfriendly workloads.
For purposes of explanation, specific numbers, materials and configurations are set forth in order to provide a thorough understanding of the illustrative implementations. However, it will be apparent to one skilled in the art that the present disclosure may be practiced without the specific details or/and that the present disclosure may be practiced with only some of the described aspects. In other instances, well known features are omitted or simplified in order not to obscure the illustrative implementations.
Further, references are made to the accompanying drawings that form a part hereof, and in which are shown, by way of illustration, embodiments that may be practiced. It is to be understood that other embodiments may be utilized, and structural or logical changes may be made without departing from the scope of the present disclosure. Therefore, the following detailed description is not to be taken in a limiting sense.
Various operations may be described as multiple discrete actions or operations in turn, in a manner that is most helpful in understanding the claimed subject matter. However, the order of description should not be construed as to imply that these operations are necessarily order dependent. In particular, these operations may not be performed in the order of presentation. Operations described may be performed in a different order from the described embodiment. Various additional operations may be performed or described operations may be omitted in additional embodiments.
For the purposes of the present disclosure, the phrase “A or B” or the phrase "A and/or B" means (A) , (B) , or (A and B) . For the purposes of the present disclosure, the phrase “A, B, or C” or the phrase "A, B, and/or C" means (A) , (B) , (C) , (A and B) , (A and C) , (B and C) , or (A, B, and C) . The term "between, " when used with reference to measurement ranges, is inclusive of the ends of the measurement ranges.
The description uses the phrases "in an embodiment" or "in embodiments, " which may each refer to one or more of the same or different embodiments. The terms "comprising, " "including, " "having, " and the like, as used with respect to embodiments of  the present disclosure, are synonymous. The disclosure may use perspective-based descriptions such as "above, " "below, " "top, " "bottom, " and "side" to explain various features of the drawings, but these terms are simply for ease of discussion, and do not imply a desired or required orientation. The accompanying drawings are not necessarily drawn to scale. Unless otherwise specified, the use of the ordinal adjectives “first, ” “second, ” and “third, ” etc., to describe a common object, merely indicates that different instances of like objects are being referred to and are not intended to imply that the objects so described must be in a given sequence, either temporally, spatially, in ranking or in any other manner.
In the following detailed description, various aspects of the illustrative implementations will be described using terms commonly employed by those skilled in the art to convey the substance of their work to others skilled in the art.
The terms “substantially, ” “close, ” “approximately, ” “near, ” and “about, ” generally refer to being within +/-20%of a target value as described herein or as known in the art. Similarly, terms indicating orientation of various elements, e.g., “coplanar, ” “perpendicular, ” “orthogonal, ” “parallel, ” or any other angle between the elements, generally refer to being within +/-5-20%of a target value as described herein or as known in the art.
In addition, the terms “comprise, ” “comprising, ” “include, ” “including, ” “have, ” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a method, process, device, or DNN accelerator that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such method, process, device, or DNN accelerators. Also, the term “or” refers to an inclusive “or” and not to an exclusive “or. ”
The systems, methods and devices of this disclosure each have several innovative aspects, no single one of which is solely responsible for all desirable attributes disclosed herein. Details of one or more implementations of the subject matter described in this specification are set forth in the description below and the accompanying drawings.
Example DNN
FIG. 1 illustrates an example DNN 100, in accordance with various embodiments. The DNN 100 may be executed by a DNN accelerator, e.g., the DNN accelerator 302 in FIG. 3. In an example, the DNN 100 may be a convolution-based DNN. In other examples, the DNN 100 may be other types of DNNs. For the purpose of illustration, the DNN 100 includes a  sequence of layers comprising a plurality of convolutional layers 110 (individually referred to as “convolutional layer 110” ) , a plurality of pooling layers 120 (individually referred to as “pooling layer 120” ) , and a plurality of fully-connected layers 130 (individually referred to as “fully-connected layer 130” ) . In other embodiments, the DNN 100 may include fewer, more, or different layers. In an execution of the DNN 100, the layers of the DNN 100 execute tensor computation that includes many tensor operations, such as matrix multiplications, convolutions (e.g., multiply-accumulate (MAC) operations, etc. ) , pooling operations, elementwise operations (e.g., elementwise addition, elementwise multiplication, etc. ) , other types of tensor operations, or some combination thereof.
The convolutional layers 110 summarize the presence of features in inputs to the DNN 100. The convolutional layers 110 function as feature extractors. The first layer of the DNN 100 is a convolutional layer 110. In an example, a convolutional layer 110 performs a convolution on an input tensor 140 (also referred to as IFM 140) and a filter 150. As shown in FIG. 1, the IFM 140 is represented by a 7×7×3 three-dimensional (3D) matrix. The IFM 140 includes 3 input channels, each of which is represented by a 7×7 two-dimensional (2D) matrix. The 7×7 2D matrix includes 7 input elements (also referred to as input points) in each row and 7 input elements in each column. The filter 150 is represented by a 3×3×3 3D matrix. The filter 150 includes 3 kernels, each of which may correspond to a different input channel of the IFM 140. A kernel is a 2D matrix of weights, where the weights are arranged in columns and rows. A kernel can be smaller than the IFM. In the embodiments of FIG. 1, each kernel is represented by a 3×3 2D matrix. The 3×3 kernel includes 3 weights in each row and 3 weights in each column. Weights can be initialized and updated by backpropagation using gradient descent. The magnitudes of the weights can indicate importance of the filter 150 in extracting features from the IFM 140.
The convolution includes MAC operations with the input elements in the IFM 140 and the weights in the filter 150. The convolution may be a standard convolution 163 or a depthwise convolution 183. In the standard convolution 163, the whole filter 150 slides across the IFM 140. All the input channels are combined to produce an output tensor 160 (also referred to as OFM 160) . The OFM 160 is represented by a 5×5 2D matrix. The 5×5 2D matrix includes 5 output elements (also referred to as output points) in each row and 5 output elements in each column. For the purpose of illustration, the standard convolution  includes one filter in the embodiments of FIG. 1. In embodiments where there are multiple filters, the standard convolution may produce multiple output channels in the OFM 160.
The multiplication applied between a kernel-sized patch of the IFM 140 and a kernel may be a dot product. A dot product is the elementwise multiplication between the kernel-sized patch of the IFM 140 and the corresponding kernel, which is then summed, always resulting in a single value. Because it results in a single value, the operation is often referred to as the “scalar product. ” Using a kernel smaller than the IFM 140 is intentional as it allows the same kernel (set of weights) to be multiplied by the IFM 140 multiple times at different points on the IFM 140. Specifically, the kernel is applied systematically to each overlapping part or kernel-sized patch of the IFM 140, left to right, top to bottom. The result from multiplying the kernel with the IFM 140 one time is a single value. As the kernel is applied multiple times to the IFM 140, the multiplication result is a 2D matrix of output elements. As such, the 2D output matrix (i.e., the OFM 160) from the standard convolution 163 is referred to as an OFM.
In the depthwise convolution 183, the input channels are not combined. Rather, MAC operations are performed on an individual input channel and an individual kernel and produce an output channel. As shown in FIG. 1, the depthwise convolution 183 produces a depthwise output tensor 180. The depthwise output tensor 180 is represented by a 5×5×3 3D matrix. The depthwise output tensor 180 includes 3 output channels, each of which is represented by a 5×5 2D matrix. The 5×5 2D matrix includes 5 output elements in each row and 5 output elements in each column. Each output channel is a result of MAC operations of an input channel of the IFM 140 and a kernel of the filter 150. For instance, the first output channel (patterned with dots) is a result of MAC operations of the first input channel (patterned with dots) and the first kernel (patterned with dots) , the second output channel (patterned with horizontal strips) is a result of MAC operations of the second input channel (patterned with horizontal strips) and the second kernel (patterned with horizontal strips) , and the third output channel (patterned with diagonal stripes) is a result of MAC operations of the third input channel (patterned with diagonal stripes) and the third kernel (patterned with diagonal stripes) . In such a depthwise convolution, the number of input channels equals the number of output channels, and each output channel corresponds to a different input channel. The input channels and output channels are referred to collectively as depthwise  channels. After the depthwise convolution, a pointwise convolution 193 is then performed on the depthwise output tensor 180 and a 1×1×3 tensor 190 to produce the OFM 160.
The OFM 160 is then passed to the next layer in the sequence. In some embodiments, the OFM 160 is passed through an activation function. An example activation function is rectified linear unit (ReLU) . ReLU is a calculation that returns the value provided as input directly, or the value zero if the input is zero or less. The convolutional layer 110 may receive several images as input and calculate the convolution of each of them with each of the kernels. This process can be repeated several times. For instance, the OFM 160 is passed to the subsequent convolutional layer 110 (i.e., the convolutional layer 110 following the convolutional layer 110 generating the OFM 160 in the sequence) . The subsequent convolutional layers 110 perform a convolution on the OFM 160 with new kernels and generate a new feature map. The new feature map may also be normalized and resized. The new feature map can be kernelled again by a further subsequent convolutional layer 110, and so on.
In some embodiments, a convolutional layer 110 has four hyperparameters: the number of kernels, the size F kernels (e.g., a kernel is of dimensions F×F×D pixels) , the S step with which the window corresponding to the kernel is dragged on the image (e.g., a step of one means moving the window one pixel at a time) , and the zero-padding P (e.g., adding a black contour of P pixels thickness to the input image of the convolutional layer 110) . The convolutional layers 110 may perform various types of convolutions, such as 2-dimensional convolution, dilated or atrous convolution, spatial separable convolution, depthwise separable convolution, transposed convolution, and so on. The DNN 100 includes 16 convolutional layers 110. In other embodiments, the DNN 100 may include a different number of convolutional layers.
The pooling layers 120 down-sample feature maps generated by the convolutional layers, e.g., by summarizing the presence of features in the patches of the feature maps. A pooling layer 120 is placed between two convolution layers 110: a preceding convolutional layer 110 (the convolution layer 110 preceding the pooling layer 120 in the sequence of layers) and a subsequent convolutional layer 110 (the convolution layer 110 subsequent to the pooling layer 120 in the sequence of layers) . In some embodiments, a pooling layer 120  is added after a convolutional layer 110, e.g., after an activation function (e.g., ReLU, etc. ) has been applied to the OFM 160.
A pooling layer 120 receives feature maps generated by the preceding convolution layer 110 and applies a pooling operation to the feature maps. The pooling operation reduces the size of the feature maps while preserving their important characteristics. Accordingly, the pooling operation improves the efficiency of the DNN and avoids over-learning. The pooling layers 120 may perform the pooling operation through average pooling (calculating the average value for each patch on the feature map) , max pooling (calculating the maximum value for each patch of the feature map) , or a combination of both. The size of the pooling operation is smaller than the size of the feature maps. In various embodiments, the pooling operation is 2×2 pixels applied with a stride of two pixels, so that the pooling operation reduces the size of a feature map by a factor of 2, e.g., the number of pixels or values in the feature map is reduced to one quarter the size. In an example, a pooling layer 120 applied to a feature map of 6×6 results in an output pooled feature map of 3×3. The output of the pooling layer 120 is inputted into the subsequent convolution layer 110 for further feature extraction. In some embodiments, the pooling layer 120 operates upon each feature map separately to create a new set of the same number of pooled feature maps.
The fully-connected layers 130 are the last layers of the DNN. The fully-connected layers 130 may be convolutional or not. The fully-connected layers 130 receive an input operand. The input operand defines the output of the convolutional layers 110 and pooling layers 120 and includes the values of the last feature map generated by the last pooling layer 120 in the sequence. The fully-connected layers 130 apply a linear combination and an activation function to the input operand and generate a vector. The vector may contain as many elements as there are classes: element i represents the probability that the image belongs to class i. Each element is therefore between 0 and 1, and the sum of all is worth one. These probabilities are calculated by the last fully-connected layer 130 by using a logistic function (binary classification) or a SoftMax function (multi-class classification) as an activation function. In some embodiments, the fully-connected layers 130 multiply each input element by weight, make the sum, and then apply an activation function (e.g., logistic if N=2, SoftMax if N>2) . This is equivalent to multiplying the input operand by the matrix containing the weights.
Example Convolution
FIG. 2 illustrates an example convolution, in accordance with various embodiments. The convolution may be a deep learning operation in a convolutional layer of a DNN, e.g., a convolutional layer 110 in FIG. 1. The convolution can be executed on an activation tensor 210 and filters 220 (individually referred to as “filter 220” ) . The filters may constitute a weight tensor of the convolution. The result of the convolution is an output tensor 230. In some embodiments, the convolution is performed by a DNN accelerator. An example of the DNN accelerator may be the DNN accelerator 302 in FIG. 3. For instance, the convolution may be performed by the sparse cell array 370 in the DNN accelerator 302.
In the embodiments of FIG. 2, the activation tensor 210 includes activations (also referred to as “input activations, ” “elements, ” or “input elements” ) arranged in a 3D matrix. The activation tensor 210 may also be referred to as an input tensor of the convolution. An input element is a data point in the activation tensor 210. The activation tensor 210 has a spatial size Hin×Win×Cin, where Hin is the height of the 3D matrix (i.e., the length along the Y axis, which indicates the number of activations in a column in the 3D matrix of each input channel) , Win is the width of the 3D matrix (i.e., the length along the X axis, which indicates the number of activations in a row in the 2D matrix of each input channel) , and Cin is the depth of the 3D matrix (i.e., the length along the Z axis, which indicates the number of input channels) . For the purpose of simplicity and illustration, the activation tensor 210 has a spatial size of 7×7×3, i.e., the activation tensor 210 includes three input channels and each input channel has a 7×7 2D matrix. Each input element in the activation tensor 210 may be represented by a (X, Y, Z) coordinate. In other embodiments, the height, width, or depth of the activation tensor 210 may be different.
Each filter 220 includes weights arranged in a 3D matrix. The values of the weights may be determined through training the DNN. A filter 220 has a spatial size Hf×Wf×Cf, where Hf is the height of the filter (i.e., the length along the Y axis, which indicates the number of weights in a column in each kernel) , Wf is the width of the filter (i.e., the length along the X axis, which indicates the number of weights in a row in each kernel) , and Cf is the depth of the filter (i.e., the length along the Z axis, which indicates the number of channels) . In some embodiments, Cf equals Cin. For purpose of simplicity and illustration, each filter 220 in FIG. 2 has a spatial size of 2×3×3, i.e., the filter 220 includes 2 convolutional  kernels with a spatial size of 2×3. In other embodiments, the height, width, or depth of the filter 220 may be different. The spatial size of the convolutional kernels is smaller than the spatial size of the 2D matrix of each input channel in the activation tensor 210.
An activation or weight may take one or more bytes in a memory. The number of bytes for an activation or weight may depend on the data format. For example, when the activation or weight has an INT8 format, the activation takes one byte. When the activation or weight has a FP16 format, the activation or weight takes two bytes. Other data formats may be used for activations or weights.
In the convolution, each filter 220 slides across the activation tensor 210 and generates a 2D matrix for an output channel in the output tensor 230. In the embodiments of FIG. 2, the 2D matrix has a spatial size of 5×5. The output tensor 230 includes activations (also referred to as “output activations, ” “elements, ” or “output element” ) arranged in a 3D matrix. An output activation is a data point in the output tensor 230. The output tensor 230 has a spatial size Hout×Wout×Cout, where Hout is the height of the 3D matrix (i.e., the length along the Y axis, which indicates the number of output activations in a column in the 2D matrix of each output channel) , Wout is the width of the 3D matrix (i.e., the length along the X axis, which indicates the number of output activations in a row in the 2D matrix of each output channel) , and Cout is the depth of the 3D matrix (i.e., the length along the Z axis, which indicates the number of output channels) . Cout may equal the number of filters 220 in the convolution. Hout and Wout may depend on the heights and weights of the activation tensor 210 and each filter 220. In an example where the kernel size is 1×1, Hout and Wout may equal to Hin and Win, respectively.
As a part of the convolution, MAC operations can be performed on a 2×3×3 subtensor 215 (which is highlighted with a dotted pattern in FIG. 2) in the activation tensor 210 and each filter 220. The result of the MAC operations on the subtensor 215 and one filter 220 is an output activation. In some embodiments (e.g., embodiments where the convolution is an integral convolution) , an output activation may include 8 bits, e.g., one byte. In other embodiments (e.g., embodiments where the convolution is a floating-point convolution) , an output activation may include more than one byte. For instance, an output element may include two bytes.
After the MAC operations on the subtensor 215 and all the filters 220 are finished, a vector 235 is produced. The vector 235 is highlighted with slashes in FIG. 2. The vector 235 includes a sequence of output activations, which are arranged along the Z axis. The output activations in the vector 235 have the same (X, Y) coordinate, but the output activations correspond to different output channels and have different Z coordinates. The dimension of the vector 235 along the Z axis may equal the total number of output channels in the output tensor 230. After the vector 235 is produced, further MAC operations are performed to produce additional vectors till the output tensor 230 is produced.
In some embodiments, the MAC operations on a 2×3×3 subtensor (e.g., the subtensor 215) and a filter 220 may be performed by a plurality of MAC units. One or more MAC units may receive an input operand (e.g., an activation operand 217 shown in FIG. 2) and a weight operand (e.g., the weight operand 227 shown in FIG. 2) . The activation operand 217 includes a sequence of activations having the same (x, y) coordinate but different z coordinates. The activation operand 217 includes an activation from each of the input channels in the activation tensor 210. The weight operand 227 includes a sequence of weights having the same (x, y) coordinate but different z coordinates. The weight operand 227 includes a weight from each of the channels in the filter 220. Activations in the activation operand 217 and weights in the weight operand 227 may be sequentially fed into a MAC unit. The MAC unit may receive an activation and a weight ( “an activation-weight pair” ) at a time and multiple the activation and the weight. The position of the activation in the activation operand 217 may match the position of the weight in the weight operand 227. The activation and weight may correspond to the same channel.
Activations or weights may be floating-point numbers. Floating-point numbers may have various data formats, such as FP32, FP16, BF16, and so on. A floating-point number may be a positive or negative number with a decimal point. A floating-point number may be represented by a sequence of bits that includes one or more bits representing the sign of the floating-point number (e.g., positive or negative) , bits representing an exponent of the floating-point number, and bits representing a mantissa of the floating-point number. The mantissa is the part of a floating-point number that represents the significant digits of that number. The mantissa is multiplied by the base raised to the exponent to give the actual value of the floating-point number.
In some embodiments, the output activations in the output tensor 230 may be further processed based on one or more activation functions before they are stored or inputted into the next layer of the DNN. The processing based on the one or more activation functions may be at least part of the post processing of the convolution. In some embodiments, the post processing may include one or more other computations, such as offset computation, bias computation, and so on. The results of the post processing may be stored in a local memory of the compute block and be used as input to the next DNN layer. In some embodiments, the input activations in the activation tensor 210 may be results of post processing of the previous DNN layer.
Example DNN System
FIG. 3 is a block diagram of a DNN system 300, in accordance with various embodiments. The whole DNN system 300 or a part of the DNN system 300 may be implemented in one or more computing devices, such as the computing device 1400 in FIG. 14. The DNN system 300 can generate and execute DNNs, such as Transformer-based models, convolution-based models, and so on. As shown in FIG. 3, the DNN system 300 includes a DNN module 301 and a DNN accelerator 302. In other embodiments, alternative configurations, different or additional components may be included in the DNN system 300. For instance, the DNN system 300 may include multiple DNN modules or multiple DNN accelerators. Further, functionality attributed to a component of the DNN system 300 may be accomplished by a different component included in the DNN system 300 or a different system. In some embodiments, the DNN module 301 and DNN accelerator 302 may include different types of processing units. In an example, the DNN module 301 may be implemented by one or more central processing units (CPUs) . The DNN accelerator 302 may also be referred to as an AI accelerator or an AI processor. The DNN module 301 and DNN accelerator 302 may be implemented in the same chip or separate chips.
The DNN module 301 facilitates generation and deployment of DNNs. In some embodiments, the DNN module 301 may generate and train DNNs. For instance, the DNN module 301 can define the layered architecture of a DNN. The DNN module 301 can also determine the internal parameters of the DNN through a DNN training process. The DNN module 301 may also determine one or more hyperparameters that define how the DNN is  trained. An example hyperparameter is a sparsity ratio that defines the sparsity level of one or more deep learning tensors for the DNN.
The DNN module 301 may also compress DNNs, e.g., during or after training. In some embodiments, the DNN module 301 may prune weights in one or more layers of a DNN by changing nonzero valued weight to zeros. The DNN module 301 may prune weights based on a target weight sparsity ratio. A weight sparsity ratio may be the ratio of the number of zero-valued weights to the total number of weights. In an example where the DNN module 301 prunes weight during DNN training, the DNN module 301 may prune weight of a layer to achieve a target sparsity ratio after one or more epochs. The DNN module 301 may prevent the pruned weights from changing values during the rest of the training process. Alternatively, the DNN module 301 may allow the pruned weights to change values so that a pruned, zero-valued weight may have a nonzero value after further training. The DNN module 301 may prune weights of the layer again after one or more additional epochs.
The DNN module 301 may deploy trained, compressed, or validated DNNs for use in deep learning applications. In some embodiments, the DNN module 301 may distribute trained, compressed, or validated DNNs to devices or systems which may use the DNNs to perform tasks (e.g., image classification, motion planning, etc. ) for which the DNNs were trained. In other embodiments, the DNN module 301 may facilitate deployment of the DNNs using the DNN accelerator 302. For instance, the DNN module 301 may receive data from a device or system coupled with the DNN system 300 and input the received data (or data generated by the DNN module 301, e.g., based on the received data) into a DNN. The DNN module 301 may generate instructions (e.g., configuration files) that control the operation of the DNN accelerator 302 during the DNN execution. The DNN module 301 may receive an output of the DNN from the DNN accelerator 302. The DNN module 301 may transmit the output of the DNN (or a result of processing the output of the DNN by the DNN module 301) to the device or system.
The DNN module 301 may control execution processes of trained, compressed, or validated DNNs. The DNN module 301 may function as a deep learning complier for DNNs executed by the DNN accelerator 302. In some embodiments, the DNN module 301 facilitates reshaping deep learning operations to be executed by the DNN accelerator 302 so that the spatial shape of the deep learning operations may be aligned with the  configurations (e.g., stencil settings) of the DNN accelerator 302. For instance, the DNN module 301 may change the shape of activation tensors of convolutions so that the spatial heights or widths of output tensors may be multiples of the height or width of one or more MAC arrays in the DNN accelerator 302. The DNN module 301 may cause the DNN accelerator 302 to execute reshaped convolutions. The DNN module 301 may receive output tensors of reshaped convolution from the DNN accelerator and reshape the output tensors to generate output tensors of the original convolutions. Certain aspects of the DNN module 301 are provided below in conjunction with FIG. 6.
The DNN accelerator 302 executes DNNs provided by the DNN module 301. For instance, the DNN accelerator 302 can execute a DNN by running deep learning operations in the DNN. The process of carrying out a deep learning operation is also referred to as a process of executing the deep learning operation or performing the deep learning operation. The execution of the DNN may be for training the DNN or for using the DNN to perform AI tasks. In some embodiments, the DNN accelerator 302 includes components designed for optimal efficiency in running convolution-based DNNs. As shown in FIG. 3, the DNN accelerator 302 includes a memory 310, a DMA (direct memory access) engine 320, and compute blocks 330 (individually referred to as “compute block 330” ) . In other embodiments, alternative configurations, different or additional components may be included in the DNN accelerator 302. For example, the DNN accelerator 302 may include more than one memory 310 or DMA engine 320. As another example, the DNN accelerator 302 may include a single compute block 330. Further, functionality attributed to a component of the DNN accelerator 302 may be accomplished by a different component included in the DNN accelerator 302 or by a different system. A component of the DNN accelerator 302 may be implemented in hardware, software, firmware, or some combination thereof.
The memory 310 stores data associated with deep learning operations performed by the DNN accelerator 302. In some embodiments, the memory 310 may store data to be used by the compute blocks 330 for DNN execution. The memory 310 may store weights, such as weights of convolutional layers, which are determined by training DNNs. The memory 310 may further store inputs to DNN layers or outputs of DNN layers, such as data generated by the compute blocks 330 from performing deep learning operations in DNNs. Example deep  learning operations include convolutions (also referred to as “convolutional operations” ) , matrix multiplication operations, pooling operations, elementwise operations, activation functions, other types of deep learning operations, or some combination thereof. The memory 310 may be a main memory of the DNN accelerator 302. In some embodiments, the memory 310 includes one or more dynamic random-access memories (DRAMs) .
The DMA engine 320 facilitates data transfer between the memory 310 and local memories of the compute blocks 330. For example, the DMA engine 320 can read data from the memory 310 and write data into a local memory of a compute block 330. As another example, the DMA engine 320 can read data from a local memory of a compute block 330and write data into the memory 310. The DMA engine 320 provides a DMA feature that allows the compute block 330 to initiate data transfer between the memory 310 and the local memories of the compute blocks 330 and to perform other operations while the data transfer is being conducted. In some embodiments, the DMA engine 320 may read tensors from the memory 310, modify the tensors in a way that is optimized for the compute block 330 before it writes the tensors into the local memories of the compute blocks 330.
The compute blocks 330 can perform deep learning operations in DNNs. For instance, a compute block 330 may execute a DNN layer by running one or more deep learning operations in the DNN layer. A compute block 330 may execute a layer, or a portion of a layer, at a time. The compute blocks 330 may be capable of running various types of deep learning operations, such as convolution, pooling, elementwise operation, linear operation, nonlinear operation, and so on. In an example, a compute block 330 may perform convolutions, e.g., standard convolution or depthwise convolution. In some embodiments, the compute block 330 receives an input tensor and one or more convolutional kernels and performs a convolution with the input tensor and convolutional kernels. The result of the convolution may be an output tensor, which can be further computed, e.g., by the compute block 330 or another compute block 330. In some embodiments, the operations of the DNN layers may be run by multiple compute blocks 330 in parallel. For instance, multiple compute blocks 330 may each perform a portion of a workload for a convolution. Data may be shared between the compute blocks 330. A compute block 330 may also be referred to as a compute tile. In some embodiments, each compute block 330 may be a processing unit.
A compute block 330 may also referred to as a tile of the DNN accelerator 302. In the embodiments of FIG. 3, each compute block 330 includes a local memory 340, a sparsity mode module 350, a load module 360, a sparse cell array 370 (also referred to as a data processing unit) , and a drain module 380. Some or all the components of the compute block 330 can be implemented on the same chip. In other embodiments, alternative configurations, different or additional components may be included in the compute block 330. Further, functionality attributed to a component of the compute block 330 may be accomplished by a different component included in the compute block 330, a different compute block 330, another component of the DNN accelerator 302, or a different system. A component of the compute block 330 may be implemented in hardware, software, firmware, or some combination thereof.
The local memory 340 is local to the corresponding compute block 330. In the embodiments of FIG. 3, the local memory 340 is inside the compute block 330. In other embodiments, the local memory 340 may be outside the compute block 330. Data in the local memory 340 may be transferred to or from the memory 310, e.g., through the DMA engine 320. In some embodiments, data in the local memory 340 may be transferred to or from the local memory of another compute block 330. The local memory 340 may store data received, used, or generated by the sparsity mode module 350, the load module 360, the sparse cell array 370, or the drain module 380. Examples of the data may include input activations, weights, output activations, sparsity bitmaps, and so on.
In some embodiments, the local memory 340 may store dense tensors (e.g., dense activation tensors, dense weight tensors, etc. ) , sparse tensors (e.g., sparse activation tensors, sparse weight tensors, etc. ) , and so on. A dense tensor may be a tensor from which zero-valued elements (if any) are not removed. A dense tensor may be converted to a sparse tensor by removing one or more zero-valued elements in the dense tensor. A sparse tensor may also be referred to as a compressed tensor or packed tensor. The process of converting a dense tensor to a sparse tensor may be referred to as sparsity encoding. Sparsity encoding may also generate a sparsity tensor. Each element in the sparsity tensor may correspond to a different element in the dense tensor and indicate whether the element in the dense tensor is zero or not. The sparsity tensor may indicate positions of elements of the sparse tensor in the dense tensor. The sparsity tensor may be a sparsity bitmap, each element of  which is a bit. A sparse tensor may be converted to a dense tensor through a densifying process, in which one or more zeros may be added to the sparse tensor based on the sparsity tensor.
In some embodiments, the local memory 340 includes one or more static random-access memories (SRAMs) . The local memory 340 may be byte-addressable, and each memory address identifies a single byte (eight bits) of storage. In some embodiments, the local memory 340 may include memory banks. The number of data banks in the local memory 340 may be 16, 64, 128, 356, 512, 1024, 3048, or other numbers. A memory bank may include a plurality of storage units. In an example, a data bank may include 8, 16, 64, or a different number of storage units. A memory bank or a storage unit in a memory bank may have a memory address. In an example, a storage unit may store a single byte, and data larger than a single byte may be stored in storage units with consecutive memory addresses, i.e., adjacent storage units. For instance, a storage unit can store an integer number in the INT8 format, versus two storage units may be needed to store a number in the FP16 or BF16 format, which has 16 bits. In some embodiments, 16 bits can be transferred from the local memory 340 in a single read cycle. In other embodiments, 16 bits can be transferred from the local memory 340 in multiple read cycles, such as two cycles.
The sparsity mode module 350 determines sparsity modes in which the compute block 330 operates to execute DNN layers. For instance, the sparsity mode module 350 may determine whether to accelerate a layer based on weight sparsity, activation sparsity, or both. The sparsity mode module 350 select the sparsity mode for a layer from a group of sparsity modes that includes, for example, combined sparsity mode in which the layer is accelerated based on both weight sparsity and activation sparsity, activation sparsity mode in which the layer is accelerated based on activation sparsity but not based on weight sparsity, weight sparsity mode in which the layer is accelerated based on weight sparsity but not based on activation sparsity, and a dense mode in which the layer is not accelerated based on sparsity. In some embodiments (e.g., embodiments where a layer is executed by multiple compute blocks 330) , the sparsity mode module 350 may determine the sparsity mode for all the compute blocks 330 that executes the layer. In some embodiments, the sparsity mode module 350 may receive configuration parameters from the DNN module 301. A configuration parameter may correspond to a layer and indicate whether to accelerate the  layer based on weight sparsity. The sparsity mode module 350 may determine the sparsity mode of the layer based on the configuration parameter.
The load module 360 loads data from the local memory 340 to the sparse cell array 370. The load module 360 may read tensors from the local memory 340. The tensors may include sparse activation tensors, sparse weight tensors, activation sparsity tensors, weight sparsity tensors, and so on. In some embodiments, the load module 360 may load data based on the sparsity mode determined by the sparsity mode module 350. The load module 360 may select different data to transmit to the sparse cell array 370 in different sparsity modes. For instance, the load module 360 may transmit an activation sparsity tensor and a weight sparsity tensor of a layer to the sparse cell array 370 in the combined sparsity mode, while transmit the activation sparsity tensor but not the weight sparsity tensor to the sparse cell array 370 in the activation sparsity mode and transmit the weight sparsity tensor but not the activation sparsity tensor to the sparse cell array 370 in the weight sparsity mode. In the dense mode, the load module 360 does not transmit either the activation sparsity tensor or the weight sparsity tensor to the sparse cell array 370.
In some embodiments, the load module 360 may process (e.g., densify) data stored in the local memory 340 before providing the data to the sparse cell array 370. In an example, the load module 360, while operating in the weight sparsity mode, may densify sparse activation tensors to generate dense activation tensors based on corresponding activation sparsity tensors. For instance, the load module 360 may add one or more zeros into a sparse activation tensor based on an activation sparsity tensor associated with the sparse activation tensor to generate the dense activation tensor. The dense activation tensor includes one or more elements than the sparse activation tensor. The additional element (s) are zero-valued. The load module 360 may identify one or more elements in the activation sparsity tensor that correspond to the zero-valued element (s) , determine the position of each of the zero-valued element (s) in the dense activation tensor, and insert the zero-valued element (s) into the sparse activation tensor based on the determined positions. After the densification, the load module 360 may transmit the dense activation tensors to the sparse cell array 370. The load module 360 may also transmit corresponding sparse weight tensors and weight sparsity tensors to the sparse cell array 370. Activation sparsity tensor of the dense activation tensors may not be loaded to the sparse cell array 370.
In another example, the load module 360, while operating in the activation sparsity mode, may densify sparse weight tensors to generate dense weight tensors based on corresponding weight sparsity tensors by inserting zeros into sparse weight tensors. The densification of sparse weight tensors may be similar to the densification of sparse activation tensors described above. After the densification, the load module 360 may transmit the dense weight tensors to the sparse cell array 370. The load module 360 may also transmit corresponding sparse activation tensors and activation sparsity tensors to the sparse cell array 370. Weight sparsity tensor of the dense weight tensors may not be loaded to the sparse cell array 370.
In yet another example, the load module 360, while operating in the dense mode, may densify both sparse weight tensors and sparse activation tensors. The load module 360 may generate the input tensor and weight tensor of the layer and transmit the tensors to the sparse cell array 370 for executing the layer without sparsity acceleration.
The sparse cell array 370 may include one or more data processing cells. Each data processing cell may include one or more MAC units that can perform MAC operations. The MAC units in a data processing cell may be arranged in an array that includes rows and columns. The data processing cells may be arranged in one or more rows and one or more columns in the sparse cell array 370. All the MAC units in the sparse cell array 370 may constitute a bigger array that includes more rows and columns. In some embodiments (e.g., embodiments where the compute block 330 executes a convolutional layer) , a computation in an MAC unit may be an MAC operation on an activation operand and a weight operand. The activation operand may be an activation tensor that may include one or more activations in the input tensor of the convolution. Different activations may be in different input channels. The weight operand may be a weight tensor that may include one or more weights in the filter of the convolution. The values of the weights are determined through training the DNN. The weights in the weight operand may be in different input channels.
In some embodiments, an MAC unit includes one or more multipliers for performing multiplications. An MAC unit may also include one or more accumulators ( “adders” ) for performing accumulations. A column of MAC units is referred to as an MAC column. An MAC column may be associated with one or more MAC lanes. An MAC lane is a path for loading data e.g., by the load module 360, into an MAC column. An MAC lane may be also referred  to as a data transmission lane or data loading lane. An MAC column may have multiple MAC lanes. The loading bandwidth of the MAC column is an aggregation of the loading bandwidths of all the MAC lanes associated with the MAC column. With a certain number of MAC lanes, data can be fed into the same number of independent MAC units simultaneously. In some embodiments where an MAC column has four MAC lanes for feeding activations or weights into the MAC column and each MAC lane may have a bandwidth of 16 bytes, the four MAC lanes can have a total loading bandwidth of 64 bytes.
In some embodiments, the sparse cell array 370 may be capable of depthwise convolution, standard convolution, or both. In a depthwise convolution, an MAC unit may perform an MAC operation that includes a sequence of multiplications for an input operand and a weight operand. Each multiplication in the sequence (also referred to as a cycle) is a multiplication of a different activation in the input operand with a different weight in the weight operand. The activation and weight in the same cycle may correspond to the same channel. The sequence of multiplication produces a product operand that includes a sequence of products. The MAC operation may also include accumulations in which multiple product operands are accumulated to produce an output operand of the MAC unit. The sparse cell array 370 may output multiple output operands at a time, each of which is generated by a different MAC unit. In a standard convolution, MAC operations may include accumulations across the channels. For instance, as opposed to generating an output operand, a MAC unit may accumulate products across different channels to generate a single output point.
In some embodiments, the sparse cell array 370 may perform MAC operations in quantized deep learning operations, such as MAC operations in a quantized convolution. In some embodiments, an MAC unit in the sparse cell array 370 may receive quantized activation and quantized weights and compute a quantized MAC result. The quantized MAC result may be a quantized value in an integer format and may be the output of the MAC unit. In some embodiments, the MAC unit may also include a quantization multiplier that can multiply a quantization scale with the quantized MAC result, and the output of the MAC unit may be a real value in a floating-point format. The MAC unit may include no quantization subtractors as zero-point offsetting is not needed for the MAC operations in quantized deep learning operations.
In some embodiments, the sparse cell array 370 may include sparsity acceleration logic for facilitating sparsity acceleration. For instance, each data processing cell in the sparse cell array 370 may include one or more sparsity modules. In an example, each MAC column or each MAC row may have a corresponding sparsity module that accelerates MAC operations in the MAC column or MAC row. In some embodiments, a sparsity module accelerates computations in the sparse cell array 370 based on sparsity in activations, sparsity in weights, or both. The sparsity module may include a storage unit that stores a sparsity tensor, which may be loaded to the storage unit by the load module 360. The sparsity tensor may be an activation sparsity tensor, a weight sparsity tensor, or a combined sparsity tensor.
An activation sparsity tensor may be the sparsity tensor of an activation tensor and has the same number of elements as the activation tensor. An element in the activation sparsity tensor may indicate whether the corresponding element in the activation tensor is zero or not. For instance, a zero-valued in the activation sparsity tensor may indicate that the corresponding element in the activation tensor is zero. A one-valued in the activation sparsity tensor may indicate that the corresponding element in the activation tensor is nonzero. A weight sparsity tensor may be the sparsity tensor of a weight tensor and has the same number of elements as the weight tensor. An element in the weight sparsity tensor may indicate whether the corresponding element in the weight tensor is zero or not. For instance, a zero-valued in the weight sparsity tensor may indicate that the corresponding element in the weight tensor is zero. A one-valued in the weight sparsity tensor may indicate that the corresponding element in the weight tensor is nonzero. The sparsity module may generate a combined sparsity tensor using an activation sparsity tensor and a weight sparsity tensor. For instance, the sparsity module may multiply an element of the activation sparsity tensor with a corresponding element of the weight sparsity tensor to compute an element of the combined sparsity tensor. The positions of the three elements in their corresponding sparsity tensors may match. In some embodiments, each element in a sparsity tensor may be a bit, and the sparsity tensor may be referred to as a sparsity bitmap.
The sparsity module may use the sparsity tensor to identify activations and weights to be used in MAC operations by the MAC units. In an embodiment where the sparse cell array 370 operates in the combined sparsity mode, the sparsity module may identify  activations and weights that correspond to nonzero valued elements of a combined sparsity tensor. In an embodiment where the sparse cell array 370 operates in the activation sparsity mode, the sparsity module may identify activations and weights that correspond to nonzero valued elements of an activation sparsity tensor. In an embodiment where the sparse cell array 370 operates in the weight sparsity mode, the sparsity module may identify activations and weights that correspond to nonzero valued elements of a weight sparsity tensor. The sparsity module may be bypassed in the dense mode as no sparsity acceleration would be conducted.
The drain module 380 drains data from the sparse cell array 370 and writes the data to the local memory 340. The data may be outputs of MAC operations performed by MAC units in the sparse cell array 370. In some embodiments, the drain module 380 may drain data on a cell level. For each data processing cell, the drain module 380 may drain outputs of MAC units in the data processing cell based on a row index or column index of each MAC unit. For instance, the drain module 380 may use a sequence of cycles to drain data from a data processing cell. The drain module 380 may drain the output of some of the MAC units in each cycle. The sequence of the cycles may be configured based on a configuration parameter indicating the operation mode of the load module 360.
In some embodiments, the drain module 380 may determine whether to drain the output of an MAC unit based on the column index of the MAC unit when the load module operates in the activation sparsity mode versus based on the row index of the MAC unit when the load module operates in the weight sparsity mode. For instance, for MAC operations where the load module 360 operates in the activation sparsity mode, the drain module 380 may drain the output of a different MAC column in each cycle. The sequence of cycles may start with the first MAC column (e.g., the MAC column on the left side of the data processing cell) and end with the last MAC column (e.g., the MAC column on the right side of the data processing cell) . For MAC operations where the load module 360 operates in the weight sparsity mode, the drain module 380 may drain the output of a different MAC row in each cycle. The sequence of cycles may start with the first MAC row (e.g., the MAC row at the top of the data processing cell) and end with the last MAC row (e.g., the MAC column at the bottom of the data processing cell) . In other embodiments, the drain module 380 may determine whether to drain the output of an MAC unit based on the row index of the MAC  unit when the load module operates in the activation sparsity mode versus based on the column index of the MAC unit when the load module operates in the weight sparsity mode.
The drain module 380 may also include sparsity encoding logic that can convert outputs of the sparse cell array 370 from a dense format to a sparse format. For instance, the drain module 380 may be implemented with one or more sparsity encoders. A sparsity encoder converts dense data to compressed data based on sparsity in the dense data. For instance, the sparsity encoder may remove zeros in an activation tensor computed by the sparse cell array 370 to convert the activation tensor to a compressed activation tensor. The sparsity encoder may also generate sparsity tensors, including activation sparsity tensors.
In some embodiments, the data drained from the sparse cell array 370 may be at least part of an output tensor (e.g., the output tensor 230 in FIG. 2) of a deep learning operation. The sparsity encoder may generate a compressed version of the output tensor. The sparsity encoder may identify every zero-valued activation in the output tensor and remove these activations from the output tensor to generate a compressed activation tensor (aka “sparse activation tensor” ) . The sparsity encoder may also generate one or more sparsity tensors for the output tensor. A sparsity tensor may correspond to a portion of the output tensor (e.g., the vector 235 in FIG. 2) . The sparsity tensor may include sparsity elements (e.g., bits) , each of which corresponds to a different activation in the vector and indicates whether the corresponding activation is zeroed or not.
The drain module 380 may write the compressed activation tensor and the one or more sparsity tensors into the local memory 340. The sparse activation tensor and the one or more sparsity tensors may be further loaded to the memory 310, e.g., through the DMA engine 320. Additionally or alternatively, the sparse activation tensor and the one or more sparsity tensors may be loaded by the load module 360 to the sparse cell array for further computation, e.g., for performing a deep learning operation in the next layer.
FIG. 4 illustrates an example data processing cell 400, in accordance with various embodiments. The data processing cell 400 may be in a sparse cell array, e.g., the sparse cell array 370 in FIG. 3. The data processing cell 400 includes 16 MAC units 410 (individually referred to as “MAC unit 410” ) , which constitutes a MAC array having four rows and four columns. The MAC array has a spatial shape of 4x4, meaning the height of the MAC array is four and the width of the MAC array is also 4. The data processing cell 400 also includes 16  weight register files 420 (individually referred to as “weight register file 420” ) , 16 activation register files 430 (individually referred to as “activation register file 430” ) , four row buffers 440 (individually referred to as “row buffer 440” ) , and sparsity modules 460 (individually referred to as “sparsity module 460” ) . In other embodiments, the data processing cell 400 may include fewer, more, or different components. For example, the data processing cell 400 may include a different number of MAC units 410, weight register files 420, activation register files 430, row buffers 440, or sparsity modules 460. As another example, the data processing cell 400 may include column buffers in lieu of or in addition to the row buffers 440. Also, the shape (e.g., the height or width) of the MAC array may be different.
The MAC units 410 are configured to perform MAC operations. Each MAC unit 410 may include one or more multipliers and one or more adders. A multiplier may multiply an activation with a weight at a time to compute a product. In some embodiments (e.g., embodiments where the MAC unit 410 includes multiple multipliers) , the multipliers may operate simultaneously to process multiple activation-weight pairs and compute multiple products in one cycle. An adder may accumulate products computed by the multipliers. Even though not shown in FIG. 4, the data processing cell may include an adder tree including a plurality of adder tiers. The first tier may receive outputs of a plurality of MAC units 410. The number of adders in the first tier may be half of the number of the MAC units 410, and each adder may accumulate the outputs of two MAC units 410. The second tier may receive outputs of adders in the first tier. The number of adders in the second tier may be half of the number of adders in the first tier, and each adder in the second tier may accumulate the outputs of two adders in the first tier. The adder tree may include one or more other tiers. The last tier may include a single adder that accumulates outputs of adders in the second last tier to compute a partial sum of the data processing cell 400.
The weight register files 420 store weights to be processed in MAC operations. In the embodiments of FIG. 4, four weight register files 420 are grouped into a storage set that stores data to be used by a column of MAC units 410. There are four storage sets corresponding to the four columns of MAC units 410. In some embodiments, a weight register file 420 may correspond to a MAC unit 410 and store data to be processed by the MAC unit. In some embodiments, all the 16 weight register files 420 constitute a weight storage unit.
The activation register files 430 stores activations to be processed in MAC operations. In the embodiments of FIG. 4, four activation register files 430 are grouped into a storage set that stores data to be used by a row of MAC units 410. There are four storage sets corresponding to the four rows of MAC units 410. In some embodiments, an activation register file 430 may correspond to a MAC unit 410 and store data to be processed by the MAC unit. In some embodiments, all the 16 activation register files 430 constitute an activation storage unit. The row buffers 440 store outputs of the MAC units 410. Each row buffer 440 may drain outputs of a single row of MAC units 410.
The sparsity module 460 facilitates dynamic sparsity-based acceleration in the data processing cell 400. In the embodiments of FIG. 4, each sparsity module 460 includes a sparsity tensor storage unit 465 and a control logic 467. The sparsity tensor storage unit 465 stores combined sparsity tensors. A combined sparsity tensor stored in the sparsity tensor storage unit 465 may correspond to an activation tensor and a weight tensor. A nonzero element in the combined sparsity tensor may correspond to a nonzero activation-weight pair that includes a nonzero activation and a nonzero weight. The position of the nonzero activation in the activation tensor may match the position of the nonzero weight in the weight tensor. The product of the nonzero activation and nonzero weight would be nonzero.
The control logic 467 may control transmission of activations and weights stored from the weight register files 420 and the activation register files 430 to the MAC units 410 based on sparsity tensors. For instance, the control logic 467 may select a subset of the weights stored in the weight register files 420 and select a subset of activations stored in the activation register files 430 based on a combined sparsity tensor. The selected weights and activations constitute nonzero activation-weight pairs. The control logic 467 may transmit the selected weights and activations to the MAC units 410 for performing MAC operations. The other weights stored in the weight register files 420 and the other activations stored in the activation register files 430 are skipped from computation. In the embodiments of FIG. 4, each sparsity module 460 controls sparsity acceleration in a respective MAC unit 410. As the sparsity acceleration is either based on both weight sparsity and activation sparsity, 16 sparsity modules 460 are used for acceleration computations in the 16 MAC units 410.
As shown in FIG. 4, the data processing cell 400 is associated with multiplexers (MUXs) 403, 404, 405, and 406. In other embodiments, the data processing cell 400 may be  associated with a different number of MUXs or other devices. The MUX 403 facilitates loading weights, e.g., from the local memory 340, into the weight register files 420. The MUX 404 facilitates loading activations, e.g., from the local memory 340, into the activation register files 430. The MUX 405 facilitates loading sparsity tensors into the sparsity tensor storage unit 465. The MUX 406 may be a drain MUX that can facilitate draining outputs of the MAC units 410, e.g., to the local memory 340.
In some embodiments, the data processing cell 400 may also execute matrix multiplications converted from Fourier transform operations. For an example Fourier transform operation, the MAC units 410 may perform MAC operations in the two sequences of matrix multiplications converted from the Fourier transform operation. The weight register files 420 may be used to store data points in transformation tensor of the Fourier transform operation. The activation register file 430 may be used to store data points in the input tensor of the Fourier transform operation. The row buffers 440 may store data points in the output tensor of the Fourier transform operation.
FIG. 5 illustrates a data processing unit 500, in accordance with various embodiments. The data processing unit 500 may be an example of the sparse cell array 370 in FIG. 3. In FIG. 5, the data processing unit 500 includes data processing cells 510 (individually referred to as “data processing cell 510” ) arranged in four columns and four rows, an activation memory 520, and a weight memory 530. The data processing unit 500 may also be referred to as a data processing unit. In other embodiments, the data processing unit 500 may include fewer, more, or different components. For instance, the data processing unit 500 may include a different number of columns, rows, or data processing cells 510.
Each data processing cell 510 may perform sparsity accelerated MAC operations. The data processing cells 510 may facilitate dynamic sparsity mode. For instance, the sparsity modes of a data processing cell 510 may be dynamically changed between a combined sparsity mode, an activation sparsity mode, a weight sparsity mode, and a dense mode. An embodiment of a data processing cell 510 may be the data processing cell 400 in FIG. 4. The activation memory 520 stores activations, such as activations in input tensors of deep learning operations. Activations may be loaded from the activation memory 520 to data processing cells 510. The weight memory 530 stores weights, such as weights in filters of  deep learning operations. Weights may be loaded from the weight memory 530 to data processing cells 510. The activation memory 520 or weight memory 530 may be a buffer. In other embodiments, the data processing unit 500 may include a dense data memory and a sparse data memory in lieu of the activation memory 520 and weight memory 530. The dense data memory may store dense tensors, e.g., dense tensors generated by the load module 360. The sparse data memory may store sparse tensors.
The data processing unit 500 may also execute matrix multiplications in Fourier transform operations. The activation memory 520 may be used to store input tensors of the Fourier transform operations. The weight memory 530 may be used to store transformation matrices of the Fourier transform operations.
FIG. 6 is a block diagram of a DNN module 600, in accordance with various embodiments. The DNN module 600 facilitates transformation of matrix multiplications to convolutions. The DNN module 600 may be an embodiment of the DNN module 301 in FIG. 3. As shown in FIG. 6, the DNN module 600 includes an interface module 610, a training module 620, a compressing module 630, a validating module 640, a reshaping module 650, and a datastore 660. In other embodiments, alternative configurations, different or additional components may be included in the DNN module 600. Further, functionality attributed to a component of the DNN module 600 may be accomplished by a different component included in the DNN module 600 or a different module or system.
The interface module 610 facilitates communications of the DNN module 600 with other modules or systems. For example, the interface module 610 establishes communications between the DNN module 600 with an external database to receive data that can be used to train DNNs or input into DNNs to perform tasks. As another example, the interface module 610 supports the DNN module 600 to distribute DNNs to other systems, e.g., computing devices configured to apply DNNs to perform tasks.
The training module 620 trains DNNs by using a training dataset. The training module 620 forms the training dataset. In an embodiment where the training module 620 trains an DNN to recognize objects in images, the training dataset includes training images and training labels. The training labels describe ground-truth classifications of objects in the training images. In some embodiments, each label in the training dataset corresponds to an object in a training image. In some embodiments, a part of the training dataset may be used  to initially train the DNN, and the rest of the training dataset may be held back as a validation subset used by the validating module 640 to validate performance of a trained DNN. The portion of the training dataset not including the tuning subset and the validation subset may be used to train the DNN.
The training module 620 also determines hyperparameters for training the DNN. Hyperparameters are variables specifying the DNN training process. Hyperparameters are different from parameters inside the DNN (e.g., weights of filters) . In some embodiments, hyperparameters include variables determining the architecture of the DNN, such as number of hidden layers, etc. Hyperparameters also include variables which determine how the DNN is trained, such as batch size, number of epochs, etc. A batch size defines the number of training samples to work through before updating the parameters of the DNN. The batch size is the same as or smaller than the number of samples in the training dataset. The training dataset can be divided into one or more batches. The number of epochs defines how many times the entire training dataset is passed forward and backwards through the entire network. The number of epochs defines the number of times that the deep learning algorithm works through the entire training dataset. One epoch means that each training sample in the training dataset has had an opportunity to update the parameters inside the DNN. An epoch may include one or more batches. The number of epochs may be 1, 5, 10, 50, 100, 500, 1000, or even larger.
The training module 620 defines the architecture of the DNN, e.g., based on some of the hyperparameters. The architecture of the DNN includes an input layer, an output layer, and a plurality of hidden layers. The input layer of an DNN may include tensors (e.g., a multidimensional array) specifying attributes of the input image, such as the height of the input image, the width of the input image, and the depth of the input image (e.g., the number of bits specifying the color of a pixel in the input image) . The output layer includes labels of objects in the input layer. The hidden layers are layers between the input layer and output layer. The hidden layers include one or more convolutional layers and one or more other types of layers, such as pooling layers, fully-connected layers, normalization layers, SoftMax or logistic layers, and so on. The convolutional layers of the DNN abstract the input image to a feature map that is represented by a tensor specifying the feature map height, the feature map width, and the feature map channels (e.g., red, green, blue images include 3  channels) . A pooling layer is used to reduce the spatial volume of input image after convolution. It is used between two convolution layers. A fully-connected layer involves weights, biases, and neurons. It connects neurons in one layer to neurons in another layer. It is used to classify images between different categories by training.
In the process of defining the architecture of the DNN, the training module 620 also adds an activation function to a hidden layer or the output layer. An activation function of a layer transforms the weighted sum of the input of the layer to an output of the layer. The activation function may be, for example, a ReLU activation function, a tangent activation function, or other types of activation functions.
After the training module 620 defines the architecture of the DNN, the training module 620 inputs a training dataset into the DNN. The training dataset includes a plurality of training samples. An example of a training sample includes an object in an image and a ground-truth label of the object. The training module 620 modifies the parameters inside the DNN ( “internal parameters of the DNN” ) to minimize the error between labels of the training objects that are generated by the DNN and the ground-truth labels of the objects. The internal parameters include weights of filters in the convolutional layers of the DNN. In some embodiments, the training module 620 uses a cost function to minimize the error.
The training module 620 may train the DNN for a predetermined number of epochs. The number of epochs is a hyperparameter that defines the number of times that the deep learning algorithm will work through the entire training dataset. One epoch means that each sample in the training dataset has had an opportunity to update internal parameters of the DNN. After the training module 620 finishes the predetermined number of epochs, the training module 620 may stop updating the parameters in the DNN. The DNN having the updated parameters is referred to as a trained DNN.
The compressing module 630 compresses DNNs. For instance, the compressing module 630 may add pruning operations to DNN layers to reduce computational complexity or memory usage. A pruning operation may prune weight tensors of a DNN layer by changing one or more nonzero valued weights of the layer to zeros. The modification may be done before, during, or after training. Weights may be pruned during training, during inference, or a combination of both. The compressing module 630 may determine a sparsity ratio for a DNN layer. The sparsity ratio may be a ratio of the number of zero-valued weight  to the total number of weights in the layer. The compressing module 630 may perform the pruning operation till the sparsity ratio of the DNN layer meets a target sparsity ration, such as 10%, 20%, 30%, 60%, 50%, and so on.
In some embodiments, the compressing module 630 may select one or more layers in a DNN and modify each selected layer with a pruning operation. For instance, the compressing module 630 may select computationally complex layers, such as layers with large filters. For a pruning operation of a layer or of a type of layer, the compressing module 630 may determine a weight threshold that would not cause a loss of the accuracy of the DNN to exceed an accuracy loss constraint. A pruning operation may modify weights having absolute values above the weight threshold to zeros and leave the other weights unchanged. The weight pruning can reduce memory storage as zero-valued weights may not be stored. Also, the number of operations in the layer can be reduced as computations on zero-valued weights can be skipped without impacting the output of the layer. In some embodiments, the compressing module 630 may also measure energy saving, final DNN accuracy, or layer-wise sparsity caused by pruning operations.
After compressing a DNN, the compressing module 630 may fine tune the DNN, e.g., through a retraining process. The compressing module 630 may fine tunes DNNs after weights are pruned. In some embodiments, the fine-tuning process is a retraining or further training process. For instance, after weights in a DNN are pruned, the compressing module 630 may further train the DNN by inputting a training dataset into the DNN. The values of the unpruned weights in the DNN may be modified based on outputs of the DNN and ground-truth labels of the training samples in the training dataset. In some embodiments, the values of the pruned weights (i.e., zero) are not changed during the fine-tuning process. For instance, the compressing module 630 may place a mask over a pruned weight block and the mask can prevent values in the pruned weight blocks from being changed during the fine-tuning process. In other embodiments, the values of all weights, including the pruned weights, may be changed during the fine-tuning process. After one or more cycles of retraining and weight changing by the compressing module 630, the compressing module 630 may perform a new pruning process, e.g., by selecting weight blocks and pruning the selected weight blocks. In some embodiments, the weight pruning process may be repeated multiple times before the fine-tuning process is done.
In some embodiments, the number of epochs in the fine-tuning process may be different from the number of epochs in the training process in which the pre-pruning values of the weights are determined. For instance, the fine-tuning process may have less epochs than the training process. In an example, the number of epochs in the fine-tuning process may be relatively small, such as 2, 3, 6, 5, and so on.
The validating module 640 verifies accuracy of trained or compressed DNNs. In some embodiments, the validating module 640 inputs samples in a validation dataset into a trained DNN and uses the outputs of the DNN to determine the model accuracy. In some embodiments, a validation dataset may be formed of some or all the samples in the training dataset. Additionally or alternatively, the validation dataset includes additional samples, other than those in the training sets. In some embodiments, the validating module 640 may determine an accuracy score measuring the precision, recall, or a combination of precision and recall of the DNN. The validating module 640 may use the following metrics to determine the accuracy score: Precision = TP / (TP + FP) and Recall = TP / (TP + FN) , where precision may be how many the DNN correctly predicted (TP or true positives) out of the total it predicted (TP + FP or false positives) , and recall may be how many the DNN correctly predicted (TP) out of the total number of objects that did have the property in question (TP +FN or false negatives) . The F-score (F-score = 2 *PR / (P + R) ) unifies precision and recall into a single measure.
The validating module 640 may compare the accuracy score with a threshold score. In an example where the validating module 640 determines that the accuracy score of the DNN is less than the threshold score, the validating module 640 instructs the training module 620 to re-train the DNN. In one embodiment, the training module 620 may iteratively re-train the DNN until the occurrence of a stopping condition, such as the accuracy measurement indication that the DNN may be sufficiently accurate, or a number of training rounds having taken place.
The reshaping module 650 reshapes deep learning operations based on configurations of the DNN accelerator 302. In an example, the reshaping module 650 may reshape convolution based on configurations of MAC arrays in the DNN accelerator 302 to increase the utilization of MAC units in the MAC arrays while performing the convolution and improve the computational efficiency of the convolutions. In some embodiments, the  reshaping module 650 may analyze the shape of the output tensor of a convolution, e.g., based on the shape of the activation tensor and weigh tensor of the convolution. The reshaping module 650 may determine the height and width of the output tensor in a single output channel. In some embodiments (e.g., embodiments where the kernel size is 1×1) , the height and width of the output tensor may equal to the height and width of the activation tensor, respectively. In other embodiments, the height or width of the output tensor may be smaller than the height or width of the activation tensor, respectively.
The reshaping module 650 may determine whether the product of the height and the weight of the output tensor is a multiple of the height or width of a MAC array in the DNN accelerator. The product of the height and the weight of the output tensor may be the total number of output activations in a single output channel. In embodiments where the product of the height and the weight of the output tensor is not a multiple of the height or width of the MAC array, the reshaping module 650 may perform padding on the activation tensor to change the product of the height and the weight of the output tensor to a multiple of the height or width of the MAC array. The reshaping module 650 may perform the padding by adding one or more zeros to one or more boundaries of the activation tensor. The reshaping module 650 may add one or more columns (or rows) of zeros into the activation tensor to increase the width (or height) of the activation tensor to make the product of the height and the weight of the output tensor to a multiple of the height or width of the MAC array. In some embodiments, the reshaping module 650 may make no change to the number of input channels in the activation tensor. Even though the activation tensor is reshaped in various examples, the reshaping module 650 may reshape the weight tensor in some embodiments.
After the reshaping module 650 determines that the product of the height and the weight of the output tensor is a multiple of the height or width of the MAC array with or without padding, the reshaping module 650 may change the height or width of the activation tensor to ensure that the height or width of the output tensor is a multiple of the height or width of the MAC array. The reshaping module 650 may keep the product of the height and width of the activation tensor the same, meaning the reshaping module 650 may add no new activations into the activation tensor. Rather, the reshaping module 650 changes the arrangement of the activations in the activation tensor. In an example where the output tensor of the convolution has a spatial size of 128×2 in each channel, the reshaping module  650 may reshape the activation tensor to change the spatial size of the output tensor in each channel to 16×16 based on a determination that the MAC array is an 8×8 array.
In some embodiments, there may be multiple options to reshape the activation tensor to ensure that the output tensor has a height or width that is a multiple of the height or width of the MAC array. For instance, the 128×2 spatial size in the example described above may be changed to 16×16, 32×8, or 8×32, which are three different options for reshaping the convolution. The reshaping module 650 may determine the computational efficiency for each of the options and select the option that can achieve the best computational efficiency. In an example where the MAC array is a N×N grid, the reshaping module 650 may determine the compute efficiency using the following algorithm:
where ceil denotes the ceil function that returns the smallest integer value that is greater than or equal to the number input into the ceil function, Wout is the width of the output tensor, and Hout is the height of the output tensor. For the example above, the reshaping module 650 may determine that the spatial size 16×16 would result in the best computational efficiency and may select to changing the 128×2 spatial size to 16×16, as opposed to 32×8 or 8×32. The reshaping module 650 may also generate one or more configuration parameters that indicate how the activation tensor is reshaped. Such configuration parameters may be referred to as reshaping parameters, which the reshaping module 650 may use to generate the output tensor of the original convolution.
In some embodiments, the reshaping module 650 may also determine whether to change the stride of the convolution to ensure that the output tensor of the reshaped convolution would have the same elements (even though different arrangements of the elements) as the output tensor of the original convolution. The reshaping module 650 may determine the stride of the reshaped convolution based on the shape of the activation tensor and the shape of the weight tensor.
After the reshaping module 650 reshapes the activation tensor, the reshaping module 650 may facilitate the sparse cell array 370 to read activations in the activation tensor in accordance with the new shape of the activation tensor. In some embodiments, the sparse cell array 370 may read an activation operand at a time and send the activation operand to a MAC unit for processing in a single computational cycle. The activation operand  may also be referred to as a context. An example of the activation operand may be a vector in the activation tensor, e.g., at least part of a row or a column in a channel of the activation tensor. As the shape of the activation tensor is changed, the activations in each activation operand may change. The reshaping module 650 may generate configuration parameters that configure the sparse cell array 370 to read the activation operands in the reshaped activation tensor. In some embodiments, the reshaping module 650 may configure storage pointers associated with the memory 310 or the local memory 340. The storage pointers may indicate the boundary of the activation operands, e.g., the memory address where the first activation in every activation operand is stored, the memory address where the last activation in every activation operand is stored, and so on. The sparse cell array 370 may read activation operands based on the storage pointers. The layout of the activations in the memory 310 or the local memory 340 may remain the same.
After the sparse cell array 370 performs the reshaped convolution on the reshaped activation tenor and the weight tensor, the reshaping module 650 may receive the output tensor of the reshaped convolution from the DNN accelerator 302. The reshaping module 650 may further transform the output tensor of the reshaped convolution to generate the output tensor of the original convolution. For instance, the reshaping module 650 may modify the shape of the output tensor of the reshaped convolution based on the reshaping parameter (s) . In some embodiments, the reshaping module 650 may modify the shape of the output tensor of the reshaped convolution in each output channel. The number of output channels of the reshaped convolution may be the same as the total number of output channels of the original convolution.
The datastore 660 stores data received, generated, used, or otherwise associated with the DNN module 600. For example, the datastore 660 stores the datasets used by the training module 620 and validating module 640. The datastore 660 may also store data generated by the training module 620 and validating module 640, such as the hyperparameters for training DNNs, internal parameters of trained DNNs (e.g., weights, etc. ) , data for sparsity acceleration (e.g., sparsity bitmap, etc. ) , and so on. The datastore 660 may store tensors (e.g., transposed tensors, reshaped tensors, etc. ) , shape representations of tensors, configuration parameters, or other data generated by the reshaping module 650. In the embodiment of FIG. 6, the datastore 660 is a component of the DNN module 600. In  other embodiments, the datastore 660 may be external to the DNN module 600 and communicate with the DNN module 600 through a network.
Reshaping Convolutions
FIG. 7 illustrates a convolution before reshaping, in accordance with various embodiments. The convolution has an activation tensor 710, a weight tensor 720, and an output tensor 730. The activation tensor 710 has a spatial size Hin×Win×Cin, where Hin is the height of the 3D matrix (i.e., the length along the Y axis, which indicates the number of activations in a column in the 3D matrix of each input channel) , Win is the width of the 3D matrix (i.e., the length along the X axis, which indicates the number of activations in a row in the 2D matrix of each input channel) , and Cin is the depth of the 3D matrix (i.e., the length along the Z axis, which indicates the number of input channels) . The weight tensor 720 includes multiple filters 725. Each filter 725 includes weights arranged in a 3D matrix. The values of the weights may be determined through training the DNN. A filter 725 has a spatial size Hf×Wf×Cf, where Hf is the height of the filter (i.e., the length along the Y axis, which indicates the number of weights in a column in each kernel) , Wf is the width of the filter (i.e., the length along the X axis, which indicates the number of weights in a row in each kernel) , and Cf is the depth of the filter (i.e., the length along the Z axis, which indicates the number of channels) . In some embodiments, Cf equals Cin.
The convolution may be performed by the DNN accelerator 302 in FIG. 3, which processes the activation tensor 710 and weight tensor 720 and computes the output tensor 730. The output tensor 730 has a spatial size Hout×Wout×Cout, where Hout is the height of the 3D matrix (i.e., the length along the Y axis, which indicates the number of output activations in a column in the 2D matrix of each output channel) , Wout is the width of the 3D matrix (i.e., the length along the X axis, which indicates the number of output activations in a row in the 2D matrix of each output channel) , and Cout is the depth of the 3D matrix (i.e., the length along the Z axis, which indicates the number of output channels) . Cout may equal the number of filters 725 in the weight tensor 720. Hout and Wout may depend on the heights and weights of the activation tensor 210 and each filter 725.
The workload of performing the convolution may be represented by a 4D tensor Cout×Hout×Wout×Cin. The DNN accelerator performing the convolution may have a N× N MAC array, meaning the MAC array has N columns and N rows. The compute efficiency of the MAC array may be denoted as:
where ceil denotes the ceil function that returns the smallest integer value that is greater than or equal to the number input into the ceil function. When N is a fixed number, the computational efficiency can change as Hout or Wout changes. For instance, the computational efficiency for a convolution workload with Hout and Wout each being a multiple of N may be higher than the computational efficiency for a convolution workload with Hout or Wout not being a multiple of N for computing the same number of output elements. For some convolution workloads, the computational efficiency can be improved by changing Hout and Wout.
FIG. 8 illustrates a reshaped convolution, in accordance with various embodiments. The reshaped convolution may be converted from the convolution in FIG. 7. For the purpose of illustration, the convolution in FIG. 7 is referred to as the original convolution.
In some embodiments, the reshaped convolution has an activation tensor 810, a weight tensor 820, and an output tensor 830. The activation tensor 810 has a spatial size H′in×W′in×C′in. The weight tensor 820 includes multiple filters 825. Each filter 825 includes weights arranged in a 3D matrix. The values of the weights may be determined through training the DNN. A filter 825 has a spatial size H′f×W′f×C′f. The output tensor 830 has a spatial size H′out×W′out×C′out. The activation tensor 810 may be generated by reshaping the activation tensor 710. In an example, C′inmay remain the same, but Hin and Win are changed. Hin×Win may be the same as H′in×W′in. The weight tensor 820 may be the same as the weight tensor 720, meaning the number of filters 825 in the weight tensor 820 may equal the number of filters 725 in the weight tensor 720 and the shape of each filter 825 may be the same as the shape of each filter 725. As the activation tensor 810 is reshaped, the shape of the output tensor 830 may be different from the shape of the output tensor 730. For instance, Hout is different from H′out, and Wout may be different from W′out. H′out×W′out may be the same as Hout×Wout. The workload of performing the convolution is changed from Cout×Hout×Wout×Cin to Cout×H′out×W′out×Cin.
The reshaping may be done based on the configuration of the MAC array, e.g., based on the height or width of the MAC array. In an example where Cout×Hout×Wout×Cin is  256x2048x1x48 and the height or width of the MAC array is 4, Cout×H′ou×W′tut×Cin is set to 256x128x16x48 to increase utilization of MAC units in the MAC array and increase computation. Even though the number of output elements is the same for both the original convolution and the reshaped convolution, the reshaped convolution workload would take less computational cycles than the original convolution workload as the utilization of the MAC units in the MAC array would be higher.
FIG. 9 illustrates reshaping of an activation tensor in a channel, in accordance with various embodiments. For the purpose of illustration and simplicity, FIG. 9 shows a matrix 910, which is a 2D matrix in a single channel of the activation tensor. The matrix 910 has a spatial shape of 2048×1.2048 may be the spatial height of the activation tensor. 1 may be the spatial width of the activation tensor.
The matrix 910 is reshaped to generate a matrix 920. The matrix 920 has a spatial shape of 128×16. The total number of activations in the matrix 920 is the same as the total number of activations in the matrix 910. However, the arrangement of the activations is changed by the reshaping. In some embodiments, the reshaping may not require data movement in the memory (e.g., the memory 310 or the local memory 340) where the activations are stored. The layout of the activations in the memory may remain the same, but the MAC array may read the activations from the memory in a different pattern that matches the shape of the matrix 920. In some embodiments, the MAC array reads activations from the memory at a context level. A context may be an operand to be processed by a MAC unit in a computational cycle. For the matrix 910, a context may have one activation. For the matrix 920, a context may have 16 activations. Each context may be associated with a storage pointer that indicates a memory address of the context (e.g., the memory address where the first activation in the context is stored) . The reshaping may be performed without making any data movement in the memory. Rather, the storage pointers for reading the contexts may be changed so that the MAC array may read 16 activations per channel to compute a stencil as opposed to reading one activation per channel to compute a stencil.
FIG. 10 illustrates computing a channel of an output tensor of a convolution without reshaping, in accordance with various embodiments. The convolution may be an example of the convolution in FIG. 7. The output tensor in the embodiment of FIG. 10 may be an  example of the output tensor 730 in FIG. 7. For the purpose of illustration and simplicity, the convolution in FIG. 10 has an activation tensor 1010 with a shape of 2048×1×48 and kernels 1020 having a shape of 1×1. Each channel of the output tensor is a tenor 1030 having a shape of 2048×1.
FIG. 11 illustrates computing a channel of an output tensor of a reshaped convolution, in accordance with various embodiments. The reshaped convolution may be an example of the reshaped convolution in FIG. 8. The output tensor in the embodiment of FIG. 10 may be an example of the output tensor 830 in FIG. 8. For the purpose of illustration and simplicity, the reshaped convolution in FIG. 11 has an activation tensor 1110 with a shape of 128×16×48 and kernels 1120 having a shape of 1×1. The kernels 1120 may be the same as the kernel 1020 in FIG. 10. Each channel of the output tensor is a tenor 1130 having a shape of 128×16.
In some embodiments, the reshaped convolution in FIG. 11 is computationally identical to the convolution in FIG. 10, meaning each input channel convolves with its corresponding filter, with the result being the sum of all these convolutions. Despite the reshaping, the computation remains the same so that the computational accuracy is not impacted while the computational efficiency is improved.
FIG. 12 illustrates an example process 1200 of improving computational efficiency of deep learning operations, in accordance with various embodiments. For the purpose of illustration and simplicity, the process 1200 is a process of improving computational efficiency of a convolution in a DNN. A reshaping module 1210 receives an activation tensor 1201 of the convolution and generates a reshaped activation tensor 1202 from the activation tensor 1201. A convolution operator 1220 performs a reshaped convolution on the reshaped activation tensor 1202 and a weight tensor 1203 and computes an output tensor 1204. The reshaping module 1210 receives the output tensor 1204 and generates a reshaped output tensor 1205 from the output tensor 1204. The reshaped output tensor 1205 may be the output of the original convolution, which may be processed in the next deep learning operation of the DNN. An example of the reshaping module 1210 may be the reshaping module 650 in FIG. 6. An example of the convolution operator 1220 may include the data processing cell 900 in FIG. 9 or the data processing unit 1000 in FIG. 10.
Example Method of Executing Convolutions
FIG. 13 is a flowchart showing a method 1300 of executing a convolution, in accordance with various embodiments. The method 1300 may be performed by the DNN system 300 in FIG. 3. Although the method 1300 is described with reference to the flowchart illustrated in FIG. 13, many other methods for executing convolutions may alternatively be used. For example, the order of execution of the steps in FIG. 13 may be changed. As another example, some of the steps may be changed, eliminated, or combined.
The DNN system 300 receives 1310 an activation tensor of a convolution. The activation tensor has one or more dimensions and a shape defined by the one or more dimensions. In some embodiments, the activation tensor comprises activations arranged in the one or more dimensions.
The DNN system 300 generates 1320 a reshaped activation tensor by modifying the shape of the activation tensor based on the structure of a MAC array. The MAC array comprises MAC units. In some embodiments, the MAC units are arranged in one or more rows and one or more columns. In some embodiments, the DNN system 300 modifies the shape of the activation tensor by changing a total number of activations in a dimension of the activation tensor so that a total number of elements in a dimension of the output tensor of the reshaped convolution is a multiple of a total number of MAC units arranged in a row or column of the MAC array.
In some embodiments, the DNN system 300 modifies the shape of the activation tensor by adding one or more zero-valued activations into the activation tensor. The DNN system 300 determines the spatial size of the activation tensor in a channel of the convolution. The DNN system 300 also determines whether the spatial size is a multiple of a total number of MAC units arranged in a row or column of the MAC array. The one or more zero-valued activations are added into the activation tensor in response to a determination that the spatial size is a multiple of a total number of MAC units arranged in a row or column of the MAC array.
In some embodiments, the activation tensor comprises a plurality of input channels. An input channel corresponds to a matrix including activations arranged in one or more rows and one or more columns. The DNN system 300 modifies the shape of the activation tensor by modifying the height or width of the matrix.
The DNN system 300 computes 1330, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution. In some embodiments, a total number of channels in the activation tensor is equal to the total number of channels in the reshaped activation tensor.
In some embodiments, a stride of the reshaped convolution is different from a stride of the convolution. In some embodiments, the DNN system 300 determines the stride of the reshaped convolution based on the shape of the activation tensor and a shape of the weight tensor.
In some embodiments, the DNN system 300 stores the activations in a memory based on the shape of the activation tensor. The DNN system 300 reads, by the MAC array, the activations from the memory based on a shape of the reshaped activation tensor. In some embodiments, the DNN system 300 performs, by the MAC array, the reshaped convolution on the reshaped activation tensor and a weight tensor of the convolution.
The DNN system 300 generates 1340 an output tensor of the convolution by modifying a shape of the output tensor of the reshaped convolution. In some embodiments, the DNN system 300 modifies the shape of the output tensor of the reshaped convolution by modifying the height or width of a matrix in the output tensor of the reshaped convolution.
Example Computing Device
FIG. 14 is a block diagram of an example computing device 1400, in accordance with various embodiments. In some embodiments, the computing device 1400 can be used as at least part of the DNN system 300. A number of components are illustrated in FIG. 14 as included in the computing device 1400, but any one or more of these components may be omitted or duplicated, as suitable for the application. In some embodiments, some or all of the components included in the computing device 1400 may be attached to one or more motherboards. In some embodiments, some or all of these components are fabricated onto a single system on a chip (SoC) die. Additionally, in various embodiments, the computing device 1400 may not include one or more of the components illustrated in FIG. 14, but the computing device 1400 may include interface circuitry for coupling to the one or more components. For example, the computing device 1400 may not include a display device 1406, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which a display device 1406 may be coupled. In another set of examples, the  computing device 1400 may not include an audio input device 1418 or an audio output device 1408 but may include audio input or output device interface circuitry (e.g., connectors and supporting circuitry) to which an audio input device 1418 or audio output device 1408 may be coupled.
The computing device 1400 may include a processing device 1402 (e.g., one or more processing devices) . The processing device 1402 processes electronic data from registers and/or memory to transform that electronic data into other electronic data that may be stored in registers and/or memory. The computing device 1400 may include a memory 1404, which may itself include one or more memory devices such as volatile memory (e.g., DRAM) , nonvolatile memory (e.g., read-only memory (ROM) ) , high bandwidth memory (HBM) , flash memory, solid state memory, and/or a hard drive. In some embodiments, the memory 1404 may include memory that shares a die with the processing device 1402. In some embodiments, the memory 1404 includes one or more non-transitory computer-readable media storing instructions executable to perform operations for executing convolutions (e.g., the method 1300 described in conjunction with FIG. 13) or some operations performed by the DNN system 300. The instructions stored in the one or more non-transitory computer-readable media may be executed by the processing device 1402.
In some embodiments, the computing device 1400 may include a communication chip 1412 (e.g., one or more communication chips) . For example, the communication chip 1412 may be configured for managing wireless communications for the transfer of data to and from the computing device 1400. The term "wireless" and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data through the use of modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not.
The communication chip 1412 may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.10 family) , IEEE 802.16 standards (e.g., IEEE 802.16-2005 Amendment) , Long-Term Evolution (LTE) project along with any amendments, updates, and/or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as "3GPP2" ) , etc. ) . IEEE 802.16 compatible Broadband Wireless Access  (BWA) networks are generally referred to as WiMAX networks, an acronym that stands for worldwide interoperability for microwave access, which is a certification mark for products that pass conformity and interoperability tests for the IEEE 802.16 standards. The communication chip 1412 may operate in accordance with a Global System for Mobile Communication (GSM) , General Packet Radio Service (GPRS) , Universal Mobile Telecommunications System (UMTS) , High Speed Packet Access (HSPA) , Evolved HSPA (E-HSPA) , or LTE network. The communication chip 1412 may operate in accordance with Enhanced Data for GSM Evolution (EDGE) , GSM EDGE Radio Access Network (GERAN) , Universal Terrestrial Radio Access Network (UTRAN) , or Evolved UTRAN (E-UTRAN) . The communication chip 1412 may operate in accordance with Code-division Multiple Access (CDMA) , Time Division Multiple Access (TDMA) , Digital Enhanced Cordless Telecommunications (DECT) , Evolution-Data Optimized (EV-DO) , and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. The communication chip 1412 may operate in accordance with other wireless protocols in other embodiments. The computing device 1400 may include an antenna 1422 to facilitate wireless communications and/or to receive other wireless communications (such as AM or FM radio transmissions) .
In some embodiments, the communication chip 1412 may manage wired communications, such as electrical, optical, or any other suitable communication protocols (e.g., the Ethernet) . As noted above, the communication chip 1412 may include multiple communication chips. For instance, a first communication chip 1412 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second communication chip 1412 may be dedicated to longer-range wireless communications such as global positioning system (GPS) , EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some embodiments, a first communication chip 1412 may be dedicated to wireless communications, and a second communication chip 1412 may be dedicated to wired communications.
The computing device 1400 may include battery/power circuitry 1414. The battery/power circuitry 1414 may include one or more energy storage devices (e.g., batteries or capacitors) and/or circuitry for coupling components of the computing device 1400 to an energy source separate from the computing device 1400 (e.g., AC line power) .
The computing device 1400 may include a display device 1406 (or corresponding interface circuitry, as discussed above) . The display device 1406 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD) , a light-emitting diode display, or a flat panel display, for example.
The computing device 1400 may include an audio output device 1408 (or corresponding interface circuitry, as discussed above) . The audio output device 1408 may include any device that generates an audible indicator, such as speakers, headsets, or earbuds, for example.
The computing device 1400 may include an audio input device 1418 (or corresponding interface circuitry, as discussed above) . The audio input device 1418 may include any device that generates a signal representative of a sound, such as microphones, microphone arrays, or digital instruments (e.g., instruments having a musical instrument digital interface (MIDI) output) .
The computing device 1400 may include a GPS device 1416 (or corresponding interface circuitry, as discussed above) . The GPS device 1416 may be in communication with a satellite-based system and may receive a location of the computing device 1400, as known in the art.
The computing device 1400 may include another output device 1410 (or corresponding interface circuitry, as discussed above) . Examples of the other output device 1410 may include an audio codec, a video codec, a printer, a wired or wireless transmitter for providing information to other devices, or an additional storage device.
The computing device 1400 may include another input device 1420 (or corresponding interface circuitry, as discussed above) . Examples of the other input device 1420 may include an accelerometer, a gyroscope, a compass, an image capture device, a keyboard, a cursor control device such as a mouse, a stylus, a touchpad, a bar code reader, a Quick Response (QR) code reader, any sensor, or a radio frequency identification (RFID) reader.
The computing device 1400 may have any desired form factor, such as a handheld or mobile computer system (e.g., a cell phone, a smart phone, a mobile internet device, a music player, a tablet computer, a laptop computer, a netbook computer, an ultrabook computer, a personal digital assistant (PDA) , an ultramobile personal computer, etc. ) , a desktop  computer system, a server or other networked computing component, a printer, a scanner, a monitor, a set-top box, an entertainment control unit, a vehicle control unit, a digital camera, a digital video recorder, or a wearable computer system. In some embodiments, the computing device 1400 may be any other electronic device that processes data.
Select Examples
The following paragraphs provide various examples of the embodiments disclosed herein.
Example 1 provides a method, including receiving an activation tensor of a convolution, the activation tensor having one or more dimensions and a shape defined by the one or more dimensions; generating a reshaped activation tensor by modifying the shape of the activation tensor based on a structure of a MAC array; computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution; and generating an output tensor of the convolution by modifying a shape of the output tensor of the reshaped convolution.
Example 2 provides the method of example 1, in which modifying the shape of the activation tensor includes changing a total number of activations in a dimension of the activation tensor so that a total number of elements in a dimension of the output tensor of the reshaped convolution is a multiple of a total number of MAC units arranged in a row or column of the MAC array.
Example 3 provides the method of example 1 or 2, in which modifying the shape of the activation tensor includes adding one or more zero-valued activations into the activation tensor.
Example 4 provides the method of example 3, in which modifying the shape of the activation tensor further includes determining a spatial size of the activation tensor in a channel of the convolution; and determining whether the spatial size is a multiple of a total number of MAC units arranged in a row or column of the MAC array, in which the one or more zero-valued activations are added into the activation tensor in response to a determination that the spatial size is a multiple of a total number of MAC units arranged in a row or column of the MAC array.
Example 5 provides the method of any one of examples 1-4, in which a stride of the reshaped convolution is different from a stride of the convolution.
Example 6 provides the method of example 5, further including determining the stride of the reshaped convolution based on a shape of the activation tensor.
Example 7 provides the method of any one of examples 1-6, in which computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution includes storing activations of the activation tensor in a memory based on the shape of the activation tensor; and reading, by the MAC array, the activations from the memory based on a shape of the reshaped activation tensor.
Example 8 provides the method of any one of examples 1-7, in which computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution includes performing, by the MAC array, the reshaped convolution on the reshaped activation tensor and a weight tensor of the convolution.
Example 9 provides the method of any one of examples 1-8, in which the activation tensor includes a plurality of input channels, an input channel corresponds to a matrix including activations arranged in one or more rows and one or more columns, and modifying the shape of the activation tensor includes modifying a height or width of the matrix.
Example 10 provides the method of any one of examples 1-9, in which a total number of channels in the activation tensor is equal to the total number of channels in the reshaped activation tensor.
Example 11 provides one or more non-transitory computer-readable media storing instructions executable to perform operations, the operations including receiving an activation tensor of a convolution, the activation tensor having one or more dimensions and a shape defined by the one or more dimensions; generating a reshaped activation tensor by modifying the shape of the activation tensor based on a structure of a MAC array; computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution; and generating an output tensor of the convolution by modifying a shape of the output tensor of the reshaped convolution.
Example 12 provides the one or more non-transitory computer-readable media of example 11, in which modifying the shape of the activation tensor includes changing a total number of activations in a dimension of the activation tensor so that a total number of elements in a dimension of the output tensor of the reshaped convolution is a multiple of a total number of MAC units arranged in a row or column of the MAC array.
Example 13 provides the one or more non-transitory computer-readable media of example 11 or 12, in which modifying the shape of the activation tensor includes adding one or more zero-valued activations into the activation tensor.
Example 14 provides the one or more non-transitory computer-readable media of any one of examples 11-13, in which a stride of the reshaped convolution is different from a stride of the convolution.
Example 15 provides the one or more non-transitory computer-readable media of any one of examples 11-14, in which computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution includes storing activations of the activation tensor in a memory based on the shape of the activation tensor; and reading, by the MAC array, the activations from the memory based on a shape of the reshaped activation tensor.
Example 16 provides the one or more non-transitory computer-readable media of any one of examples 11-15, in which computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution includes performing, by the MAC array, the reshaped convolution on the reshaped activation tensor and a weight tensor of the convolution.
Example 17 provides the one or more non-transitory computer-readable media of any one of examples 11-16, in which the activation tensor includes a plurality of input channels, an input channel corresponds to a matrix including activations arranged in one or more rows and one or more columns, and modifying the shape of the activation tensor includes modifying a height or width of the matrix.
Example 18 provides an apparatus, including a computer processor for executing computer program instructions; and a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations including receiving an activation tensor of a convolution, the activation tensor having one or more dimensions and a shape defined by the one or more dimensions, generating a reshaped activation tensor by modifying the shape of the activation tensor based on a structure of a MAC array, computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution, and generating an output  tensor of the convolution by modifying a shape of the output tensor of the reshaped convolution.
Example 19 provides the apparatus of example 18, in which modifying the shape of the activation tensor includes changing a total number of activations in a dimension of the activation tensor so that a total number of elements in a dimension of the output tensor of the reshaped convolution is a multiple of a total number of MAC units arranged in a row or column of the MAC array.
Example 20 provides the apparatus of example 18 or 19, in which modifying the shape of the activation tensor includes adding one or more zero-valued activations into the activation tensor.
The above description of illustrated implementations of the disclosure, including what is described in the Abstract, is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. While specific implementations of, and examples for, the disclosure are described herein for illustrative purposes, various equivalent modifications are possible within the scope of the disclosure, as those skilled in the relevant art will recognize. These modifications may be made to the disclosure in light of the above detailed description.

Claims (20)

  1. A method, comprising:
    receiving an activation tensor of a convolution, the activation tensor having one or more dimensions and a shape defined by the one or more dimensions;
    generating a reshaped activation tensor by modifying the shape of the activation tensor based on a structure of a multiply-accumulate (MAC) array, the MAC array comprising one or more MAC units;
    computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution; and
    generating an output tensor of the convolution by modifying a shape of the output tensor of the reshaped convolution.
  2. The method of claim 1, wherein modifying the shape of the activation tensor comprises:
    changing a total number of activations in a dimension of the activation tensor so that a total number of elements in a dimension of the output tensor of the reshaped convolution is a multiple of a total number of MAC units arranged in a row or column of the MAC array.
  3. The method of claim 1, wherein modifying the shape of the activation tensor comprises:
    adding one or more zero-valued activations into the activation tensor.
  4. The method of claim 3, wherein modifying the shape of the activation tensor further comprises:
    determining a spatial size of the activation tensor in a channel of the convolution; and
    determining whether the spatial size is a multiple of a total number of MAC units arranged in a row or column of the MAC array,
    wherein the one or more zero-valued activations are added into the activation tensor in response to a determination that the spatial size is a multiple of a total number of MAC units arranged in a row or column of the MAC array.
  5. The method of claim 1, wherein a stride of the reshaped convolution is different from a stride of the convolution.
  6. The method of claim 5, further comprising:
    determining the stride of the reshaped convolution based on a shape of the activation tensor.
  7. The method of claim 1, wherein computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution comprises:
    storing activations of the activation tensor in a memory based on the shape of the activation tensor; and
    reading, by the MAC array, the activations from the memory based on a shape of the reshaped activation tensor.
  8. The method of claim 1, wherein computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution comprises:
    performing, by the MAC array, the reshaped convolution on the reshaped activation tensor and a weight tensor of the convolution.
  9. The method of claim 1, wherein the activation tensor comprises a plurality of input channels, an input channel corresponds to a matrix including activations arranged in one or more rows and one or more columns, and modifying the shape of the activation tensor comprises modifying a height or width of the matrix.
  10. The method of claim 1, wherein a total number of channels in the activation tensor is equal to the total number of channels in the reshaped activation tensor.
  11. One or more non-transitory computer-readable media storing instructions executable to perform operations, the operations comprising:
    receiving an activation tensor of a convolution, the activation tensor having one or more dimensions and a shape defined by the one or more dimensions;
    generating a reshaped activation tensor by modifying the shape of the activation tensor based on a structure of a multiply-accumulate (MAC) array, the MAC array comprising one or more MAC units;
    computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution; and
    generating an output tensor of the convolution by modifying a shape of the output tensor of the reshaped convolution.
  12. The one or more non-transitory computer-readable media of claim 11, wherein modifying the shape of the activation tensor comprises:
    changing a total number of activations in a dimension of the activation tensor so that a total number of elements in a dimension of the output tensor of the reshaped convolution is a multiple of a total number of MAC units arranged in a row or column of the MAC array.
  13. The one or more non-transitory computer-readable media of claim 11, wherein modifying the shape of the activation tensor comprises:
    adding one or more zero-valued activations into the activation tensor.
  14. The one or more non-transitory computer-readable media of claim 11, wherein a stride of the reshaped convolution is different from a stride of the convolution.
  15. The one or more non-transitory computer-readable media of claim 11, wherein computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution comprises:
    storing activations of the activation tensor in a memory based on the shape of the activation tensor; and
    reading, by the MAC array, the activations from the memory based on a shape of the reshaped activation tensor.
  16. The one or more non-transitory computer-readable media of claim 11, wherein computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution comprises:
    performing, by the MAC array, the reshaped convolution on the reshaped activation tensor and a weight tensor of the convolution.
  17. The one or more non-transitory computer-readable media of claim 11, wherein the activation tensor comprises a plurality of input channels, an input channel corresponds to a matrix including activations arranged in one or more rows and one or more columns, and modifying the shape of the activation tensor comprises modifying a height or width of the matrix.
  18. An apparatus, comprising:
    a computer processor for executing computer program instructions; and
    a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations comprising:
    receiving an activation tensor of a convolution, the activation tensor having one or more dimensions and a shape defined by the one or more dimensions,
    generating a reshaped activation tensor by modifying the shape of the activation tensor based on a structure of a multiply-accumulate (MAC) array, the MAC array comprising one or more MAC units,
    computing, by the MAC array using the reshaped activation tensor, an output tensor of a reshaped convolution, and
    generating an output tensor of the convolution by modifying a shape of the output tensor of the reshaped convolution.
  19. The apparatus of claim 18, wherein modifying the shape of the activation tensor comprises:
    changing a total number of activations in a dimension of the activation tensor so that a total number of elements in a dimension of the output tensor of the reshaped convolution is a multiple of a total number of MAC units arranged in a row or column of the MAC array.
  20. The apparatus of claim 18, wherein modifying the shape of the activation tensor comprises:
    adding one or more zero-valued activations into the activation tensor.
PCT/CN2024/081131 2024-03-12 2024-03-12 Reshaping convolution based on configuration of deep neural network accelerator Pending WO2025189339A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/CN2024/081131 WO2025189339A1 (en) 2024-03-12 2024-03-12 Reshaping convolution based on configuration of deep neural network accelerator

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2024/081131 WO2025189339A1 (en) 2024-03-12 2024-03-12 Reshaping convolution based on configuration of deep neural network accelerator

Publications (1)

Publication Number Publication Date
WO2025189339A1 true WO2025189339A1 (en) 2025-09-18

Family

ID=97062724

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/081131 Pending WO2025189339A1 (en) 2024-03-12 2024-03-12 Reshaping convolution based on configuration of deep neural network accelerator

Country Status (1)

Country Link
WO (1) WO2025189339A1 (en)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP3674982A1 (en) * 2018-12-27 2020-07-01 IMEC vzw Hardware accelerator architecture for convolutional neural network
CN115485695A (en) * 2020-08-21 2022-12-16 墨子国际有限公司 Method and system for hierarchically weighted sparse convolution processing
CN116547643A (en) * 2020-11-06 2023-08-04 墨芯国际有限公司 Method and system for convolution with active sparsity for workload balancing
US20240013040A1 (en) * 2023-09-26 2024-01-11 Intel Corporation Output drain path facilitating flexible schedule-based deep neural network accelerator
US20240046078A1 (en) * 2022-08-04 2024-02-08 Qualcomm Incorporated Desparsified convolution for sparse activations

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP3674982A1 (en) * 2018-12-27 2020-07-01 IMEC vzw Hardware accelerator architecture for convolutional neural network
CN115485695A (en) * 2020-08-21 2022-12-16 墨子国际有限公司 Method and system for hierarchically weighted sparse convolution processing
CN116547643A (en) * 2020-11-06 2023-08-04 墨芯国际有限公司 Method and system for convolution with active sparsity for workload balancing
US20240046078A1 (en) * 2022-08-04 2024-02-08 Qualcomm Incorporated Desparsified convolution for sparse activations
US20240013040A1 (en) * 2023-09-26 2024-01-11 Intel Corporation Output drain path facilitating flexible schedule-based deep neural network accelerator

Similar Documents

Publication Publication Date Title
US20240119269A1 (en) Dynamic sparsity-based acceleration of neural networks
US20240028895A1 (en) Switchable one-sided sparsity acceleration
US20230394312A1 (en) Pruning activations and weights of neural networks with programmable thresholds
US20230376765A1 (en) Performing operation in neural network with storage pointer and sparsity map
US20230229917A1 (en) Hybrid multipy-accumulation operation with compressed weights
US20230368030A1 (en) Block-wise pruning of weights in deep neural network
US20220051103A1 (en) System and method for compressing convolutional neural networks
US12572800B2 (en) Transposing memory layout of weights in deep neural networks (DNNs)
WO2025071788A1 (en) Output drain path facilitating flexible schedule-based deep neural network accelerator
WO2025091335A1 (en) Multi-precision tensor multiplication in neural network
WO2025096102A1 (en) Approximating activation functions in neural networks with programmable look-up table
WO2025122274A1 (en) Accuracy-based approximation of activation functions with programmable look-up table having area budget
WO2025136548A1 (en) Approximating activation function in neural network with look-up table having hybrid architecture
US20240020517A1 (en) Real-time inference of temporal down-sampling convolutional networks
WO2025174353A1 (en) Executing fourier transform operations with deep neural network accelerator
WO2026090973A1 (en) Converting neural network operation based on hardware configuration
WO2025184850A1 (en) Executing matrix multiplication by performing convolution with deep neural network accelerator
WO2026044571A1 (en) In-place execution of neural network operations with scatter write and gather read across memory fragments
US20240265260A1 (en) Compressing neural networks through unbiased minimum variance pruning
WO2025251247A1 (en) Converting interpolation operation in neural network to depthwise convolution
EP4738197A1 (en) Neural network accelerator performing operation with mixed-format weights
WO2025025421A1 (en) Tensor multiplication in neural network based on dequantization with shuffled data layout
WO2025207091A1 (en) Deep neural network accelerator with multifunctional data processing unit
WO2025230563A1 (en) Activation function approximation based on input range partition
WO2025230560A1 (en) Neural network accelerator with sparsity logic supporting various sparsity patterns and data precisions

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24928824

Country of ref document: EP

Kind code of ref document: A1