WO2022227671A1 - 处理器微架构、SoC芯片及低功耗智能设备 - Google Patents

处理器微架构、SoC芯片及低功耗智能设备 Download PDF

Info

Publication number
WO2022227671A1
WO2022227671A1 PCT/CN2021/142830 CN2021142830W WO2022227671A1 WO 2022227671 A1 WO2022227671 A1 WO 2022227671A1 CN 2021142830 W CN2021142830 W CN 2021142830W WO 2022227671 A1 WO2022227671 A1 WO 2022227671A1
Authority
WO
WIPO (PCT)
Prior art keywords
instruction
processor
architecture
processing
main
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2021/142830
Other languages
English (en)
French (fr)
Inventor
张慧敏
饶国明
杨鹤
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Spreadtrum Communications Shanghai Co Ltd
Original Assignee
Spreadtrum Communications Shanghai Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Spreadtrum Communications Shanghai Co Ltd filed Critical Spreadtrum Communications Shanghai Co Ltd
Priority to US18/288,627 priority Critical patent/US12510954B2/en
Publication of WO2022227671A1 publication Critical patent/WO2022227671A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F1/00Details not covered by groups G06F3/00 - G06F13/00 and G06F21/00
    • G06F1/26Power supply means, e.g. regulation thereof
    • G06F1/32Means for saving power
    • G06F1/3203Power management, i.e. event-based initiation of a power-saving mode
    • G06F1/3234Power saving characterised by the action undertaken
    • G06F1/3293Power saving characterised by the action undertaken by switching to a less power-consuming processor, e.g. sub-CPU
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F15/00Digital computers in general; Data processing equipment in general
    • G06F15/76Architectures of general purpose stored program computers
    • G06F15/78Architectures of general purpose stored program computers comprising a single central processing unit
    • G06F15/7807System on chip, i.e. computer system on a single chip; System in package, i.e. computer system on one or more chips in a single package
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F15/00Digital computers in general; Data processing equipment in general
    • G06F15/16Combinations of two or more digital computers each having at least an arithmetic unit, a program unit and a register, e.g. for a simultaneous processing of several programs
    • G06F15/163Interprocessor communication
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00Arrangements for program control, e.g. control units
    • G06F9/06Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/30Arrangements for executing machine instructions, e.g. instruction decode
    • G06F9/38Concurrent instruction execution, e.g. pipeline or look ahead
    • G06F9/3877Concurrent instruction execution, e.g. pipeline or look ahead using a secondary processor, e.g. coprocessor
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02DCLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
    • Y02D10/00Energy efficient computing, e.g. low power processors, power management or thermal management

Definitions

  • the invention relates to the technical field of intelligent terminals, in particular to a processor micro-architecture, a SoC (system-on-chip) chip and a low-power consumption intelligent device.
  • APCPU application processor main control system
  • Sensor Hub sub-control subsystem
  • Modem modem subsystem
  • BTCPU subsystem a subsystem for Bluetooth connection and control, etc.
  • the Cortex-M (a processor micro-architecture) series micro-architecture licensing scheme is generally implemented as the MCU core, but this method will bring several disadvantages: (1) Insufficient flexibility of the MCU micro-architecture: You can only choose from the optional MCU architecture, which often leads to excess or insufficient performance. For example, for the APCPU subsystem, a virtual memory system needs to be used to expand the available address space; while the ARM MCU does not have an MMU (memory management unit), so it cannot meet the requirements.
  • MMU memory management unit
  • the device core architecture cannot achieve the optimal solution: for example, in the smart watch architecture, APCPU and MMCPU (multimedia subsystem CPU) may all require AI processing acceleration capabilities.
  • APCPU and MMCPU multimedia subsystem CPU
  • a DSP co-processing unit SIMD
  • SIMD DSP co-processing unit
  • the main processor implements basic pipeline components and L1 Cache units, co-processing
  • the technical problem to be solved by the present invention is to overcome the fact that the processor micro-architecture in the prior art either has very low execution efficiency and cannot meet the requirements of high performance, or cannot realize the shared use of coprocessors, and the main processor cannot meet the actual business requirements. Flexible configuration is required, which is easy to cause the defect that the best performance of the micro-architecture cannot be achieved or the excess performance of the micro-architecture occurs.
  • a processor micro-architecture, a SoC chip and a low-power smart device are provided.
  • the present invention provides a processor micro-architecture, the processor micro-architecture includes a coprocessor and at least two main processors, each of the main processors is connected to the coprocessor through a request processing unit;
  • the request processing unit is configured to, when receiving use requests initiated by at least two main processors, determine a processing order corresponding to each main processor that initiates the request according to the first preset condition, and generate and Send feedback commands to different main processors;
  • the main processor is configured to send a processing instruction to the coprocessor for processing when the received feedback instruction indicates that the use is permitted.
  • the coprocessor includes an access interface unit and a coprocessing unit;
  • the main processor is configured to send the processing instruction to the access interface unit
  • the access interface unit is configured to transmit the processing instruction to the co-processing unit for processing
  • the request processing unit is independently arranged between the access interface unit and each of the main processors; or, the request processing unit is integrated and arranged in the access interface unit.
  • the first preset condition includes a preset processing priority corresponding to each of the main processors.
  • the request processing unit is configured to generate, according to the processing sequence, a first feedback instruction indicating permission to use and a second feedback instruction indicating continued waiting, and send the first feedback instruction to the highest ranked feedback instruction.
  • the main processor and respectively send the second feedback instruction to the other main processors in the lower order.
  • the request processing unit is further configured to send the first feedback instruction to the main processor in the next rank when the main processor being processed cancels the initiation of the use request, and send all the requests.
  • the second feedback instructions are respectively sent to the other main processors that are later in order.
  • each of the main processors corresponds to a different power domain
  • the access interface unit and the co-processing unit are divided into the same power domain.
  • the main processor is configured to send the processing instruction to the access interface unit through a command stream; and/or,
  • the co-processing unit is configured to determine to use a blocking instruction processing mode to process the processing instruction when the processing instruction satisfies the second preset condition; otherwise, use a pipelined instruction processing manner to process the processing instruction .
  • the co-processing unit is configured to send an access request of a data storage unit to the corresponding main processor according to the processing instruction and read target data, so as to perform an instruction processing operation according to the target data; the The co-processing unit is further configured to write back the calculation result of the instruction processing operation according to the target data to the original register corresponding to the processing instruction, and after receiving the instruction response from the main processor, write the original register to the original register.
  • the calculation result stored in the host processor is written back to the register of the host processor.
  • the coprocessor supports user-defined instructions and/or Vector vector instructions.
  • the main processor includes a plurality of configurable functional architectures, each of which is configured based on an open source instruction set architecture.
  • the open source instruction set architecture includes an open source instruction set architecture RISC-V based on the principle of reduced instruction set;
  • the open source instruction set architecture RISC-V supports multiple instruction sets
  • the corresponding computing unit and pipeline architecture are configured according to each of the instruction sets.
  • the instruction set includes a basic instruction set, a floating-point instruction set, a compressed instruction set or an extended instruction set; and/or,
  • the configured pipeline architecture supports a three-stage pipeline architecture or a five-stage pipeline architecture.
  • the functional architecture includes a multi-level memory structure.
  • each level of memory structure in the configured multi-level memory structure corresponds to multiple memories of different categories;
  • the memory includes L1 Cache, I-Cache, D-Cache, I-TCM, D-TCM (L1 Cache, I-Cache, D-Cache, I-TCM, D-TCM are all a kind of memory) or MMU.
  • the main processor further comprises an extended instruction interface and an extended instruction co-processing unit, the extended instruction co-processing unit is respectively connected to the extended instruction interface and the system bus in communication, and the extended instruction co-processing unit is used for The extended instruction interface and the system bus extend the instruction set.
  • the present invention also provides an SoC chip, which includes the above-mentioned processor micro-architecture.
  • the present invention also provides a low-power consumption smart device, and the low-power consumption smart device includes the above-mentioned SoC chip.
  • the low-power smart device includes a smart watch.
  • Each main processor is connected to the request processing unit (ie, the arbiter/arbitration unit), so that when multiple main processors simultaneously initiate a use request to use the coprocessor to the request processing unit, according to the pre-priority setting , only one main controller CPU and one Ready signal are returned, and a Hold signal is returned to the remaining main controller CPUs respectively; when the main controller CPU being processed cancels the use request, it returns to the next main controller CPU.
  • the request processing unit ie, the arbiter/arbitration unit
  • each main controller CPU sends out a usage request, and the arbitration unit realizes the mutual exclusive access at the instruction level, so as to realize the coordination
  • the processor is used as a shared resource, and multiple main controllers share access to the same coprocessor to meet the instruction-level delay, achieve the best balance between performance and power consumption, and effectively improve resource utilization.
  • the coprocessor is implemented by user-defined instructions and Vector vector instructions, which can realize data-level parallel processing; instruction fetching and decoding units are implemented by a unified CPU pipeline architecture; instruction execution, data read and write back are implemented by The coprocessor is completed, and the instruction transmission and data access intercommunication are realized with the main controller CPU through the dedicated interface.
  • the open source CPU project represented by RISC-V allows users to add and design instructions according to business needs to provide the best task processing capability.
  • Each subsystem CPU (ie, multiple main controllers) of the shared coprocessor unit adopts the Harvard structure and needs to implement independent L1 instruction Cache and L1 data Cache; the entire multi-core processor architecture implements a unified L2 Cache; , the coprocessor accesses the L1 D-cache of the main controller through a dedicated interface.
  • each main controller is divided into independent power supply domains, and the coprocessor also adopts an independent power supply domain design; when the coprocessor is not needed, the power supply can be cut off explicitly to improve the sharing efficiency,
  • the purpose of reducing power consumption can better meet the requirements of the smart wearable device for the performance-to-power ratio.
  • FIG. 1 is a first structural schematic diagram of a processor micro-architecture according to Embodiment 1 of the present invention.
  • FIG. 2 is a second structural schematic diagram of the processor micro-architecture according to Embodiment 1 of the present invention.
  • FIG. 3 is a third structural schematic diagram of the processor micro-architecture according to Embodiment 1 of the present invention.
  • FIG. 4 is a schematic diagram of a principle framework of a processor micro-architecture according to Embodiment 1 of the present invention.
  • FIG. 5 is a fourth schematic structural diagram of the processor micro-architecture according to Embodiment 1 of the present invention.
  • FIG. 6 is a schematic structural diagram of a main processor in a processor micro-architecture according to Embodiment 2 of the present invention.
  • FIG. 7 is a schematic diagram of a framework of an extended vector instruction and a user-defined instruction kernel architecture according to Embodiment 2 of the present invention.
  • the processor micro-architecture of this embodiment is applied in an SoC chip of a low-power smart device (eg, a smart watch), and the processor micro-architecture is the CPU micro-architecture of the low-power smart device.
  • a low-power smart device eg, a smart watch
  • the processor micro-architecture is the CPU micro-architecture of the low-power smart device.
  • multiple main processors will be designed to carry different system functions, such as APCPU for application processing, MMCPU for multimedia and Camera control, On SensorHub's SPCPU, etc.
  • the processor micro-architecture of this embodiment includes a co-processor (Co-Processor Unit) 100 and at least two main processors (CPUs) 200, each main processor 200 processes requests through The unit 300 (or called Arbitrator, arbiter/arbitration unit) is connected to the coprocessor 100 .
  • the request processing unit is provided between each main processor and the coprocessor, or in the coprocessor.
  • each main processor 200 is connected to the request processing unit 300 through a Req (request) line and a Response (response) line, and the numbers of the Req and Response lines can be designed or adjusted according to actual conditions.
  • Each CPU has CPU Core (processor core) and L1D-Cache (a kind of memory), etc.
  • the CPU sends usage requests (req) and receives feedback commands (resp, including Ready/Hold), sends processing commands (cmd) and receives command responses (cmd resp) through the CPU Core; receives Mem req (data request) and Send Mem resp (data response).
  • the request processing unit 300 is configured to, when receiving the usage requests initiated by at least two main processors 200, determine the processing order corresponding to each main processor 200 that initiates the request according to the first preset condition, and according to the processing order generating and sending feedback instructions to different main processors 200;
  • the first preset condition includes, but is not limited to, a preset processing priority corresponding to each main processor 200 .
  • the main processor 200 is configured to send the processing instruction to the coprocessor 100 for processing when the received feedback instruction indicates that the usage is allowed.
  • each subsystem CPU ie the main processor 200
  • the arbitration unit realizes the mutual exclusive access at the instruction level, and finally realizes the shared access of multiple subsystem CPUs to the same one
  • the coprocessor 100 unit, the coprocessor 100 is used as a shared resource, which effectively improves the resource utilization rate.
  • the coprocessor 100 supports user-defined extension instructions and Vector vector processing instructions. Taking the RISC-V instruction set as an example, it specifies user-defined instructions and Vector vector instructions that can be used for user extension.
  • the coprocessor 100 architecture is designed and implemented according to the specifications of these instruction sets, and can be used to process custom instructions and vector multiplication, vector addition, and the like.
  • the coprocessor 100 of this embodiment includes an access interface unit 400 and a coprocessing unit 500 .
  • the co-processing unit 500 is an Accelarator (accelerator), and specifically includes a Commad dispatch Unit (instruction distribution unit), a Data access unit (data storage unit), and the like.
  • the main processor 200 is configured to send the processing instruction to the access interface unit 400;
  • the access interface unit 400 is configured to transmit the processing instruction to the co-processing unit 500 for processing;
  • the request processing unit 300 is independently provided between the access interface unit 400 and each main processor 200 .
  • the request processing unit 300 is integrated in the access interface unit 400 .
  • the request processing unit 300 is integrated into the access interface unit 400 .
  • the request processing unit 300 is configured to generate, according to the processing sequence, a first feedback instruction (eg, Ready signal) representing permission to use, and a second feedback instruction (eg, Hold signal) representing continued waiting, and send the first feedback instruction to the sequencer
  • a first feedback instruction eg, Ready signal
  • a second feedback instruction eg, Hold signal
  • the request processing unit 300 is further configured to send the first feedback instruction to the main processor 200 in the next order when the main processor 200 being processed cancels the initiation of the use request, and send the second feedback instruction to other ordering partners respectively. After the main processor 200.
  • the main processor 200 is configured to send the processing instruction to the access interface unit 400 in a command stream manner.
  • the co-processing unit 500 is configured to determine to use a blocking instruction processing method to process the processing instruction when the processing instruction satisfies the second preset condition; otherwise, use a pipelined instruction processing manner to process the instruction. to be processed.
  • the second preset condition corresponds to the category to which the processing instruction belongs.
  • a certain category of processing instructions can be set according to actual needs to use a blocking instruction processing method, and a certain category of processing instructions requires a pipelined instruction processing method.
  • the coprocessor 100 When the blocking instruction processing method is adopted, after each instruction issued by the main controller CPU, the coprocessor 100 will return a Busy signal to notify the corresponding main controller CPU through the access interface unit 400Interface. The main controller CPU executes an instruction Only after completion can the remaining commands be sent.
  • the main controller CPU can continue to send instructions according to the flow line method without waiting.
  • the co-processing unit 500 is used to send the access request of the data storage unit to the corresponding main processor 200 according to the processing instruction and read the target data, so as to perform the instruction processing operation according to the target data;
  • the calculation result of the instruction processing operation is written back to the original register corresponding to the processing instruction, and after receiving the instruction response from the host processor 200 , the calculation result stored in the original register is written back to the register of the host processor 200 .
  • the instructions to be processed by the coprocessor 100 include user-defined instructions or vector instructions.
  • the instruction will simultaneously transmit two source register values to the coprocessor 100 through the dedicated channel.
  • the coprocessor 100 subsequently requests the main processor 200 to access the L1-Dcache data through the mem req request, reads the Mem data pointed to by the register, and transfers the instruction to access the Cache through mem_req and mem_resp.
  • the coprocessor 100 accesses the L1D-cache of the main controller through a dedicated interface.
  • each main processor 200 corresponds to a different power domain, so that each main processor 200 can be powered on and off independently; the access interface unit 400 and the co-processing unit 500 are divided into the same power domain. If the main processor 200 needs to use the coprocessor 100, the power switch of the coprocessor 100 needs to be turned on in advance.
  • each main controller is divided into independent power supply domains, and the coprocessor 100 also adopts an independent power supply domain design; when the coprocessor 100 is not needed, the power supply can be cut off explicitly, so as to improve the sharing efficiency,
  • the purpose of reducing power consumption can better meet the requirements of the smart wearable device for the performance-to-power ratio.
  • the above-mentioned main processor 200 has the requirement of performing AI processing, and the vector coprocessor 100 can provide AI computing capability conforming to the performance-to-power ratio.
  • Each main processor 200 generates an execution instruction through the program processor PC and sends it to the instruction Cache (instruction memory), and then enters the instruction distribution queue. heap), ALU execute to store the execution structure in the Data Memory data storage unit; if it belongs to the instruction that needs to be processed by the coprocessor 100, then the request processing unit 300 that is included in the access interface unit 400 (interface) is the arbiter initiated uniformly A usage request to use the coprocessor 100;
  • main processors 200 CPU1, CPU2, . ;
  • the coprocessor 100 generates a processing order according to the preset request and response priorities of the n main processors 200, first returns a Ready signal to the main processor 200 with the highest ranking, and sorts the other n-1
  • the rear main processor 200 returns the Hold signal;
  • the main processor 200 receiving the Ready signal transmits the processing instruction to the co-processing unit 500 of the co-processor 100 through the access interface unit 400 in a command stream mode for processing; wherein the processing instruction may adopt a blocking processing mode or a pipeline processing mode Way;
  • the processing instruction passes through the coprocessor 100 access unit and the instruction distribution unit of the coprocessor 100 in turn.
  • the instruction distribution unit determines whether the current processing instruction belongs to a vector processing instruction or a user-defined instruction according to the set conditions, and allocates it to the corresponding instruction after determination. the instruction unit for processing;
  • the vector processor unit analyze the input processing instruction to write back the data request instruction and sequentially pass through the coprocessor 100 access unit and the access interface unit 400 to transmit and access the data storage unit Data Memory of the main processor 200 and read the corresponding target. Data, and then complete the filling of the vector registers through the access interface unit 400 and the coprocessor 100 access unit to ensure that the vector processor unit obtains the instruction processing result based on the addition pipeline, vector register file, multiplication pipeline, etc., and finally passes the coprocessor 100.
  • the access unit and the access interface unit 400 store the instruction processing result in the data storage unit Data Memory of the main processor 200.
  • the instruction processing principle of the user-defined co-processing unit 500 is similar to the instruction processing principle of the above-mentioned vector processor unit, so it will not be repeated here.
  • the instructions to be processed by the coprocessor 100 may be user-defined instructions or vector instructions.
  • the instruction When the user-defined instruction contains register values, the instruction will simultaneously transmit two source register values to the coprocessor 100 through the dedicated channel, and the coprocessor 100 subsequently requests access to the main processor 200 through the mem req request.
  • L1-Dcache data read the Mem data pointed to by the register, and pass the instructions to access the Cache through mem_req and mem_resp.
  • the co-processing unit 500 sends the access request of the data storage unit to the corresponding main processor 200 according to the processing instruction and reads the target data, so as to perform the instruction processing operation according to the target data; write back the calculation result of the instruction processing operation according to the target data.
  • the original register corresponding to the instruction is processed, and after receiving the instruction response from the main processor 200, the calculation result stored in the original register is written back to the register of the main processor 200.
  • the request processing unit 300 If the currently processing main processor 200 cancels the initiation of the use request, the request processing unit 300 returns a Ready signal to the main processor 200 in the next order, and sends a Ready signal to the other n-2 main processors in the lower order. The processor 200 returns the Hold signal; and so on, until the Ready signal feedback of all the main processors 200 is completed.
  • the multiple main controllers sharing the coprocessor 100 adopt the Harvard structure, and need to implement independent L1 instruction cache and L1 data cache.
  • the processor architecture implements a unified L2Cache.
  • Each main controller uses the Extension Interface (extension interface) to connect with the coprocessor.
  • each main processor is connected to the request processing unit, so that when multiple main processors simultaneously initiate a use request to use the coprocessor to the request processing unit, only one main control unit is returned according to the pre-priority setting.
  • a Ready signal is sent to the controller CPU, and a Hold signal is returned to the remaining main controller CPUs respectively; when the processing main controller CPU cancels the use request, a Ready signal is returned to the next main controller CPU, and a Ready signal is sent to the remaining main controller CPUs at the same time.
  • the main controller CPUs return a Hold signal respectively, and so on, that is, they are divided according to the requirements of the usage scenarios.
  • Each main controller CPU sends out a usage request, and the arbitration unit realizes the mutual exclusive access at the instruction level, so as to realize the coprocessor as a shared resource.
  • Multiple main controllers share access to the same coprocessor to meet the instruction-level delay, achieve the best balance between performance and power consumption, and effectively improve resource utilization.
  • the processor micro-architecture of this embodiment is a further improvement to Embodiment 1, specifically:
  • each main processor 200 in this embodiment includes a plurality of configurable functional architectures, and the functional architecture includes a pipeline architecture 1, an extended instruction interface 2, and a memory set in the main Core (processor core). Architecture 3 and TEE4, and the extended instruction co-processing unit 5005 located outside the main Core.
  • Each functional architecture is configured based on the open source instruction set architecture.
  • the open source instruction set architecture includes the open source instruction set architecture RISC-V based on the principle of reduced instruction set.
  • the open source instruction set architecture RISC-V supports multiple instruction sets; the corresponding computing unit and pipeline architecture are configured according to each instruction set.
  • the open source instruction set architecture RISC-V is used for targeted design according to the actual design requirements, and flexible combination of various functional architectures can be applied in low-power smart devices to meet the different design requirements of different customers and realize the customization of the processor micro-architecture. demand.
  • the instruction sets supported by the processor microstructure include basic instruction sets, floating-point instruction sets, compressed instruction sets, extended instruction sets, and the like.
  • the basic instruction set includes addition, subtraction, multiplication, division, atomic swap, memory access and other instructions
  • the floating-point instruction set includes single-precision and double-precision floating-point calculations
  • the compressed instruction set includes 16bit
  • the extended instruction set includes vector instructions, SIMD (Single Instruction Multiple Data Stream) instruction, etc.
  • ALUs computing units
  • the processor micro-architecture of this embodiment includes an extended instruction interface and an extended instruction co-processing unit 500.
  • the extended instruction co-processing unit 500 is respectively connected to the extended instruction interface and the system bus in communication, and the extended instruction co-processing unit 500 is used for extending the instruction based on the extended instruction.
  • the interface and system bus extend the instruction set.
  • the main core in the micro-architecture design of the CPU processor is used to support instruction sets such as the basic instruction set, the floating-point instruction set, and the compressed instruction set, while the implementation of the extended instruction set needs to rely on the extended instruction co-processing outside the main core.
  • unit 500 the vector instruction co-processing unit 500 is used to process extended vector instructions, so that a logic unit with better performance and power consumption ratio can be used to process the required domain computing requirements, that is, the extended instruction interface and the extended instruction co-processing unit 500 can be used.
  • the extended instruction interface and the extended instruction co-processing unit 500 can be used.
  • a special floating-point processing pipeline can be designed for the FPU (floating-point processing processor), adding a floating-point processing computing unit, which shares the instruction prefetching and decoding unit with the integer processing pipeline; of course, it can be configured according to the needs of the product. , choose to support or not support the floating-point processing instruction set, which ensures the flexibility of configuration and better meets higher design requirements.
  • FPU floating-point processing processor
  • the functional architecture of this embodiment includes, but is not limited to, a pipeline architecture, a memory architecture, and the like.
  • the pipeline architecture is a multi-level pipeline architecture
  • the configured multi-level pipeline architecture supports a three-level pipeline architecture or a five-level pipeline architecture.
  • other levels of pipeline architecture can be used, and the configuration can be adjusted according to the actual design requirements.
  • a multi-stage pipeline architecture is designed according to the different implementation complexity.
  • a configurable multi-stage pipeline architecture method is: use the classic 5-stage pipeline architecture in the application processor with high processing performance; when the processing performance requirements are low, the power consumption area requires the optimal MCU control
  • a simple 3-stage pipeline architecture is used; due to the basic principle of pipeline design, in the 3-stage pipeline, the highest supported main frequency is lower than the 5-stage pipeline, which sacrifices the processing capacity in exchange for chip area and power. consumption optimization.
  • this embodiment can form independent configurable features for integer pipelines, floating-point pipelines, 3-stage pipelines, 5-stage pipelines, etc., and can also be combined and designed according to system complexity.
  • the functional architecture includes a multi-level memory structure.
  • Each level of memory structure in the configured multi-level memory structure corresponds to multiple memories of different categories;
  • the memory includes L1 Cache, I-Cache, D-Cache, I-TCM, D-TCM or MMU.
  • the processor micro-architecture in this embodiment supports flexible configuration of memories such as L1 Cache, I-Cache, D-Cache, I-TCM, D-TCM, or MMU. Combinations are made according to different usage requirements. For example, in the AP application processor environment, I-Cache, D-Cache, L2 Cache and MMU virtual memory need to be combined and configured; in the Sensor Hub requirements with higher low power consumption requirements , you only need to configure the I-Cache and D-TCM in combination to meet the requirements.
  • the processor micro-architecture in this embodiment supports the MMU architecture.
  • the MMU is a memory management unit, located between the CPU core and the external main memory, and performs memory management by loading page tables, mainly realizing the transformation from virtual addresses to actual physical addresses.
  • the technology of virtual memory can be realized through MMU, which is very effective in expanding embedded systems with insufficient memory (such as smart watches).
  • the MPU microprocessor
  • the MPU can realize the access protection of the main memory space by different co-processing units 500 and different MCUs. As long as the main memory space is divided into different areas, and the read and write permissions of the MPU are configured, requests for unauthorized access will be effectively blocked.
  • processor microarchitecture supports a Trusted Execution Environment TEE.
  • the processor micro-architecture in this embodiment supports the implementation of TEE design, and is designed in a privileged mode.
  • a special instruction is set, the system enters the privileged mode.
  • a trusted operating system is executed in a hardware environment that is completely isolated from the normal mode, including independent registers, independent and isolated storage spaces, independent and isolated devices, and TOS (trusted operating system).
  • the Extension Interface in the processor core in Figure 1 represents the extension instruction interface
  • FIQ-CTL represents the fast post-interrupt request control module
  • IRQ-CTL represents the interrupt request control module
  • Debug represents the debug module
  • JTAG represents the debug interface
  • IRQ_src represents the corresponding interrupt module
  • Timer represents the timing module
  • Per1, Pern, Dev1 all represent the terminals connected to the bus
  • SRAM represents static random access memory
  • ExtMEM (a kind of register).
  • the processor micro-architecture in this embodiment can be designed through the above-mentioned instruction set design, pipeline design, register design, Cache design, extended instruction unit design, etc. Under the same conditions, the power consumption capability of the processor micro-architecture is bound to be better than the current one.
  • the power consumption capability of the processor micro-architecture is bound to be better than the current one.
  • the extended instructions of the redesigned processor micro-architecture in this embodiment only a smaller number of instructions can be used to achieve the same function; because the execution clock cycle is reduced, the corresponding power consumption is also optimized.
  • the combination and combination of unit modules, as well as user-defined extended instructions and other optimized designs effectively reduce the chip area and operating frequency, so as to achieve the purpose of power consumption optimization.
  • a CPU micro-architecture in which multiple main processor CPUs share the same co-processor unit can be designed according to customer requirements, and at the same time, based on the actual requirements, different functional structures (such as implementing L1 Cache) are implemented based on the open source instruction set architecture RISC-V , L2 Cache, MMU, TEE, floating-point operation unit, vector operation unit and other pipeline architectures, memory architectures, etc.) can be flexibly configured and combined, which can be used to achieve the MCU requirements of multi-functional subsystems in complex SoC systems, thus providing
  • the CPU micro-architecture that can be precisely configured, can adjust the functional characteristics, performance and power consumption optimal solution, meets the customizable requirements for the processor CPU micro-architecture, and then meets the product configuration requirements of the low-power smart watch system.
  • the SoC chip of this embodiment includes the processor micro-architecture in Embodiment 1 or 2.
  • the SoC chip of this embodiment includes the above-mentioned processor micro-architecture, which can be customized according to customer requirements.
  • the CPU micro-architecture in which multiple main processor CPUs share the same co-processor unit is designed, and the open-source instruction set architecture RISC-V is based on actual requirements.
  • Flexible configuration and combination of different functional structures can be used to meet the MCU requirements of multi-functional subsystems in complex SoC systems, thereby providing a CPU micro-architecture that can be precisely configured, adjustable in functional characteristics, performance and power consumption. , to meet the customizable requirements for the processor CPU micro-architecture, and then meet the product configuration requirements of the low-power smart watch system.
  • the low-power smart device in this embodiment includes an SoC chip.
  • the low-power smart devices include smart watches.
  • the low-power smart device of this embodiment includes the above-mentioned SoC chip, which can be customized according to customer requirements.
  • the CPU micro-architecture in which multiple main processor CPUs share the same co-processor unit is designed, and the RISC-based open source instruction set architecture is based on actual requirements.
  • V can flexibly configure and combine different functional structures, which can be used to meet the MCU requirements of multi-functional subsystems in complex SoC systems, thereby providing a CPU microcomputer that can be precisely configured, adjustable in functional characteristics, performance and power consumption.
  • the architecture can meet the customizable requirements for the processor CPU micro-architecture, and then meet the product configuration requirements of the low-power smart watch system.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Computer Hardware Design (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Microelectronics & Electronic Packaging (AREA)
  • Advance Control (AREA)

Abstract

本发明公开了一种处理器微架构、SoC芯片及低功耗智能设备,该处理器微架构包括协处理器和和至少两个主处理器,每个主处理器通过请求处理单元与协处理器连接;请求处理单元用于在接收到两个主处理器发起的使用请求时确定发起请求的主处理器的处理顺序,并根据处理顺序生成并发送反馈指令至不同的主处理器;主处理器用于在反馈指令表征允许使用时,则将处理指令发送至协处理器进行处理。本发明根据各主控制器CPU发出使用请求,由仲裁单元实现指令级互斥访问,从而实现多个主控制器共享访问同一个协处理器,满足指令级的延时,达到性能和功耗的最佳平衡,且实现对处理器CPU微架构的可定制化需求,满足低功耗智能系统的产品配置要求。

Description

处理器微架构、SoC芯片及低功耗智能设备
本申请要求申请日为2021年4月30日的中国专利申请CN202110485283.3的优先权。本申请引用上述中国专利申请的全文。
技术领域
本发明涉及智能终端技术领域,特别涉及一种处理器微架构、SoC(系统级芯片)芯片及低功耗智能设备。
背景技术
目前,在低功耗智能终端(如智能手表)的系统架构设计中,需要在芯片中集成多个MCU(微控制单元),以达到通过多个功能子模块共同实现复杂功能的目的。例如用于用户应用程序控制的APCPU(应用处理器主控系统)、用于低功耗控制及传感器处理的Sensor Hub(副控子系统)、用于无线蜂窝通信的Modem(调制解调器)子系统、用于蓝牙连接与控制的BTCPU子系统(一种子系统)等。
在传统的CPU(中央处理器)架构实现中,一般会以Cortex-M(一种处理器微架构)系列微架构授权方案作为MCU核实现,但这样做的方式会带来几个不利因素:(1)MCU微架构灵活性不足:只能在可选的MCU架构上进行选择,这样往往导致性能过剩或者性能不足。例如针对APCPU子系统来说,需要采用虚拟内存体系以扩展可用地址空间;而ARM MCU不具备MMU(内存管理单元),因而无法满足要求。针对BTCPU子系统,要求功耗最低,CodeSize(代码长度)最小,性能上并不追求更高,而ARM MCU往往会因多方面受限于,无法采用最合适的MCU核;(2)多处理器核架构无法达到最优解:例如在智能手表架构中,APCPU、MMCPU(多媒体子系统CPU)都可能需要AI处理加速能力。ARMMCU架构中,为每个core都提供了一个DSP协处理单元(SIMD),但无法实现DSP资源共享,从而造成资源浪费;另外,主处理器实现基本的流水线部件和L1 Cache等单元,协处理器基于需求进行单独设计并采用标准的总线接口;主处理器、协处理器、DDR(双倍速率)主存储器等都挂载在总线(如AXI等)上;这样架构设计的好处在于设计简单,能够快速扩充协处理器功能;但缺点是总线传输延时非常大,执行效率非常低,远非指令级的交互延时,无法达到高性能的要求;(3)无法根据客户需求扩展自定义指令,传统基于ARM的MCU只能处理通用指令,而无法根据客户需求自行扩展指令集。
发明内容
本发明要解决的技术问题是为了克服现有技术中处理器微架构要么存在执行效率非常低,无法达到高性能的要求,要么无法实现协处理器的共享使用,且主处理器无法根据实际业务需求进行灵活配置,易造成无法实现微架构最佳性能或者发生微架构性能过剩的缺陷,提供一种处理器微架构、SoC芯片及低功耗智能设备。
本发明是通过下述技术方案来解决上述技术问题:
本发明提供一种处理器微架构,所述处理器微架构包括协处理器和和至少两个主处理器,每个所述主处理器通过请求处理单元与所述协处理器连接;
所述请求处理单元用于在接收到至少两个主处理器发起的使用请求时,根据第一预设条件确定发起请求的每个主处理器对应的处理顺序,并根据所述处理顺序生成并发送反馈指令至不同的主处理器;
所述主处理器用于在接收的所述反馈指令表征允许使用时,则将处理指令发送至所述协处理器进行处理。
较佳地,所述协处理器包括访问接口单元和协处理单元;
所述主处理器用于将所述处理指令发送至所述访问接口单元;
所述访问接口单元用于将所述处理指令传送至所述协处理单元进行处理;
其中,所述请求处理单元独立设置在所述访问接口单元和每个所述主处理器之间;或,所述请求处理单元集成设置在所述访问接口单元中。
较佳地,所述第一预设条件包括预先设定的每个所述主处理器对应的处理优先级。
较佳地,所述请求处理单元用于根据所述处理顺序生成表征允许使用的第一反馈指令和表征继续等待的第二反馈指令,并将所述第一反馈指令发送至排序最靠前的所述主处理器,且将所述第二反馈指令分别发送至其他排序靠后的所述主处理器。
较佳地,所述请求处理单元还用于在正在处理的所述主处理器取消发起使用请求时,将所述第一反馈指令发送至排序在下一位的所述主处理器,且将所述第二反馈指令分别发送至其他排序靠后的所述主处理器。
较佳地,每个所述主处理器对应不同的电源域;和/或,
所述访问接口单元和所述协处理单元划分在同一电源域中。
较佳地,所述主处理器用于通过命令流方式将所述处理指令发送至所述访问接口单元;和/或,
所述协处理单元用于在所述处理指令满足第二预设条件时确定采用阻塞式指令处理方式对所述处理指令进行处理;否则,采用流水线式的指令处理方式对所述处理指令进 行处理。
较佳地,所述协处理单元用于根据所述处理指令向对应的所述主处理器发送数据存储单元的访问请求并读取目标数据,以根据所述目标数据进行指令处理操作;所述协处理单元还用于将根据所述目标数据进行指令处理操作的计算结果写回至所述处理指令对应的原始寄存器,并在接收到所述主处理器的指令响应后,将所述原始寄存器中存储的所述计算结果写回至所述主处理器的寄存器中。
较佳地,所述协处理器支持用户自定义指令和/或Vector向量指令。
较佳地,所述主处理器包括多个可配置的功能架构,每个所述功能架构基于开源指令集架构进行配置。
较佳地,所述开源指令集架构包括基于精简指令集原则的开源指令集架构RISC-V;
其中,所述开源指令集架构RISC-V支持多种指令集;
根据每种所述指令集配置对应的计算单元和流水线架构。
较佳地,所述指令集包括基础指令集、浮点指令集、压缩指令集或扩展指令集;和/或,
配置后的所述流水线架构支持三级流水线架构或五级流水线架构。
较佳地,所述功能架构包括多级存储器结构。
较佳地,配置后的所述多级存储器结构中的每级存储器架构对应不同类别的多个存储器;
其中,所述存储器包括L1 Cache、I-Cache、D-Cache、I-TCM、D-TCM(L1 Cache、I-Cache、D-Cache、I-TCM、D-TCM均为一种存储器)或MMU。
较佳地,所述主处理器还包括扩展指令接口和扩展指令协处理单元,所述扩展指令协处理单元分别与扩展指令接口和系统总线通信连接,所述扩展指令协处理单元用于基于所述扩展指令接口和所述系统总线对所述指令集进行扩展。
本发明还提供一种SoC芯片,所述SoC芯片包括上述的处理器微架构。
本发明还提供一种低功耗智能设备,所述低功耗智能设备包括上述的SoC芯片。
较佳地,所述低功耗智能设备包括智能手表。
在符合本领域常识的基础上,所述各优选条件,可任意组合,即得本发明各较佳实施例。
本发明的积极进步效果在于:
(1)通过每个主处理器与请求处理单元(即仲裁器/仲裁单元)连接,使得在多个主处理器同时向请求处理单元发起使用协处理器的使用请求时,根据预先优先级设置,只 返回一个主控制器CPU一个Ready信号,向其余的主控制器CPU分别返回一个Hold信号;当正在处理的主控制器CPU取消使用请求时则向排序在下一位的主控制器CPU返回一个Ready信号,同时向其余的主控制器CPU分别返回一个Hold信号,依次类推,即根据使用场景需求进行划分,各主控制器CPU发出使用请求,由仲裁单元实现指令级互斥访问,从而实现协处理器作为共享资源,多个主控制器共享访问同一个协处理器,满足指令级的延时,达到性能和功耗的最佳平衡,有效地提高了资源利用率。
(2)协处理器由用户自定义指令和Vector向量指令实现,可实现数据级并行处理;指令取指和译码单元由统一的CPU流水线架构所实现;指令执行、数据读取和写回由协处理器完成,并通过专用接口与主控制器CPU实现指令传输和数据访问互通。该CPU微架构设计中,以RISC-V为代表的开源CPU项目,允许用户根据业务需求自行增加并设计指令,以提供最佳的任务处理能力。
(3)共享协处理器单元的各子系统CPU(即多个主控制器),采用哈佛结构,需要实现独立的L1指令Cache和L1数据Cache;整个多核处理器架构实现统一的L2 Cache;另外,协处理器通过专用接口访问主控制器的L1 D-cache。
(4)用低功耗设计,各个主控制器划分为独立的供电域,协处理器也采用独立的供电域设计;在无需使用协处理器时可显性切断电源,以达到提升共享效率、降低功耗的目的,能够更好地满足智能穿戴设备对性能功耗比的要求。
(5)根据实际需求基于开源指令集架构RISC-V对不同功能结构(如等流水线架构、存储器架构等)进行灵活配置,搭配组合,可用于实现复杂的SoC系统中多功能子系统的MCU需求,从而提供了可精确配置、可调整功能特性、性能与功耗最优解的CPU微架构,满足对处理器CPU微架构的可定制化需求,进而满足低功耗智能手表系统的产品配置要求。
附图说明
图1为本发明实施例1的处理器微架构的第一结构示意图。
图2为本发明实施例1的处理器微架构的第二结构示意图。
图3为本发明实施例1的处理器微架构的第三结构示意图。
图4为本发明实施例1的处理器微架构的原理框架示意图。
图5为本发明实施例1的处理器微架构的第四结构示意图。
图6为本发明实施例2的处理器微架构中主处理器的架构示意图。
图7为本发明实施例2的扩展向量指令和用户自定义指令内核架构的框架示意图。
具体实施方式
下面通过实施例的方式进一步说明本发明,但并不因此将本发明限制在所述的实施例范围之中。
实施例1
本实施例的处理器微架构应用在低功耗智能设备(如智能手表)的SoC芯片中,该处理器微架构即为低功耗智能设备的CPU微架构。以智能手表为代表的可穿戴芯片设计中,基于不同的功能需求,会设计多个主处理器以承载不同的系统功能,例如用于应用处理的APCPU、用于多媒体和Camera控制的MMCPU,用于SensorHub的SPCPU等。
如图1和图2所示,本实施例的处理器微架构包括协处理器(Co-Processor Unit)100和和至少两个主处理器(CPU)200,每个主处理器200通过请求处理单元300(或称为Arbitrator,仲裁器/仲裁单元)与协处理器100连接。请求处理单元设置在每个主处理器和协处理器之间,或者设置在协处理器中。
具体地,每个主处理器200通过Req(请求)线和Response(响应)线连接到请求处理单元300,Req线和Response线的数量可以根据实际情况进行设计或者调整。
每个CPU设有CPU Core(处理器核)和L1D-Cache(一种存储器)等。CPU通过CPU Core发送使用请求(req)以及接收反馈指令(resp,包括Ready/Hold)、发送处理指令(cmd)以及接收指令响应(cmd resp);通过L1D-Cache接收Mem req(数据请求)以及发送Mem resp(数据响应)。具体地:请求处理单元300用于在接收到至少两个主处理器200发起的使用请求时,根据第一预设条件确定发起请求的每个主处理器200对应的处理顺序,并根据处理顺序生成并发送反馈指令至不同的主处理器200;
其中,第一预设条件包括但不限于预先设定的每个主处理器200对应的处理优先级。
主处理器200用于在接收的反馈指令表征允许使用时,则将处理指令发送至协处理器100进行处理。
根据使用场景需求进行划分,各子系统CPU(即主处理器200)同时发出使用协处理器100的使用请求,由仲裁单元实现指令级互斥访问,最终实现支持多个子系统CPU共享访问同一个协处理器100单元,协处理器100作为共享资源,有效地提高了资源利用率。
在一可实施的方案中,协处理器100支持自定义扩展指令和Vector向量处理指令,以RISC-V指令集为例,其规定了可用于用户扩展的自定义指令和Vector向量指令。根据这些指令集的规范设计实现协处理器100架构,可用于处理自定义指令和向量乘法, 向量加法等。
如图3所示,本实施例的协处理器100包括访问接口单元400和协处理单元500。协处理单元500即为Accelarator(加速器),具体包括Commad dispatch Unit(指令分配单元)以及Data access unit(数据存储单元)等。
主处理器200用于将处理指令发送至访问接口单元400;
访问接口单元400用于将处理指令传送至协处理单元500进行处理;
在一可实施例的方案中,请求处理单元300独立设置在访问接口单元400和每个主处理器200之间。在一可实施例的方案中,请求处理单元300集成设置在访问接口单元400中。优选地,将请求处理单元300集成设置访问接口单元400中。
具体地,请求处理单元300用于根据处理顺序生成表征允许使用的第一反馈指令(如Ready信号)和表征继续等待的第二反馈指令(如Hold信号),并将第一反馈指令发送至排序最靠前的主处理器200,且将第二反馈指令分别发送至其他排序靠后的主处理器200。
请求处理单元300还用于在正在处理的主处理器200取消发起使用请求时,将第一反馈指令发送至排序在下一位的主处理器200,且将第二反馈指令分别发送至其他排序靠后的主处理器200。
在一可实施例的方案中,主处理器200用于通过命令流方式将处理指令发送至访问接口单元400。
在一可实施例的方案中,协处理单元500用于在处理指令满足第二预设条件时确定采用阻塞式指令处理方式对处理指令进行处理;否则,采用流水线式的指令处理方式对处理指令进行处理。
其中,第二预设条件对应处理指令所属类别,具体可以根据实际需求设定某种类别的处理指令需要采用阻塞式指令处理方式,某种类别的处理指令需要采用流水线式的指令处理方式。
在采用阻塞式指令处理方式时,主控制器CPU每发一条指令后,协处理器100将返回Busy信号,以通过访问接口单元400Interface通知对应的主控制器CPU,主控制器CPU在一条指令执行完毕后才能继续发送剩余指令。
在采用流水线式指令处理方式时,主控制器CPU可以持续按照流程线方式发送指令,无需等待。
协处理单元500用于根据处理指令向对应的主处理器200发送数据存储单元的访问请求并读取目标数据,以根据目标数据进行指令处理操作;协处理单元500还用于将根据目标数据进行指令处理操作的计算结果写回至处理指令对应的原始寄存器,并在接收 到主处理器200的指令响应后,将原始寄存器中存储的计算结果写回至主处理器200的寄存器中。
协处理器100待处理的指令包括用户自定义指令或者向量指令。当用户自定义指令中包含有寄存器值时,该指令将通过专用通道同时传送2个源寄存器值到协处理器100中。协处理器100后继通过mem req请求向主处理器200请求访问L1-Dcache数据,读取寄存器指向的Mem数据,访问Cache的指令传递通过mem_req和mem_resp来进行。
另外,协处理器100通过专用接口访问主控制器的L1D-cache。另外,每个主处理器200对应不同的电源域,以实现每个主处理器200的单独上下电;访问接口单元400和协处理单元500划分在同一电源域中。若有主处理器200需要使用协处理器100时,则需要提前将协处理器100的电源开关打开。
即采用低功耗设计,各个主控制器划分为独立的供电域,协处理器100也采用独立的供电域设计;在无需使用协处理器100时可显性切断电源,以达到提升共享效率、降低功耗的目的,能够更好地满足智能穿戴设备对性能功耗比的要求。
另外,由于AI应用的扩展,以上主处理器200都有进行AI处理的需求,而向量协处理器100能够提供符合性能功耗比的AI计算能力。
下面结合图4具体说明本实施例的处理器微架构的工作原理:
(1)每个主处理器200通过程序处理器PC产生执行指令并发送至instruction Cache(指令存储器),然后进入指令分发队列,如属于主处理器200自身完成的指令则分发至register file(寄存器堆)、ALU执行以将执行结构存储至Data Memory数据存储单元中;若属于需要协处理器100处理的指令,则统一向访问接口单元400(interface)中包含的请求处理单元300即仲裁器发起使用协处理器100的使用请求;
(2)N个主处理器200中的n个主处理器200(CPUl、CPU2…、CPUn)同时发起使用协处理器100的使用请求;(N、n均为正数器且N≥n);
(3)协处理器100根据n个主处理器200预设的请求响应优先级生成处理顺序,首先给排序最靠前的一个主处理器200返回一个Ready信号,并向其他n-1个排序靠后的主处理器200返回Hold信号;
(4)接收Ready信号的主处理器200采用命令流方式将处理指令通过访问接口单元400传送至协处理器100的协处理单元500进行处理;其中处理指令可以采用阻塞式处理方式或流水线式处理方式;
处理指令依次经过协处理器100的协处理器100存取单元和指令分配单元,指令分配单元根据设定条件判断当前处理指令属于向量处理指令还是用户自定义指令,并在确 定后分别分配至对应的指令单元进行处理;
以向量处理器单元为例,分析输入的处理指令写回数据请求指令依次通过协处理器100存取单元和访问接口单元400传输访问主处理器200的数据存储单元Data Memory并读取对应的目标数据,进而通过访问接口单元400和协处理器100存取单元完成向量寄存器的填充,以保证向量处理器单元基于加法流水线、向量寄存器堆、乘法流水线等得到指令处理结果,最后通过协处理器100存取单元和访问接口单元400将该指令处理结果存储至主处理器200的数据存储单元Data Memory。
对于用户自定义协处理单元500的指令处理原理与上述向量处理器单元的指令处理原理类似,因此此处就不再赘述。
(5)协处理器100待处理的指令可能为用户自定义指令或者向量指令。当用户自定义指令中包含有寄存器值时,该指令将通过专用通道,同时传送2个源寄存器值到协处理器100中,协处理器100后继通过mem req请求,向主处理器200请求访问L1-Dcache数据,读取寄存器指向的Mem数据,访问Cache的指令传递通过mem_req和mem_resp来进行。
协处理单元500根据处理指令向对应的主处理器200发送数据存储单元的访问请求并读取目标数据,以根据目标数据进行指令处理操作;将根据目标数据进行指令处理操作的计算结果写回至处理指令对应的原始寄存器,并在接收到主处理器200的指令响应后,将原始寄存器中存储的计算结果写回至主处理器200的寄存器中。
(6)若当前正在处理的主处理器200取消发起使用请求时,请求处理单元300则向排序在下一位的主处理器200返回一个Ready信号,并向其他n-2个排序靠后的主处理器200返回Hold信号;依次类推,直至完成所有主处理器200的Ready信号反馈。
依次类推,直至完成n个主处理器CPU1、CPU2…、CPUn的处理指令的处理操作,以实现多个主控制器共享访问同一个协处理器100。
在一可实施的方案中,如图5所示的多核共享协处理器架构,共享协处理器100的多个主控制器采用哈佛结构,需要实现独立的L1指令Cache和L1数据Cache,整个多核处理器架构实现统一的L2Cache。每个主控制器采用Extension Interface(扩展接口)与协处理器进行连接。
本实施例中,通过每个主处理器与请求处理单元连接,使得在多个主处理器同时向请求处理单元发起使用协处理器的使用请求时,根据预先优先级设置,只返回一个主控制器CPU一个Ready信号,向其余的主控制器CPU分别返回一个Hold信号;当正在处理的主控制器CPU取消使用请求时则向排序在下一位的主控制器CPU返回一个Ready 信号,同时向其余的主控制器CPU分别返回一个Hold信号,依次类推,即根据使用场景需求进行划分,各主控制器CPU发出使用请求,由仲裁单元实现指令级互斥访问,从而实现协处理器作为共享资源,多个主控制器共享访问同一个协处理器,满足指令级的延时,达到性能和功耗的最佳平衡,有效地提高了资源利用率。
实施例2
本实施例的处理器微架构是对实施例1的进一步改进,具体地:
如图6所示,本实施例的每个主处理器200包括多个可配置的功能架构,该功能架构包括设置在主Core(处理器核)中的流水线架构1、扩展指令接口2、存储器架构3和TEE4,以及设于主Core外的扩展指令协处理单元5005。每个功能架构基于开源指令集架构进行配置。
其中,开源指令集架构包括基于精简指令集原则的开源指令集架构RISC-V,开源指令集架构RISC-V支持多种指令集;根据每种指令集配置对应的计算单元和流水线架构。
采用开源指令集架构RISC-V根据实际设计需求进行针对性设计、灵活搭配各种功能架构组合以应用在低功耗智能设备中,以满足不同客户的不同设计需求,实现处理器微架构的定制化需求。
处理器微结构可支持的指令集包括基础指令集、浮点指令集、压缩指令集、扩展指令集等。其中,基础指令集包括加法、减法、乘法、除法、原子交换、访问存储器等指令;浮点指令集包括单精度、双精度浮点计算;压缩指令集包括16bit,扩展指令集包括向量指令、SIMD(单指令多数据流)指令等。具体地,根据上述不同的指令集设计不同的ALU(计算单元),并设计不同的流水线架构。
另外,本实施例的处理器微架构包括扩展指令接口和扩展指令协处理单元500,扩展指令协处理单元500分别与扩展指令接口和系统总线通信连接,扩展指令协处理单元500用于基于扩展指令接口和系统总线对指令集进行扩展。
具体地,CPU处理器微架构设计中的主Core用于支持基础指令集、浮点指令集和压缩指令集等指令集,而扩展指令集的实现则需要依赖于主Core外部的扩展指令协处理单元500。例如向量指令协处理单元500就是用于处理扩展的向量指令,这样可以采用性能功耗比更优的逻辑单元来处理所需求的领域计算要求,即可以通过扩展指令接口以及扩展指令协处理单元500实现用户自定义优化指令集的目的。
可针对FPU(浮点运算的处理器)设计专门的浮点处理流水线,增加了浮点处理的计算单元,其与整型处理流水线共享指令预取和译码单元;当然可根据产品配置的需求,选择支持或不支持浮点处理指令集,保证了配置的灵活性,更好地满足更高的设计需求。
本实施例的功能架构包括但不限于流水线架构、存储器结构等。
具体地,流水线架构为多级流水线架构,配置后的多级流水线架构支持三级流水线架构或五级流水线架构。当然可以为其他级别的流水线架构,具体可以根据实际设计需求进行调整配置。
如图7所示,根据实现复杂度的不同,设计多级流水线架构。例如:一种可配置的多级流水线架构的方式为:在处理性能较高的应用处理器中使用经典的5级流水线架构;在处理性能要求较低,功耗面积要求达到最优的MCU控制器设计中,使用简单的3级流水线架构;由于流水线设计的基本原理,在3级流水线中,最高可支持的主频率低于5级流水线,这样牺牲了处理的能力,换来芯片面积和功耗的优化。
其中,本实施例可以对整型流水线、浮点流水线、3级流水线、5级流水线等形成独立可配置特性,也可以根据系统复杂度进行搭配组合设计。
功能架构包括多级存储器结构。配置后的多级存储器结构中的每级存储器架构对应不同类别的多个存储器;
其中,存储器包括L1 Cache、I-Cache、D-Cache、I-TCM、D-TCM或、MMU。
本实施例中的处理器微架构支持L1 Cache、I-Cache、D-Cache、I-TCM、D-TCM或MMU等存储器的灵活配置。根据不同的使用需求进行组合,例如在AP应用处理器环境中,需要将I-Cache、D-Cache、L2 Cache和MMU虚拟内存等进行组合配置;在低功耗要求更高的Sensor Hub需求中,则仅需要将I-Cache和D-TCM进行组合配置即可满足需求。
本实施例中的处理器微架构支持MMU架构,MMU是内存管理单元,位于CPU core与外部主存储器之间,通过加载页表的方式进行内存管理,主要实现虚拟地址到实际物理地址的变换。通过MMU可以实现虚拟内存的技术,这在扩展内存不足的嵌入式系统(如智能手表)中非常有效。MPU(微处理器)可以实现不同协处理单元500和不同MCU对主存空间的访问保护。只要将主存空间划分为不同的区域,并且配置MPU的读写权限,越权访问的请求则会被有效阻止。
另外,处理器微架构支持可信执行环境TEE。
本实施例中的处理器微架构支持实现TEE设计,采用特权模式设计,当设置特殊的指令时,系统则进入特权模式。此时在与正常模式完全隔离的硬件环境中执行可信的运行系统,包括独立的寄存器,独立且隔离的存储空间,独立且隔离的器件和TOS(可信操作系统)等。
对于SoC芯片系统,图1中处理器核中的Extension Interface表示扩展指令接口、 FIQ-CTL表示快速后中断请求控制模块、IRQ-CTL表示中断请求控制模块、Debug表示调试模块、JTAG表示调试接口、IRQ_src表示中断相应模块、Timer表示计时模块;Per1、Pern、Dev1均表示与总线连接的接线端、SRAM表示静态随机存取存储器、ExtMEM(一种寄存器)。
本实施例中的处理器微架构可以经上述的指令集设计、流水线设计、寄存器设计、Cache设计、扩展指令单元设计等,在同等条件下,本处理器微架构的功耗能力必然优于现有其他的CPU架构的实现。例如,对于本实施例的处理器微架构对应对的特定应用场景(如智能手表的心率检测功能 )计算需求,采用现有的ARM架构编译指令,可能需要多条汇编指令。而通过本实施例重新设计的处理器微架构的扩展指令,仅需采用更小的指令数就可以实现相同的功能;由于执行时钟周期减小,对应的功耗也随之得到优化,通过功能单元模块的搭配组合,以及用户自定义扩展指令等优化设计,有效降低了芯片面积和运行主频,从而达到功耗优化的目的。
对于复杂的嵌入式SoC系统,例如用于低功耗智能手表的SoC设计,很适合采用本实施例的处理器微架构进行模块化组合和定制化设计。下表为例如低功耗智能手表的SoC设计中各项参数的配置数据:
Figure PCTCN2021142830-appb-000001
本实施例中,可根据客户需求定制,设计多主处理器CPU共享同一个协处理器单元的CPU微架构,同时根据实际需求基于开源指令集架构RISC-V对不同功能结构(如实现L1 Cache、L2 Cache、MMU、TEE、浮点运算单元、向量运算单元等流水线架构、存储器架构等)进行灵活配置,搭配组合,可用于实现复杂的SoC系统中多功能子系统的MCU需求,从而提供了可精确配置、可调整功能特性、性能与功耗最优解的CPU微架构,满足对处理器CPU微架构的可定制化需求,进而满足低功耗智能手表系统的产品配置要求。
实施例3
本实施例的SoC芯片包括实施例1或2中的处理器微架构。
本实施例的SoC芯片包括上述的处理器微架构,可根据客户需求定制,设计多主处理器CPU共享同一个协处理器单元的CPU微架构,同时根据实际需求基于开源指令集 架构RISC-V对不同功能结构进行灵活配置,搭配组合,可用于实现复杂的SoC系统中多功能子系统的MCU需求,从而提供了可精确配置、可调整功能特性、性能与功耗最优解的CPU微架构,满足对处理器CPU微架构的可定制化需求,进而满足低功耗智能手表系统的产品配置要求。
实施例4
本实施例的低功耗智能设备包括SoC芯片。其中,低功耗智能设备包括智能手表。
本实施例的低功耗智能设备包括上述的SoC芯片,可根据客户需求定制,设计多主处理器CPU共享同一个协处理器单元的CPU微架构,同时根据实际需求基于开源指令集架构RISC-V对不同功能结构进行灵活配置,搭配组合,可用于实现复杂的SoC系统中多功能子系统的MCU需求,从而提供了可精确配置、可调整功能特性、性能与功耗最优解的CPU微架构,满足对处理器CPU微架构的可定制化需求,进而满足低功耗智能手表系统的产品配置要求。
虽然以上描述了本发明的具体实施方式,但是本领域的技术人员应当理解,这仅是举例说明,本发明的保护范围是由所附权利要求书限定的。本领域的技术人员在不背离本发明的原理和实质的前提下,可以对这些实施方式做出多种变更或修改,但这些变更和修改均落入本发明的保护范围。

Claims (18)

  1. 一种处理器微架构,其特征在于,所述处理器微架构包括协处理器和和至少两个主处理器,每个所述主处理器通过请求处理单元与所述协处理器连接;
    所述请求处理单元用于在接收到至少两个主处理器发起的使用请求时,根据第一预设条件确定发起请求的每个主处理器对应的处理顺序,并根据所述处理顺序生成并发送反馈指令至不同的主处理器;
    所述主处理器用于在接收的所述反馈指令表征允许使用时,则将处理指令发送至所述协处理器进行处理。
  2. 如权利要求1所述的处理器微架构,其特征在于,所述协处理器包括访问接口单元和协处理单元;
    所述主处理器用于将所述处理指令发送至所述访问接口单元;
    所述访问接口单元用于将所述处理指令传送至所述协处理单元进行处理;
    其中,所述请求处理单元独立设置在所述访问接口单元和每个所述主处理器之间;或,所述请求处理单元集成设置在所述访问接口单元中。
  3. 如权利要求1或2所述的处理器微架构,其特征在于,所述第一预设条件包括预先设定的每个所述主处理器对应的处理优先级。
  4. 如权利要求3所述的处理器微架构,其特征在于,所述请求处理单元用于根据所述处理顺序生成表征允许使用的第一反馈指令和表征继续等待的第二反馈指令,并将所述第一反馈指令发送至排序最靠前的所述主处理器,且将所述第二反馈指令分别发送至其他排序靠后的所述主处理器。
  5. 如权利要求4所述的处理器微架构,其特征在于,所述请求处理单元还用于在正在处理的所述主处理器取消发起使用请求时,将所述第一反馈指令发送至排序在下一位的所述主处理器,且将所述第二反馈指令分别发送至其他排序靠后的所述主处理器。
  6. 如权利要求2所述的处理器微架构,其特征在于,每个所述主处理器对应不同的电源域;和/或,
    所述访问接口单元和所述协处理单元划分在同一电源域中。
  7. 如权利要求2所述的处理器微架构,其特征在于,所述主处理器用于通过命令流方式将所述处理指令发送至所述访问接口单元;和/或,
    所述协处理单元用于在所述处理指令满足第二预设条件时确定采用阻塞式指令处理方式对所述处理指令进行处理;否则,采用流水线式的指令处理方式对所述处理指令进 行处理。
  8. 如权利要求2所述的处理器微架构,其特征在于,所述协处理单元用于根据所述处理指令向对应的所述主处理器发送数据存储单元的访问请求并读取目标数据,以根据所述目标数据进行指令处理操作;所述协处理单元还用于将根据所述目标数据进行指令处理操作的计算结果写回至所述处理指令对应的原始寄存器,并在接收到所述主处理器的指令响应后,将所述原始寄存器中存储的所述计算结果写回至所述主处理器的寄存器中。
  9. 如权利要求1所述的处理器微架构,其特征在于,所述协处理器支持用户自定义指令和/或Vector向量指令。
  10. 如权利要求1所述的处理器微架构,其特征在于,所述主处理器包括多个可配置的功能架构,每个所述功能架构基于开源指令集架构进行配置。
  11. 如权利要求10所述的处理器微架构,其特征在于,所述开源指令集架构包括基于精简指令集原则的开源指令集架构RISC-V;
    其中,所述开源指令集架构RISC-V支持多种指令集;
    根据每种所述指令集配置对应的计算单元和流水线架构。
  12. 如权利要求11所述的处理器微架构,其特征在于,所述指令集包括基础指令集、浮点指令集、压缩指令集或扩展指令集;和/或,
    配置后的所述流水线架构支持三级流水线架构或五级流水线架构。
  13. 如权利要求10所述的处理器微架构,其特征在于,所述功能架构包括多级存储器结构。
  14. 如权利要求13所述的处理器微架构,其特征在于,配置后的所述多级存储器结构中的每级存储器架构对应不同类别的多个存储器;
    其中,所述存储器包括L1 Cache、I-Cache、D-Cache、I-TCM、D-TCM或MMU。
  15. 如权利要求10所述的处理器微架构,其特征在于,所述主处理器还包括扩展指令接口和扩展指令协处理单元,所述扩展指令协处理单元分别与扩展指令接口和系统总线通信连接,所述扩展指令协处理单元用于基于所述扩展指令接口和所述系统总线对所述指令集进行扩展。
  16. 一种SoC芯片,其特征在于,所述SoC芯片包括权利要求1-15中任一项所述的处理器微架构。
  17. 一种低功耗智能设备,其特征在于,所述低功耗智能设备包括权利要求16所述的SoC芯片。
  18. 如权利要求17所述的低功耗智能设备,其特征在于,所述低功耗智能设备包括智 能手表。
PCT/CN2021/142830 2021-04-30 2021-12-30 处理器微架构、SoC芯片及低功耗智能设备 Ceased WO2022227671A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US18/288,627 US12510954B2 (en) 2021-04-30 2021-12-30 Processor micro-architecture, SoC chip and low-power-consumption intelligent device

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202110485283.3 2021-04-30
CN202110485283.3A CN113312303B (zh) 2021-04-30 2021-04-30 处理器微架构系统、SoC芯片及低功耗智能设备

Publications (1)

Publication Number Publication Date
WO2022227671A1 true WO2022227671A1 (zh) 2022-11-03

Family

ID=77371469

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2021/142830 Ceased WO2022227671A1 (zh) 2021-04-30 2021-12-30 处理器微架构、SoC芯片及低功耗智能设备

Country Status (3)

Country Link
US (1) US12510954B2 (zh)
CN (1) CN113312303B (zh)
WO (1) WO2022227671A1 (zh)

Families Citing this family (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113312303B (zh) * 2021-04-30 2022-10-21 展讯通信(上海)有限公司 处理器微架构系统、SoC芯片及低功耗智能设备
CN114356836B (zh) * 2021-11-29 2025-05-30 山东领能电子科技有限公司 基于risc-v的三维互联众核处理器架构及其工作方法
US11714649B2 (en) 2021-11-29 2023-08-01 Shandong Lingneng Electronic Technology Co., Ltd. RISC-V-based 3D interconnected multi-core processor architecture and working method thereof
CN114629665B (zh) * 2022-05-16 2022-07-29 百信信息技术有限公司 一种用于可信计算的硬件平台
CN115017087A (zh) * 2022-06-08 2022-09-06 深圳鲲云信息科技有限公司 一种传送dma控制信息的方法、装置、电子设备和存储介质
CN116245149B (zh) * 2022-12-20 2026-03-31 南京大学 一种基于risc-v指令集拓展的加速计算装置及方法
CN120011150B (zh) * 2024-12-27 2025-11-14 宁波甬华创芯科技发展有限责任公司 一种电力电子设备的异构双核处理器系统及电力电子设备
CN120256136B (zh) * 2025-06-03 2025-10-17 芯来智融半导体科技(上海)有限公司 多核处理器中指令集运算单元的配置方法和装置

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20090024834A1 (en) * 2007-07-20 2009-01-22 Nec Electronics Corporation Multiprocessor apparatus
CN101667165A (zh) * 2009-09-28 2010-03-10 中国电力科学研究院 一种分布式多主cpu共享总线的方法及其装置
CN102110072A (zh) * 2009-12-29 2011-06-29 中兴通讯股份有限公司 一种多处理器完全互访的方法及系统
CN102402422A (zh) * 2010-09-10 2012-04-04 北京中星微电子有限公司 处理器组件及该组件内存共享的方法
US20200097395A1 (en) * 2018-09-24 2020-03-26 Hewlett Packard Enterprise Development Lp Exception handling in wireless access points
CN113312303A (zh) * 2021-04-30 2021-08-27 展讯通信(上海)有限公司 处理器微架构、SoC芯片及低功耗智能设备

Family Cites Families (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR100243100B1 (ko) * 1997-08-12 2000-02-01 정선종 다수의 주프로세서 및 보조 프로세서를 갖는 프로세서의구조 및 보조 프로세서 공유 방법
US7996592B2 (en) * 2001-05-02 2011-08-09 Nvidia Corporation Cross bar multipath resource controller system and method
CN101187908A (zh) * 2007-09-27 2008-05-28 上海大学 单芯片多处理器共享数据存储空间的访问方法
US9678758B2 (en) * 2014-09-26 2017-06-13 Qualcomm Incorporated Coprocessor for out-of-order loads
CA2982785C (en) * 2015-04-14 2023-08-08 Capital One Services, Llc Systems and methods for secure firmware validation
US10642617B2 (en) * 2015-12-08 2020-05-05 Via Alliance Semiconductor Co., Ltd. Processor with an expandable instruction set architecture for dynamically configuring execution resources
US10235176B2 (en) * 2015-12-17 2019-03-19 The Charles Stark Draper Laboratory, Inc. Techniques for metadata processing
KR102563648B1 (ko) * 2018-06-05 2023-08-04 삼성전자주식회사 멀티 프로세서 시스템 및 그 구동 방법
US11119788B2 (en) * 2018-09-04 2021-09-14 Apple Inc. Serialization floors and deadline driven control for performance optimization of asymmetric multiprocessor systems
CN112130901A (zh) * 2020-09-11 2020-12-25 山东云海国创云计算装备产业创新中心有限公司 基于risc-v的协处理器、数据处理方法及存储介质

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20090024834A1 (en) * 2007-07-20 2009-01-22 Nec Electronics Corporation Multiprocessor apparatus
CN101667165A (zh) * 2009-09-28 2010-03-10 中国电力科学研究院 一种分布式多主cpu共享总线的方法及其装置
CN102110072A (zh) * 2009-12-29 2011-06-29 中兴通讯股份有限公司 一种多处理器完全互访的方法及系统
CN102402422A (zh) * 2010-09-10 2012-04-04 北京中星微电子有限公司 处理器组件及该组件内存共享的方法
US20200097395A1 (en) * 2018-09-24 2020-03-26 Hewlett Packard Enterprise Development Lp Exception handling in wireless access points
CN113312303A (zh) * 2021-04-30 2021-08-27 展讯通信(上海)有限公司 处理器微架构、SoC芯片及低功耗智能设备

Also Published As

Publication number Publication date
US12510954B2 (en) 2025-12-30
US20240211020A1 (en) 2024-06-27
CN113312303A (zh) 2021-08-27
CN113312303B (zh) 2022-10-21

Similar Documents

Publication Publication Date Title
WO2022227671A1 (zh) 处理器微架构、SoC芯片及低功耗智能设备
US12020031B2 (en) Methods, apparatus, and instructions for user-level thread suspension
US12554495B2 (en) Processor having multiple cores, shared core extension logic, and shared core extension utilization instructions
TWI628594B (zh) 用戶等級分叉及會合處理器、方法、系統及指令
EP2549382B1 (en) Virtual GPU
US9262353B2 (en) Interrupt distribution scheme
CN108885586B (zh) 用于以有保证的完成将数据取出到所指示的高速缓存层级的处理器、方法、系统和指令
JP5977094B2 (ja) フレキシブルフラッシュコマンド
Al-Shaikh et al. A Comparative Study on the Performance of 64-bit ARM Processors
CN104969182A (zh) 高动态范围软件-透明异构计算元件处理器、方法及系统
CN111522585A (zh) 基于平台热以及功率预算约束,对于给定工作负荷的最佳逻辑处理器计数和类型选择
JP2013025794A (ja) フラッシュインタフェースの有効利用
Kozyrakis A media-enhanced vector architecture for embedded memory systems
US9032099B1 (en) Writeback mechanisms for improving far memory utilization in multi-level memory architectures
CN103294449B (zh) 发散操作的预调度重演
US12487762B2 (en) Flexible provisioning of coherent memory address decoders in hardware
US11886910B2 (en) Dynamic prioritization of system-on-chip interconnect traffic using information from an operating system and hardware
Natvig et al. Multi‐and Many‐Cores, Architectural Overview for Programmers
Hussain Memory resources aware run-time automated scheduling policy for multi-core systems
US20260050469A1 (en) Strategy for instruction scheduling of multiple waves based on instruction status
He et al. A RISC-V Heterogeneous SoC and Its Co-Scheduling Optimization Approach for Digital Signal Processing
CN117667211A (zh) 指令同步控制方法、同步控制器、处理器、芯片和板卡
Itou et al. The Instruction Execution Mechanism for Responsive Multithreaded Processor.
Uchiyama et al. Chip Implementations
Chen et al. CoDMA: Buffer Avoided Data Exchange in Distributed Memory Systems

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21939130

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 18288627

Country of ref document: US

WWE Wipo information: entry into national phase

Ref document number: 202327079427

Country of ref document: IN

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 21939130

Country of ref document: EP

Kind code of ref document: A1

WWG Wipo information: grant in national office

Ref document number: 18288627

Country of ref document: US