WO2025001344A1 - Cxl数据传输板卡及控制数据传输的方法 - Google Patents

Cxl数据传输板卡及控制数据传输的方法 Download PDF

Info

Publication number
WO2025001344A1
WO2025001344A1 PCT/CN2024/083114 CN2024083114W WO2025001344A1 WO 2025001344 A1 WO2025001344 A1 WO 2025001344A1 CN 2024083114 W CN2024083114 W CN 2024083114W WO 2025001344 A1 WO2025001344 A1 WO 2025001344A1
Authority
WO
WIPO (PCT)
Prior art keywords
cxl
processor
host
control chip
memory module
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2024/083114
Other languages
English (en)
French (fr)
Inventor
赵建杰
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Suzhou Metabrain Intelligent Technology Co Ltd
Original Assignee
Suzhou Metabrain Intelligent Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Suzhou Metabrain Intelligent Technology Co Ltd filed Critical Suzhou Metabrain Intelligent Technology Co Ltd
Priority to US19/122,952 priority Critical patent/US12547579B2/en
Publication of WO2025001344A1 publication Critical patent/WO2025001344A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F13/00Interconnection of, or transfer of information or other signals between, memories, input/output devices or central processing units
    • G06F13/38Information transfer, e.g. on bus
    • G06F13/42Bus transfer protocol, e.g. handshake; Synchronisation
    • G06F13/4282Bus transfer protocol, e.g. handshake; Synchronisation on a serial bus, e.g. I2C bus, SPI bus
    • G06F13/4291Bus transfer protocol, e.g. handshake; Synchronisation on a serial bus, e.g. I2C bus, SPI bus using a clocked protocol
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F13/00Interconnection of, or transfer of information or other signals between, memories, input/output devices or central processing units
    • G06F13/38Information transfer, e.g. on bus
    • G06F13/40Bus structure
    • G06F13/4004Coupling between buses
    • G06F13/4022Coupling between buses using switching circuits, e.g. switching matrix, connection or expansion network
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L49/00Packet switching elements
    • H04L49/10Packet switching elements characterised by the switching fabric construction
    • H04L49/111Switch interfaces, e.g. port details
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F2213/00Indexing scheme relating to interconnection of, or transfer of information or other signals between, memories, input/output devices or central processing units
    • G06F2213/0026PCI express
    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02DCLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
    • Y02D10/00Energy efficient computing, e.g. low power processors, power management or thermal management

Definitions

  • the embodiments of the present application relate to the field of computers, and more particularly to a CXL data transmission board and a method for controlling data transmission.
  • the related technology has the problem of being unable to reallocate resources according to different workloads.
  • the embodiments of the present application provide a CXL data transmission board and a method for controlling data transmission, so as to at least solve the problem in the related art that resources cannot be reallocated according to different workloads.
  • a CXL data transmission board comprising: a control chip, the control chip is a chip supporting the open interconnection standard CXL protocol, an uplink port and a downlink port are deployed on the control chip; the uplink port is connected to a host, wherein the host is a high-speed serial computer expansion bus standard PCIe protocol device or a device supporting the above-mentioned CXL protocol; the above-mentioned downstream port is connected to the processor and/or the memory module, and the above-mentioned downstream port is correspondingly connected to the above-mentioned upstream port, and the above-mentioned host is configured to transmit data to the corresponding above-mentioned downstream port through the above-mentioned upstream port, so as to transmit the above-mentioned data to the above-mentioned processor and/or the above-mentioned memory module through the above-mentioned downstream port, wherein the above-mentioned processor and the above-mentioned processor and the above-menti
  • a method for controlling data transmission comprising: receiving a data request sent by a host through an upstream port, wherein the data request includes data to be transmitted, and the host is a device supporting the PCIe protocol or a device supporting the CXL protocol; in response to the data request, searching for routing information of the host from a routing table; transmitting the data to a processor and/or a memory module connected to a downstream port corresponding to the upstream port according to the routing information, wherein the downstream port is correspondingly connected to the upstream port, and the processor and the memory module are both devices supporting the PCIe protocol or devices supporting the CXL protocol.
  • the link switching control system includes the CXL data transmission board.
  • a non-volatile readable storage medium in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
  • an electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
  • the control chip in the CXL data transmission board is a chip that supports the CXL protocol, and an upstream port and a downstream port are deployed;
  • the upstream port is connected to the host, wherein the host is a device that supports the PCIe protocol or a device that supports the CXL protocol;
  • the downstream port is connected to the processor and/or the memory module, and the downstream port is connected to the upstream port accordingly, so that the host can be connected to both the processor and the memory module, so that the number of processors and memory modules can be flexibly configured according to actual needs. Therefore, the problem that resources cannot be reallocated according to different workloads in the related art can be solved, and the effect of realizing flexible resource allocation can be achieved.
  • FIG1 is a schematic diagram of the structure of a CXL data transmission board according to an embodiment of the present application.
  • FIG2 is a result schematic diagram of a CXL data transmission board according to an embodiment of the present application.
  • FIG3 is a connection diagram at the switch level according to an embodiment of the present application.
  • FIG4 is a second flowchart of a method for controlling signal transmission according to an embodiment of the present application.
  • FIG. 5 is a schematic diagram of an external manager managing a control chip according to an embodiment of the present application.
  • FIG6 is a flow chart of a method for controlling data transmission according to an embodiment of the present application.
  • FIG. 7 is a schematic diagram of an electronic device according to an embodiment of the present application.
  • CXL (Compute Express Link) is an open industrial standard for high-bandwidth and low-latency device interconnection. It can be used to connect devices such as CPU and GPU, memory and smart network card.
  • FIG. 1 is a schematic diagram of the structure of a CXL data transmission board according to an embodiment of the present application. As shown in FIG. 1 , the CXL data transmission board includes:
  • Control chip is a chip that supports the open interconnection standard CXL protocol.
  • the control chip is deployed with an uplink port and a downlink port;
  • the uplink port is connected to a host, wherein the host is a device supporting the high-speed serial computer expansion bus standard PCIe (Peripheral Component Interconnect express) protocol or a device supporting the CXL protocol;
  • PCIe Peripheral Component Interconnect express
  • the downstream port is connected to the processor and/or memory module, and the downstream port is correspondingly connected to the upstream port.
  • the host is configured to transmit data to the corresponding downstream port through the upstream port, so as to transmit data to the processor and/or memory module through the downstream port, wherein the processor and the memory module are both devices supporting the PCIe protocol or devices supporting the CXL protocol.
  • the CXL data transmission board in this embodiment can be applied to scenarios that require high-speed data transmission and large-scale data processing, such as data centers, high-performance computing, artificial intelligence cloud computing, and the like.
  • control chip is a chip that supports the CXL protocol, including a high-speed interconnection architecture CXL Switch Fabric.
  • CXL Switch Fabric can connect multiple CXL devices to form a computing platform that shares resources.
  • CXL Switch Fabric provides a scalable and flexible infrastructure for CXL devices, which can achieve communication and data transmission between devices by providing high-speed point-to-point connections between CXL devices.
  • CXL devices can be hosts or devices such as processors and memory modules.
  • the CXL Switch Fabric architecture is divided into three layers: transaction layer, link layer (Link Layer) and physical layer (Physical Layer).
  • a Flex Bus port is provided in the CXL Switch Fabric architecture, allowing devices to choose between PCIe devices and CXL devices.
  • the CXL.cache and CXL.mem protocols are combined to share a common transaction layer and link layer, while CXL.io has its own transaction layer and link layer.
  • the CXL link layer interacts with the CXL ARB or MUX to interleave traffic from two logical streams.
  • the physical layer also contains two sublayers, namely the logical sublayer and the electrical sublayer. Among them, the logical sublayer can switch between PCIe mode and CXL mode, while the electrical sublayer follows the PCIe specification.
  • CXL.io provides an incoherent load or store interface for I/O devices.
  • the processor can be a CPU or a GPU.
  • the CPU and GPU can bypass the PCIe protocol and use the CXL protocol to share and access each other's memory resources. Through the CXL protocol, the CPU and GPU are connected to form a single huge stack memory pool.
  • the GPU BOX can support both the PCIE protocol and the CXL protocol.
  • the CXL data transmission board can be a CXL Switch board
  • the control chip can be a CXL Switch chip
  • the processor can be a GPU
  • the upstream device can be connected to multiple hosts HOST
  • the downstream device can be connected to the GPU BOX, and can also be connected to the memory module, with flexible configuration.
  • each host can be connected to different downstream devices, and the connection relationship between the host HOST and the downstream device can be changed by changing the burning information of the CXL Switch chip.
  • the number of connected upstream devices and downstream devices can be determined by the function of the CXL Switch chip.
  • the GPU BOX supports 16 GPU slots inside and is interconnected with the CXL Switch chip through a CDFP cable.
  • the onboard CPLD Complex Programmable Logic Device
  • BMC Bus Master Controller
  • the memory module is a device based on the CXL bus protocol, which can realize remote memory expansion and multi-host shared memory.
  • the CDFP interface of the memory board is interconnected with the CXL Switch board through a CDFP cable.
  • the onboard BMC and CPLD are responsible for the management of the CXL Switch board, including heat dissipation control, power on and off, status indication, etc.
  • the control chip in the CXL data transmission board is a chip that supports the CXL protocol, and an upstream port and a downstream port are deployed;
  • the upstream port is connected to the host, wherein the host is a device that supports the PCIe protocol or a device that supports the CXL protocol;
  • the downstream port is connected to the processor and/or the memory module, and the downstream port is connected to the upstream port accordingly, so that the host can be connected to both the processor and the memory module, so that the number of processors and memory modules can be flexibly configured according to actual needs. Therefore, the problem that resources cannot be reallocated according to different workloads in the related art can be solved, and the effect of realizing flexible resource allocation can be achieved.
  • N there are N upstream ports and N downstream ports, where N is a natural number greater than 1.
  • the value of N is determined by the function of the control chip.
  • N can be 16 or 4.
  • 6 upstream and downstream ports are set in the CXL Switch chip.
  • control chip includes at least one of the following:
  • a control unit is configured to control a topology of a CXL link between an upstream port and a downstream port, wherein the CXL link is configured to connect corresponding upstream ports and downstream ports.
  • the management unit is configured to manage routing information between the host and the processor and/or the memory module.
  • An identification unit is connected to the processor and/or the memory module, and the identification unit is configured to identify the protocol type supported by the processor and/or the memory module when the host establishes a connection with the processor and/or the memory module.
  • control unit when the control chip is a chip of CXL Switch Fabric architecture, the control unit can be composed of multiple CXL Switch Fanouts.
  • CXL switch Fanout refers to the process of copying a CXL link to multiple other links in CXL Switch Fabric. This process allows multiple devices to access and share devices on the same CXL link at the same time, thereby improving system performance and scalability.
  • the control chip can implement CXL switch Fanout because CXL supports point-to-point topology, which allows multiple devices to communicate with each other through a CXL link and transfer data packets in the link.
  • the main purpose of CXL switch Fanout is to improve the connection efficiency and data throughput between multiple devices in the system.
  • the management unit can be composed of multiple CXL Switch Roots.
  • CXL devices are connected to CXL Switch Fabric through CXL ports.
  • CXL Switch Root maintains a routing table to store and manage routing information between nodes in the network.
  • the routing table selects the best path and forwards the data to the destination CXL device, for example, transferring the data in HOST1 to the memory module for storage.
  • the traffic manager of the CXL port in the CXL data transmission board stores the data in the buffer in sequence for fast processing and transmission.
  • CXL Switch Fabric selects the best path in the buffer according to the information of the input port and the routing table, and forwards the data to the port of the target CXL device.
  • CXL Switch Fabric While forwarding data, CXL Switch Fabric also needs to ensure the stability of the internal communication speed of the control chip and the consistency of performance. When the CXL device needs to return data, CXL Switch Fabric will also find the best path according to the routing information and the status of the buffer to transmit the data to the target device.
  • the identification unit can use BIOS commands to identify whether the port is connected to a PCLE device or a CXL device. For example, taking the Intel Eagle Stream series CPU as an example, the CXL interconnection is completely dependent on the PCIe 5.0 electrical layer, and all topologies rely on PCIe. Eagle Stream CPU can support CXL 4X16 channels. CXL and PCI Express devices can work simultaneously on a given port, with 8 lanes for each device. Any x16 PCI Express port can be connected to a PCI Express device or a CXL device. The identification unit can automatically detect whether the other end is a PCI Express card or a CXL device, and dynamically configure the link.
  • the CXL mode When connecting a CXL device, the CXL mode should be set in advance in the BIOS, and BIOS commands can be used to identify whether the port is connected to a PCLE device or a CXL device.
  • This embodiment can control the link topology and routing information, as well as identify the type of device, through the control chip, thereby Resources can be deployed quickly.
  • the number of hosts is M
  • the number of processors is P
  • the number of memory modules is K
  • M, P, and K are all natural numbers greater than or equal to 1.
  • M hosts share K memory modules.
  • N 16
  • M is greater than or equal to 2 and less than or equal to 7.
  • P 7
  • the processors and hosts correspond one to one, or one host corresponds to multiple processors, wherein data is transmitted between the host and the processor via the PCIe protocol.
  • K is greater than 1
  • one host corresponds to multiple memory modules, and multiple memory modules form a memory resource pool of a host, wherein data is transmitted between the host and the memory module via the CXL protocol, and the number of hosts is proportional to the number of memory modules.
  • control chip when the control chip uses the CXL 2.0 switch chip, the control chip is compatible with PCIe5.0 and supports switching of different firmware (programs written in EPROM (erasable programmable read-only memory) or EEPROM (electrically erasable programmable read-only memory)).
  • the control chip has a total of 16 ports, each of which can be configured as PCIe or CXL, supporting the mixed use of PCIe and CXL.
  • a maximum of 7 HOSTs can be connected, and a minimum of 2 can be connected.
  • the upstream device can be connected to 7 HOSTs, and the downstream device can be connected to 7 GPUs.
  • Each HOST is connected to a GPU, and data can be processed in parallel in multiple channels.
  • each host has 4 X16 signals
  • HOST1 is connected to 4 GPUs and is set to accelerate data processing
  • HOST2 is connected to a 4X16 memory module and is set to read and store data; it can also be connected to 4 hosts upstream and 8X16 memory modules downstream, using the CXL protocol.
  • Each host has two groups of X16 ports, connected to two X16 memory modules.
  • the control chip can be composed of multiple VCXLs, which can be mapped to a specified physical port according to the host's instruction requirements, and can support the allocation of up to 16 logical storage spaces.
  • Each VCXL has its own ID, which is used to bind or unbind the mapping of VCXL to the physical port under the control of the host to achieve configuration between different ports.
  • the ports in this embodiment can be flexibly configured based on actual application requirements, and the number of hosts and the port resource information of the hosts can also be flexibly configured, thereby achieving the purpose of flexible allocation of resources.
  • the CXL data transmission board also includes: a connector, the connector connects the control chip and the manager, wherein the manager is configured to manage the operation of the control chip through the connector.
  • the connector may be a cable assembly connector MCIO, as shown in FIG2
  • the CXL Switch chip is connected to an external manager CPU through the connector MCIO.
  • the manager CPU manages the I2C, RESET, status indication, etc. of the CXL Switch chip. This embodiment manages the control chip through an external manager, and can accurately control the working state of the control chip.
  • the CXL data transmission board further includes: a clock generator connected to the control chip, configured to generate a first initialization clock signal to start the control chip, and configured to start the control chip at the control chip. After the chip is started, a first clock signal of the clock frequency of the coordinated control chip is generated.
  • the clock generator is connected to the control chip, the processor and the memory module through a preset interface, and is configured to generate a second initialization clock signal for starting the control chip, the processor and the memory module respectively, and is configured to generate a second clock signal of the clock frequency of the coordinated control chip, the processor and the memory module after the control chip, the processor and the memory module are started.
  • the CXL data transmission board supports the RJ45 interface, which is the BMC management network port.
  • Each CDFP interface is equipped with a set of red and green two-color lights.
  • the red light is not on when the data transmission is normal, and it is always on after a fault; the green light is always on when there is no data transmission, and flashes at a frequency of 1Hz during normal data transmission.
  • This embodiment can accurately monitor the transmission of data through the clock generator.
  • the CXL data transmission board also includes: a debugging device, which is connected to the control chip and is configured to debug the CXL performance of the control chip.
  • the debugging scope of the debugging device includes: CXL margin testing, which is the same as PCIe debugging; CXL link speed, width, and compliance testing; enabling or disabling interleaving mode for system stress or performance or bandwidth testing.
  • CXL compliance testing CXL stress testing, CXL delay testing, etc.
  • link and protocol error handling and PCIe ERR messages are used to record and display CXL.io and CXL.io protocol errors. This embodiment can ensure the correctness of data transmission through the debugging of the debugging device.
  • FIG. 5 is a hardware structure block diagram of a mobile terminal of a method for controlling data transmission in an embodiment of the present application.
  • the mobile terminal may include one or more (only one is shown in FIG. 5 ) processors 502 (the processor 502 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 504 configured to store data, wherein the mobile terminal may also include a transmission device 506 and an input/output device 508 configured to have a communication function.
  • processors 502 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA
  • a memory 504 configured to store data
  • the mobile terminal may also include a transmission device 506 and an input/output device 508 configured to have a communication function.
  • FIG. 5 is for illustration only and does not limit the structure of the mobile terminal.
  • the mobile terminal may also include more or fewer components than those shown in FIG. 1 , or have a configuration different from that shown in FIG. 5 .
  • the memory 504 may be configured to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for controlling data transmission in the embodiment of the present application.
  • the processor 502 executes various functional applications and data processing by running the computer program stored in the memory 504, that is, to implement the above method.
  • the memory 504 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory.
  • the memory 504 may include a memory remotely arranged relative to the processor 502, and these remote memories may be connected to the mobile terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
  • the transmission device 506 is configured to receive or send data via a network. Examples of the above network may include mobile
  • the transmission device 506 may be a wireless network provided by a communication provider of the mobile terminal.
  • the transmission device 506 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet.
  • the transmission device 506 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
  • RF radio frequency
  • FIG. 6 is a flow chart of the method for controlling data transmission according to an embodiment of the present application. As shown in FIG. 6 , the flow chart includes the following steps:
  • Step S602 receiving a data request sent by a host through an uplink port, wherein the data request includes data to be transmitted, and the host is a device supporting the PCIe protocol or a device supporting the CXL protocol;
  • Step S604 responding to the data request, searching the routing information of the host from the routing table;
  • Step S606 transmitting the data to the processor and/or memory module connected to the downlink port corresponding to the uplink port according to the routing information, wherein the downlink port is connected to the uplink port correspondingly, and the processor and the memory module are both devices supporting the PCIe protocol or devices supporting the CXL protocol.
  • the execution subject of the above steps may be a processor provided in a control chip, or a processor or processing device provided relatively independently from a terminal or a server, but is not limited thereto.
  • This embodiment can be applied to scenarios that require high-speed data transmission and large-scale data processing, such as data centers, high-performance computing, artificial intelligence cloud computing, and the like.
  • control chip is a chip that supports the CXL protocol, including a high-speed interconnection architecture CXL Switch Fabric.
  • CXL Switch Fabric can connect multiple CXL devices to form a computing platform that shares resources.
  • CXL Switch Fabric provides a scalable and flexible infrastructure for CXL devices, which can achieve communication and data transmission between devices by providing high-speed point-to-point connections between CXL devices.
  • CXL devices can be hosts or devices such as processors and memory modules.
  • the CXL Switch Fabric architecture is divided into three layers: transaction layer, link layer, and physical layer.
  • a Flex Bus port is provided in the CXL Switch Fabric architecture, allowing devices to choose between PCIe devices and CXL devices.
  • CXL.cache and CXL.mem protocols are combined to share a common transaction layer and link layer, while CXL.io has its own transaction layer and link layer.
  • the CXL link layer interacts with CXL ARB or MUX to interweave traffic from two logical streams.
  • the physical layer also contains two sublayers, namely the logical sublayer and the electrical sublayer. Among them, the logical sublayer can switch between PCIe mode and CXL mode, while the electrical sublayer follows the PCIe specification.
  • CXL.io provides a non-coherent load or store interface for I/O devices.
  • the processor can be a CPU or a GPU.
  • the CPU and GPU can bypass the PCIe protocol and use the CXL protocol to share and access each other's memory resources. Through the CXL protocol, the CPU and GPU are connected as a single A huge stack memory pool.
  • the GPU BOX can support both the PCIE protocol and the CXL protocol.
  • the CXL data transmission board can be a CXL Switch board
  • the control chip can be a CXL Switch chip
  • the processor can be a GPU
  • the upstream device can be connected to multiple hosts HOST
  • the downstream device can be connected to the GPU BOX, or it can be connected to a memory module, and the configuration is flexible.
  • Each host can be connected to different downstream devices, and the connection relationship between the host HOST and the downstream device can be changed by changing the burning information of the CXL Switch chip.
  • the number of connected upstream devices and downstream devices can be determined by the function of the CXL Switch chip.
  • the GPU BOX supports 16 GPU slots internally and is interconnected with the CXL Switch chip through a CDFP cable.
  • the onboard CPLD is set to control the power-on timing of the BOX
  • the onboard bus master controller BMC Bus Master Controller
  • the memory module is a device based on the CXL bus protocol, which can realize remote memory expansion and multi-host shared memory.
  • the CDFP interface of the memory board is interconnected with the CXL Switch board through a CDFP cable.
  • the onboard BMC and CPLD are responsible for the management of the CXL Switch board, including heat dissipation control, power on and off, status indication, etc.
  • the control chip in the CXL data transmission board is a chip that supports the CXL protocol, and an upstream port and a downstream port are deployed;
  • the upstream port is connected to the host, where the host is a device that supports the PCIe protocol or a device that supports the CXL protocol;
  • the downstream port is connected to the processor and/or memory module, and the downstream port is connected to the upstream port accordingly, so that the host can connect to both the processor and the memory module, so that the number of processors and memory modules can be flexibly configured according to actual needs. Therefore, the problem that resources cannot be reallocated according to different workloads in the related art can be solved, and the effect of flexible resource allocation can be achieved.
  • N upstream ports and N downstream ports there are N upstream ports and N downstream ports, where N is a natural number greater than 1.
  • N upstream ports and N downstream ports where N is a natural number greater than 1.
  • the value of N is determined by the function of the control chip. For example, N can be 16 or 4.
  • 6 upstream and downstream ports are set in the CXL Switch chip.
  • the value of the flexible device N can flexibly set the connected upstream and downstream devices to achieve flexible allocation of resources.
  • the method in response to a data request, before searching for the routing information of the host from the routing table, the method further includes: setting the topology of the CXL link between N upstream ports and N downstream ports, wherein the CXL link is set to connect the corresponding upstream ports and downstream ports; determining the route between the host and the processor and/or memory module based on the topology of the CXL link to obtain a routing table.
  • transmitting the data to the processor and/or memory module connected to the downstream port corresponding to the upstream port according to the routing information includes: caching the data in the cache unit through the downstream port; selecting the target route from the routing information; transmitting the data in the cache unit to the processor and/or memory module connected to the downstream port according to the target route.
  • the control unit when the control chip is a chip of CXL Switch Fabric architecture, the control unit may be The CXL switch Fanout is composed of multiple CXL Switch Fanouts.
  • CXL switch Fanout refers to the process of copying a CXL link to multiple other links in the CXL Switch Fabric. This process allows multiple devices to access and share devices on the same CXL link at the same time, thereby improving system performance and scalability.
  • the control chip can implement CXL switch Fanout because CXL supports point-to-point topology, which allows multiple devices to communicate with each other through a CXL link and transfer data packets in the link.
  • the main purpose of CXL switch Fanout is to improve the connection efficiency and data throughput between multiple devices in the system.
  • CXL devices are connected to the CXL Switch Fabric through CXL ports.
  • the CXL Switch Root maintains a routing table to store and manage routing information between nodes in the network.
  • the routing table selects the best path and forwards the data to the destination CXL device, for example, transferring the data in HOST1 to the memory module for storage.
  • the traffic manager of the CXL port in the CXL data transmission board stores the data in the buffer in sequence for fast processing and transmission.
  • the CXL Switch Fabric selects the best path in the buffer and forwards the data to the port of the target CXL device.
  • CXL Switch Fabric While forwarding data, CXL Switch Fabric also needs to ensure the stability of the communication speed and performance consistency within the control chip. When the CXL device needs to return data, CXL Switch Fabric will also find the best path based on the routing information and the status of the cache area to transmit the data to the target device.
  • the method before transmitting data to a processor and/or memory module connected to a downstream port corresponding to an upstream port according to routing information, the method further includes: when a host establishes a connection with the processor and/or memory module, identifying the protocol type supported by the processor and/or memory module, so as to transmit data to the protocol type supported by the processor and/or memory module according to the protocol type supported by the processor and/or memory module.
  • a BIOS command is used to identify whether the port is connected to a PCLE device or a CXL device.
  • the host takes the CPU of the Intel Eagle Stream series as an example, and the CXL interconnection completely relies on the PCIe 5.0 electrical layer, and all topologies rely on PCIe.
  • the Eagle Stream CPU can support CXL 4X16 channels.
  • CXL and PCI Express devices can work simultaneously under a given port, with 8 lanes for each device.
  • Any x16 PCI Express port can be connected to a PCI Express device or a CXL device.
  • the identification unit can automatically detect whether the other end is a PCI Express card or a CXL device, and dynamically configure the link.
  • the CXL mode should be set in advance in the BIOS, and the BIOS command can be used to identify whether the port is connected to a PCLE device or a CXL device.
  • This embodiment can control the link topology and routing information and identify the type of device through the control chip, so that resources can be quickly allocated.
  • the host includes M
  • the processor includes P
  • the memory module includes K
  • M, P, and K are all natural numbers greater than or equal to 1.
  • M is greater than 1
  • the M hosts share the K memory modules.
  • N 16
  • M is greater than or equal to 2 and less than or equal to 7.
  • N 16
  • M is 7,
  • P is 7, and the processors and hosts correspond one to one, or one host corresponds to multiple processors, wherein the host and Data is transmitted between processors via the PCIe protocol.
  • control chip when the control chip uses the CXL 2.0 switch chip, the control chip is compatible with PCIe 5.0 and supports the switching of different firmware (programs written in EPROM (erasable programmable read-only memory) or EEPROM (electrically erasable programmable read-only memory)).
  • the control chip has a total of 16 ports, each of which can be configured as PCIe or CXL, supporting the mixed use of PCIe and CXL. It can connect up to 7 hosts and at least 2.
  • the upstream device can connect to 7 hosts and the downstream device can connect to 7 GPUs. Each host connects to one GPU, and data can be processed in multiple channels in parallel.
  • each host can also connect to 2 hosts, each host outputs 4 X16 signals, HOST1 connects to 4 GPUs and is set to accelerate data processing, and HOST2 connects to 4X16 memory modules and is set to read and store data; it can also connect to 4 hosts upstream and 8X16 memory modules downstream, using the CXL protocol.
  • Each host has two groups of X16 ports, connected to two X16 memory modules.
  • the control chip can be composed of multiple VCXLs, which can be mapped to a specified physical port according to the host's instruction requirements, and can support the allocation of up to 16 logical storage spaces.
  • Each VCXL has its own ID, which is used to bind or unbind the mapping of VCXL to the physical port under the control of the host to achieve configuration between different ports.
  • the ports in this embodiment can be flexibly configured based on actual application requirements, and the number of hosts and the port resource information of the hosts can also be flexibly configured, thereby achieving the purpose of flexible allocation of resources.
  • one host corresponds to multiple memory modules, and the multiple memory modules constitute a memory resource pool of the host, wherein data is transmitted between the host and the memory module via the CXL protocol, and the number of hosts is directly proportional to the number of memory modules.
  • the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
  • the technical solution of the present application, or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a non-volatile readable storage medium (such as ROM/RAM, magnetic disk, optical disk), including a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.
  • a non-volatile readable storage medium such as ROM/RAM, magnetic disk, optical disk
  • a terminal device which can be a mobile phone, computer, server, or network device, etc.
  • control system for link switching includes the above-mentioned CXL data transmission board.
  • the device is configured to implement the above-mentioned embodiment and optional implementation modes, and those that have been explained will not be repeated.
  • the embodiment of the present application further provides a non-volatile readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute any one of the above method embodiments when it is run. steps.
  • the above-mentioned non-volatile readable storage medium may include, but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other non-volatile readable storage media that can store computer programs.
  • An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
  • FIG7 is a schematic diagram of an electronic device according to an embodiment of the present application.
  • the electronic device includes a memory and a processor.
  • a computer program is stored in the memory.
  • the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
  • the electronic device may further include a transmission device and an input/output device, wherein the transmission device is connected to the processor, and the input/output device is connected to the processor.
  • modules or steps of the present application can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation.
  • the present application is not limited to any specific combination of hardware and software.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Signal Processing (AREA)
  • Mathematical Physics (AREA)
  • Computer Hardware Design (AREA)
  • Multi Processors (AREA)
  • Communication Control (AREA)

Abstract

一种CXL数据传输板卡及控制数据传输的方法,CXL数据传输板卡包括:控制芯片,控制芯片是支持开放式互连标准CXL协议的芯片,控制芯片上部署了上行端口和下行端口;上行端口与主机连接,其中,主机是支持高速串行计算机扩展总线标准PCIe协议的设备或支持CXL协议的设备;下行端口与处理器和/或内存模组连接,下行端口与上行端口对应连接,主机被设置为通过上行端口将数据传输至对应的下行端口,以通过下行端口将数据传输至处理器和/或内存模组。通过本申请,解决了相关技术中存在的无法根据不同的工作负载进行资源的重新调配的问题,达到实现灵活调配资源的效果。

Description

CXL数据传输板卡及控制数据传输的方法
相关申请的交叉引用
本申请要求于2023年06月28日提交中国专利局,申请号为202310776401.5,申请名称为“CXL数据传输板卡及控制数据传输的方法”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请实施例涉及计算机领域,特别涉及一种CXL数据传输板卡及控制数据传输的方法。
背景技术
伴随云计算应用的发展,信息化逐渐覆盖到社会的各个领域,人们的日常工作越来越多的通过网络来交流,网络数据量也在爆发式的增长,服务器作为处理和存储数据的核心设备,对性能和配置的要求也越来越高。当今服务器技术的发展,正处于内存架构的瓶颈问题:内存通道数量的增长,已经赶不上CPU(Central Processing Unit,中央处理器)核心数量的增长,从而导致每个核心可以处理的内存带宽降低,限制了处理器的性能。
目前内存和处理器是紧耦合的,内存都部署在服务器节点内。内存成本是占整个服务器成本的比例很高,但是在实际使用中,内存的使用效率并不高,有的内存空间根本没有被访问,有的内存空间则存放了一些比较冷的数据,它访问的频率其实很低。这部分的内存没能很好地发挥它的价值。相关技术中的CXL(Compute Express Link,开放式互连标准)内存的扩展并不能根据不同的工作负载进行资源的重新调配。
由此可见,相关技术中存在无法根据不同的工作负载进行资源的重新调配的问题。
发明内容
本申请实施例提供了一种CXL数据传输板卡及控制数据传输的方法,以至少解决相关技术中存在的无法根据不同的工作负载进行资源的重新调配的问题。
根据第一方面,提供了一种CXL数据传输板卡,包括:控制芯片,上述控制芯片是支持开放式互连标准CXL协议的芯片,上述控制芯片上部署了上行端口和下行端口;上述上行端口与主机连接,其中,上述主机是支持高速串行计算机扩展总线标准PCIe协议 的设备或支持上述CXL协议的设备;上述下行端口与处理器和/或内存模组连接,上述下行端口与上述上行端口对应连接,上述主机被设置为通过上述上行端口将数据传输至对应的上述下行端口,以通过上述下行端口将上述数据传输至上述处理器和/或上述内存模组,其中,上述处理器和上述内存模组均是支持上述PCIe协议的设备或支持上述CXL协议的设备。
根据第二方面,提供了一种控制数据传输的方法,包括:通过上行端口接收主机发送的数据请求,其中,上述数据请求中包括待传输的数据,上述主机是支持PCIe协议的设备或支持CXL协议的设备;响应上述数据请求,从路由表中查找上述主机的路由信息;按照上述路由信息将上述数据传输至与上述上行端口对应的下行端口连接的处理器和/或内存模组中,其中,上述下行端口与上述上行端口对应连接,上述处理器和上述内存模组均是支持上述PCIe协议的设备或支持上述CXL协议的设备。
根据第三方面,提供了一种链路交换的控制系统,上述链路交换的控制系统包括上述的CXL数据传输板卡。
根据第四方面,还提供了一种非易失性可读存储介质,非易失性可读存储介质中存储有计算机程序,其中,计算机程序被设置为运行时执行上述任一项方法实施例中的步骤。
根据第五方面,还提供了一种电子设备,包括存储器和处理器,存储器中存储有计算机程序,处理器被设置为运行计算机程序以执行上述任一项方法实施例中的步骤。
通过本申请,CXL数据传输板卡中的控制芯片是支持CXL协议的芯片,并部署了上行端口和下行端口;上行端口与主机连接,其中,主机是支持PCIe协议的设备或支持CXL协议的设备;下行端口与处理器和/或内存模组连接,下行端口与上行端口对应连接,使得主机即可以接处理器又可以接内存模组,从而可以根据实际需要灵活配置处理器和内存模组的数量。因此,可以解决相关技术中存在的无法根据不同的工作负载进行资源的重新调配的问题,达到实现灵活调配资源的效果。
附图说明
图1是根据本申请实施例的CXL数据传输板卡的结构示意图;
图2是根据本申请实施例的CXL数据传输板卡的结果示意图;
图3是根据本申请实施例的switch层面的连接图;
图4是根据本申请实施例的控制信号传输的方法的流程图二;
图5是根据本申请实施例的外部管理器对控制芯片进行管理的示意图;
图6是根据本申请实施例的控制数据传输的方法的流程图;
图7是根据本申请实施例的电子设备的示意图。
具体实施方式
下文中将参考附图并结合实施例来详细说明本申请的实施例。
需要说明的是,本申请的说明书和权利要求书及上述附图中的术语“第一”、“第二”等是用于区别类似的对象,而不必用于描述特定的顺序或先后次序。
下面对本实施例中的相关技术解释如下:
CXL:(Compute Express Link),是一种开放工业标准用于高带宽低延迟的设备互联。它可以用来连接CPU和GPU,内存和智能网卡等类型的设备。
GPU,图像处理器(Graphic Processing Unit);
CXL SW,计算高速链路交换机(Compute Express Link switch);
CPU,中央处理器(Central Processing Unit)。
在本实施例中提供了一种CXL数据传输板卡,图1是根据本申请实施例的CXL数据传输板卡的结构示意图,如图1所示,该CXL数据传输板卡包括:
控制芯片,控制芯片是支持开放式互连标准CXL协议的芯片,控制芯片上部署了上行端口和下行端口;
上行端口与主机连接,其中,主机是支持高速串行计算机扩展总线标准PCIe(Peripheral Component Interconnect express)协议的设备或支持CXL协议的设备;
下行端口与处理器和/或内存模组连接,下行端口与上行端口对应连接,主机被设置为通过上行端口将数据传输至对应的下行端口,以通过下行端口将数据传输至处理器和/或内存模组,其中,处理器和内存模组均是支持PCIe协议的设备或支持CXL协议的设备。
本实施例中的CXL数据传输板卡可以应用于需要高速数据传输和大规模数据处理的场景中,例如数据中心,高性能计算,人工智能云计算等场景。
在本实施例中,控制芯片是支持CXL协议的芯片,包括高速互连架构CXL Switch Fabric。CXL Switch Fabric可以连接多个CXL设备,形成一个共享资源的计算平台。CXL Switch Fabric给CXL设备提供了一个可扩展、灵活的基础架构,可以通过在CXL设备之间提供高速的点对点连接来实现设备之间的通信和数据传输。CXL设备可以是主机也可以是处理器、内存模组等设备。
可选地,CXL Switch Fabric架构分为三层:事务层(Transaction Layer)、链路层 (Link Layer)和物理层(Physical Layer)。为了同时兼容PCIe协议和CXL协议,CXL Switch Fabric架构中提供了Flex Bus端口,使设备可以在PCIe设备和CXL设备中进行选择。CXL.cache和CXL.mem协议被组合在一起共享一个公共的事务层和链路层,而CXL.io拥有自己的事务层和链路层。CXL链路层与CXL ARB或MUX交互,从而交织来自两个逻辑流的流量。物理层中还包含两个子层,分别是逻辑子层和电气子层。其中,逻辑子层可以在PCIe模式和CXL模式之间切换,而电气子层则遵循PCIe规范。在事务层中,CXL.io为I/O设备提供了一个非一致性的加载或存储接口。
可选地,处理器可以是CPU,也可以是GPU,CPU与GPU之间可以绕过PCIe协议,用CXL协议来共享,互取对方的内存资源。透过CXL协议,CPU与GPU之间形同连成单一个庞大的堆栈内存池,本实施例中的,GPU BOX(盒)既可以支持PCIE协议,又可以支持CXL协议。如图2所示,CXL数据传输板卡可以是CXL Switch板,控制芯片可以是CXL Switch芯片,处理器可以是GPU,上行设备可以连接多个主机HOST,下行设备可以连接GPU BOX,也可以接内存模组,配置灵活。且每个主机均可以和不同的下行设备连接,通过改变CXL Switch芯片的烧录信息来改变主机HOST和下行设备的连接关系。连接的上行设备和下行设备的数量可以由CXL Switch芯片的功能决定。
可选地,在下行设备为GPU BOX时,GPU BOX内部支持16个GPU插槽,与CXL Switch芯片通过CDFP线缆进行互联,板载CPLD(Complex Programmable Logic Device,复杂可编程逻辑器件)被设置为控制BOX的上电时序,板载总线主控制器BMC(Bus Master Controller)被设置为管理GPU的工作状态。
可选地,内存模组是基于CXL总线协议的设备,可以实现内存远端拓展,多主机共享内存,内存板的CDFP接口通过CDFP线缆与CXL Switch板卡互联,板载BMC和CPLD负责CXL Switch板卡d的管理,包括散热控制、上下电、状态指示等。
通过本申请,CXL数据传输板卡中的控制芯片是支持CXL协议的芯片,并部署了上行端口和下行端口;上行端口与主机连接,其中,主机是支持PCIe协议的设备或支持CXL协议的设备;下行端口与处理器和/或内存模组连接,下行端口与上行端口对应连接,使得主机即可以接处理器又可以接内存模组,从而可以根据实际需要灵活配置处理器和内存模组的数量。因此,可以解决相关技术中存在的无法根据不同的工作负载进行资源的重新调配的问题,达到实现灵活调配资源的效果。
在一个示例性实施例中,上行端口和下行端口均为N个,N是大于1的自然数。N的取值由控制芯片的功能确定,例如,N可以是16,也可以是4等,如图2所示,CXL Switch芯片中设置6个上下行端口。本实施例通过灵活设置N的取值,可以灵活的设置 连接的上下行设备,实现对资源的灵活调配。
在一个示例性实施例中,控制芯片包括以下至少之一:
控制单元,控制单元被设置为控制上行端口和下行端口之间的CXL链路的拓扑,其中,CXL链路被设置为连接对应的上行端口和下行端口。
管理单元,管理单元被设置为管理主机与处理器和/或内存模组之间的路由信息。
识别单元,识别单元与处理器和/或内存模组连接,识别单元被设置为在主机与处理器和/或内存模组建立连接的情况下,识别处理器和/或内存模组支持的协议类型。
在本实施例中,在控制芯片是CXL Switch Fabric架构的芯片时,控制单元可以是多个CXL Switch Fanout组成,CXL switch Fanout指的是在CXL Switch Fabric中,将一条CXL链路复制到多个另外的链路的过程。这个过程允许多个设备能够同时访问和共享同一个CXL链路上的设备,从而提高系统性能和可扩展性。控制芯片能够实现CXL switch Fanout是因为CXL支持点对点拓扑结构,可以允许多个设备通过一条CXL链路相互通信,在链路中转数据包。CXL switch Fanout的主要用途是提高系统中多个设备之间的连接效率和数据吞吐量。
管理单元可以是多个CXL Switch Root组成,CXL设备通过CXL端口与CXL Switch Fabric连接。CXL Switch Root维护着一个路由表,用于存储和管理网络中各个节点之间的路由信息。当CXL设备发送数据请求时,路由表选择最佳路径,将数据转发到目的CXL设备,例如,将HOST1中的数据传输至内存模组中存储。CXL数据传输板卡中的CXL端口的流量管理器将数据按顺序存储在缓存区中,以便快速处理和传输。CXL Switch Fabric根据输入端口和路由表的信息,在缓冲区中选择最佳路径,并将数据转发到目标CXL设备的端口。在转发数据的同时,CXL Switch Fabric还需要确保控制芯片内部通信速度的稳定和性能的一致性。当CXL设备需要返回数据时,CXL Switch Fabric同样会根据路由信息和缓存区的状态,找到最佳的路径,将数据传输到目标设备。
识别单元可以通过BIOS命令去识别端口连接的是PCLE设备还是CXL设备。例如,主机以IntelEagle Stream系列的CPU为例,CXL互连完全依赖于PCIe 5.0电气层,所有拓扑都依赖于PCIe。Eagle Stream CPU可以支持CXL4X16个通道。CXL和PCI Express设备可以同时工作在给定端口下,每个设备8lane。任何x16 PCI Express端口都可以连接到PCI Express设备或CXL设备。识别单元能够自动检测另一端是PCI Express卡还是CXL设备,并动态配置链接。在连接CXL设备的时候,应该在BIOS提前进行设置CXL模式,还可以通过BIOS命令去识别端口连接的是PCLE设备还是CXL设备。本实施例通过控制芯片可以控制链路拓扑和路由信息,以及识别设备的类型,从而 可以快速的对资源进行调配。
在一个示例性实施例中,主机包括M个,处理器包括P个,内存模组包括K个,其中,M、P以及K均是大于或等于1的自然数。
在本实施例中,在M大于1的情况下,M个主机共享K个内存模组。在N为16时,M大于或等于2且小于或等于7。在N为16,M为7,P为7时,处理器和主机一一对应,或者,一个主机对应多个处理器,其中,主机和处理器之间通过PCIe协议传输数据。在K大于1的情况下,一个主机对应多个内存模组,多个内存模组组成一个主机的内存资源池,其中,主机与内存模组之间通过CXL协议传输数据,主机的数量和内存模组的数量呈正比关系。例如,在控制芯片选用CXL 2.0switch芯片的情况下,控制芯片兼容PCIe5.0,支持不同的firmware(写入EPROM(可擦写可编程只读存储器)或EEPROM(电可擦可编程只读存储器)中的程序)切换,控制芯片共计16个端口port,每个port都可以配置成PCIe或者CXL,支持PCIe与CXL混用的场景。可以最多接7个HOST,最少接2个。在应用于高速计算的场景中时,上行设备可以连接7个HOST,下行设备可以连接7个GPU,每个HOST接一个GPU,可以多路并行处理数据。也可以接2个主机,每个主机出4路X16信号,HOST1接4个GPU被设置为加速处理数据,HOST2接4X16的内存模组被设置为读取存储数据;也可以上行连接4个主机,下行连接8X16的内存模组,走CXL协议。每个主机分出两组X16 port,接两个X16内存模组。
可选地,如图3所示,在控制芯片的switch层面,可以由多个VCXL组成,可以根据主机的指令要求,映射到指定的物理端口,最多可以支持分配16个逻辑存储空间。每个VCXL都有自己的ID,在主机的控制下用来绑定或者解绑VCXL到物理端口的映射,实现不同端口间的配置。
本实施例中的端口可以基于应用实际需求进行灵活配置,主机的数量和主机的端口资源信息也可以进行灵活配置。从而可以实现灵活调配资源的目的。
在一个示例性实施例中,CXL数据传输板卡还包括:连接器,连接器连接控制芯片和管理器,其中,管理器被设置为通过连接器管理控制芯片的运行。在本实施例中,连接器可以是电缆组件连接器MCIO,如图2所示,CXL Switch芯片通过连接器MCIO与外部的管理器CPU相连接。如图4所示,管理器CPU管理CXL Switch芯片的I2C、RESET、状态指示等。本实施例通过外部管理器对控制芯片进行管理,可以准确的控制控制芯片的工作状态。
在一个示例性实施例中,CXL数据传输板卡,还包括:时钟发生器,时钟发生器与控制芯片连接,被设置为产生启动控制芯片的第一初始化时钟信号,并被设置为在控制芯 片启动之后,产生协调控制芯片的时钟频率的第一时钟信号。在本实施例中,时钟发生器通过预设接口与控制芯片、处理器以及内存模组均连接,被设置为分别产生启动控制芯片、处理器以及内存模组的第二初始化时钟信号,并被设置为在控制芯片、处理器以及内存模组启动之后,产生协调控制芯片、处理器以及内存模组的时钟频率的第二时钟信号。如图4所示,CXL数据传输板卡支持RJ45接口,为BMC管理网口。每个CDFP接口配有一组红绿双色灯,按照时钟发生器产生的时钟信号,红色灯在数据传输正常时不亮、故障后常亮;绿色灯在无数据传输时常亮、正常数据传输时以1Hz的频率闪烁。本实施例通过时钟发生器可以准确的对数据的传输进行监测。
在一个示例性实施例中,CXL数据传输板卡还包括:调试设备,调试设备与控制芯片连接,被设置为对控制芯片的CXL性能进行调试。在本实施例中,调试设备调试的范围包括:CXL边际测试,与PCIe调试的方式相同;CXL链接速度、宽度、符合性测试;启用或禁用交错模式进行系统压力或性能或带宽测试。此外,还包括CXL合规性测试、CXL压力测试、CXL延迟测试等。此外,还包括链接和协议错误处理,利用PCIe ERR消息通过CXL.io、CXL.io协议错误被记录并显示出来。本实施例通过调试设备的调试可以保证数据传输的正确性。
本申请实施例中所提供的方法实施例可以在移动终端、计算机终端或者类似的运算装置中执行。以运行在移动终端上为例,图5是本申请实施例的一种控制数据传输的方法的移动终端的硬件结构框图。如图5所示,移动终端可以包括一个或多个(图5中仅示出一个)处理器502(处理器502可以包括但不限于微处理器MCU或可编程逻辑器件FPGA等的处理装置)和被设置为存储数据的存储器504,其中,上述移动终端还可以包括被设置为通信功能的传输设备506以及输入输出设备508。本领域普通技术人员可以理解,图5所示的结构仅为示意,其并不对上述移动终端的结构造成限定。例如,移动终端还可包括比图1中所示更多或者更少的组件,或者具有与图5所示不同的配置。
存储器504可被设置为存储计算机程序,例如,应用软件的软件程序以及模块,如本申请实施例中的控制数据传输的方法对应的计算机程序,处理器502通过运行存储在存储器504内的计算机程序,从而执行各种功能应用以及数据处理,即实现上述的方法。存储器504可包括高速随机存储器,还可包括非易失性存储器,如一个或者多个磁性存储装置、闪存、或者其他非易失性固态存储器。在一些实例中,存储器504可包括相对于处理器502远程设置的存储器,这些远程存储器可以通过网络连接至移动终端。上述网络的实例包括但不限于互联网、企业内部网、局域网、移动通信网及其组合。
传输设备506被设置为经由一个网络接收或者发送数据。上述的网络实例可包括移 动终端的通信供应商提供的无线网络。在一个实例中,传输设备506包括一个网络适配器(Network Interface Controller,简称为NIC),其可通过基站与其他网络设备相连从而可与互联网进行通讯。在一个实例中,传输设备506可以为射频(Radio Frequency,简称为RF)模块,其被设置为通过无线方式与互联网进行通讯。
在本实施例中提供了一种控制数据传输的方法,图6是根据本申请实施例的控制数据传输的方法的流程图,如图6所示,该流程包括如下步骤:
步骤S602,通过上行端口接收主机发送的数据请求,其中,数据请求中包括待传输的数据,主机是支持PCIe协议的设备或支持CXL协议的设备;
步骤S604,响应数据请求,从路由表中查找主机的路由信息;
步骤S606,按照路由信息将数据传输至与上行端口对应的下行端口连接的处理器和/或内存模组中,其中,下行端口与上行端口对应连接,处理器和内存模组均是支持PCIe协议的设备或支持CXL协议的设备。
其中,上述步骤的执行主体可以为控制芯片中设置的处理器,或者与终端或者服务器相对独立设置的处理器或者处理设备等,但不限于此。
本实施例可以应用于需要高速数据传输和大规模数据处理的场景中,例如数据中心,高性能计算,人工智能云计算等场景。
在本实施例中,控制芯片是支持CXL协议的芯片,包括高速互连架构CXL Switch Fabric。CXL Switch Fabric可以连接多个CXL设备,形成一个共享资源的计算平台。CXL Switch Fabric给CXL设备提供了一个可扩展、灵活的基础架构,可以通过在CXL设备之间提供高速的点对点连接来实现设备之间的通信和数据传输。CXL设备可以是主机也可以是处理器、内存模组等设备。
可选地,CXL Switch Fabric架构分为三层:事务层(Transaction Layer)、链路层(Link Layer)和物理层(Physical Layer)。为了同时兼容PCIe协议和CXL协议,CXL Switch Fabric架构中提供了Flex Bus端口,使设备可以在PCIe设备和CXL设备中进行选择。CXL.cache和CXL.mem协议被组合在一起共享一个公共的事务层和链路层,而CXL.io拥有自己的事务层和链路层。CXL链路层与CXL ARB或MUX交互,从而交织来自两个逻辑流的流量。物理层中还包含两个子层,分别是逻辑子层和电气子层。其中,逻辑子层可以在PCIe模式和CXL模式之间切换,而电气子层则遵循PCIe规范。在事务层中,CXL.io为I/O设备提供了一个非一致性的加载或存储接口。
可选地,处理器可以是CPU,也可以是GPU,CPU与GPU之间可以绕过PCIe协议,用CXL协议来共享,互取对方的内存资源。透过CXL协议,CPU与GPU之间形同连成单 一个庞大的堆栈内存池,本实施例中的,GPU BOX既可以支持PCIE协议,又可以支持CXL协议。如图2所示,CXL数据传输板卡可以是CXL Switch板,控制芯片可以是CXL Switch芯片,处理器可以是GPU,上行设备可以连接多个主机HOST,下行设备可以连接GPU BOX,也可以接内存模组,配置灵活。且每个主机均可以和不同的下行设备连接,通过改变CXL Switch芯片的烧录信息来改变主机HOST和下行设备的连接关系。连接的上行设备和下行设备的数量可以由CXL Switch芯片的功能决定。
可选地,在下行设备为GPU BOX时,GPU BOX内部支持16个GPU插槽,与CXL Switch芯片通过CDFP线缆进行互联,板载CPLD被设置为控制BOX的上电时序,板载总线主控制器BMC(Bus Master Controller)被设置为管理GPU的工作状态。
可选地,内存模组是基于CXL总线协议的设备,可以实现内存远端拓展,多主机共享内存,内存板的CDFP接口通过CDFP线缆与CXL Switch板卡互联,板载BMC和CPLD负责CXL Switch板卡d的管理,包括散热控制、上下电、状态指示等。
通过上述步骤,CXL数据传输板卡中的控制芯片是支持CXL协议的芯片,并部署了上行端口和下行端口;上行端口与主机连接,其中,主机是支持PCIe协议的设备或支持CXL协议的设备;下行端口与处理器和/或内存模组连接,下行端口与上行端口对应连接,使得主机即可以接处理器又可以接内存模组,从而可以根据实际需要灵活配置处理器和内存模组的数量。因此,可以解决相关技术中存在的无法根据不同的工作负载进行资源的重新调配的问题,达到实现灵活调配资源的效果。
在一个示例性实施例中,上行端口和下行端口均为N个,N是大于1的自然数。上行端口和下行端口均为N个,N是大于1的自然数。N的取值由控制芯片的功能确定,例如,N可以是16,也可以是4等,如图2所示,CXL Switch芯片中设置6个上下行端口。本实施例通过灵活设备N的取值,可以灵活的设置连接的上下行设备,实现对资源的灵活调配。
在一个示例性实施例中,响应数据请求,从路由表中查找主机的路由信息之前,方法还包括:设置N个上行端口和N个下行端口之间的CXL链路的拓扑,其中,CXL链路被设置为连接对应的上行端口和下行端口;基于CXL链路的拓扑确定主机与处理器和/或内存模组之间的路由,得到路由表。其中,按照路由信息将数据传输至与上行端口对应的下行端口连接的处理器和/或内存模组中,包括:通过下行端口将数据缓存至缓存单元中;从路由信息中选择目标路由;按照目标路由将缓存单元中的数据传输至下行端口连接的处理器和/或内存模组中。
在本实施例中,在控制芯片是CXL Switch Fabric架构的芯片时,控制单元可以是 多个CXL Switch Fanout组成,CXL switch Fanout指的是在CXL Switch Fabric中,将一条CXL链路复制到多个另外的链路的过程。这个过程允许多个设备能够同时访问和共享同一个CXL链路上的设备,从而提高系统性能和可扩展性。控制芯片能够实现CXL switch Fanout是因为CXL支持点对点拓扑结构,可以允许多个设备通过一条CXL链路相互通信,在链路中转数据包。CXL switch Fanout的主要用途是提高系统中多个设备之间的连接效率和数据吞吐量。CXL设备通过CXL端口与CXL Switch Fabric连接。CXL Switch Root维护着一个路由表,用于存储和管理网络中各个节点之间的路由信息。当CXL设备发送数据请求时,路由表选择最佳路径,将数据转发到目的CXL设备,例如,将HOST1中的数据传输至内存模组中存储。CXL数据传输板卡中的CXL端口的流量管理器将数据按顺序存储在缓存区中,以便快速处理和传输。CXL Switch Fabric根据输入端口和路由表的信息,在缓冲区中选择最佳路径,并将数据转发到目标CXL设备的端口。在转发数据的同时,CXL Switch Fabric还需要确保控制芯片内部通信速度的稳定和性能的一致性。当CXL设备需要返回数据时,CXL Switch Fabric同样会根据路由信息和缓存区的状态,找到最佳的路径,将数据传输到目标设备。
在一个示例性实施例中,按照路由信息将数据传输至与上行端口对应的下行端口连接的处理器和/或内存模组中之前,方法还包括:在主机与处理器和/或内存模组建立连接的情况下,识别处理器和/或内存模组支持的协议类型,以按照处理器和/或内存模组支持的协议类型将数据传输至处理器和/或内存模组支持的协议类型。在本实施例中,通过BIOS命令去识别端口连接的是PCLE设备还是CXL设备。例如,主机以IntelEagle Stream系列的CPU为例,CXL互连完全依赖于PCIe 5.0电气层,所有拓扑都依赖于PCIe。Eagle Stream CPU可以支持CXL4X16个通道。CXL和PCI Express设备可以同时工作在给定端口下,每个设备8lane。任何x16 PCI Express端口都可以连接到PCI Express设备或CXL设备。识别单元能够自动检测另一端是PCI Express卡还是CXL设备,并动态配置链接。在连接CXL设备的时候,应该在BIOS提前进行设置CXL模式,还可以通过BIOS命令去识别端口连接的是PCLE设备还是CXL设备。本实施例通过控制芯片可以控制链路拓扑和路由信息,以及识别设备的类型,从而可以快速的对资源进行调配。
在一个示例性实施例中,主机包括M个,处理器包括P个,内存模组包括K个,其中,M、P以及K均是大于或等于1的自然数。在M大于1的情况下,M个主机共享K个内存模组。在N为16的情况下,M大于或等于2且小于或等于7。在N为16的情况下,M为7,P为7,处理器和主机一一对应,或者,一个主机对应多个处理器,其中,主机和 处理器之间通过PCIe协议传输数据。
例如,在控制芯片选用CXL 2.0switch芯片的情况下,控制芯片兼容PCIe 5.0,支持不同的firmware(写入EPROM(可擦写可编程只读存储器)或EEPROM(电可擦可编程只读存储器)中的程序)切换,控制芯片共计16个端口port,每个port都可以配置成PCIe或者CXL,支持PCIe与CXL混用的场景。可以最多接7个HOST,最少接2个。在应用于高速计算的场景中时,上行设备可以连接7个HOST,下行设备可以连接7个GPU,每个HOST接一个GPU,可以多路并行处理数据。也可以接2个主机,每个主机出4路X16信号,HOST1接4个GPU被设置为加速处理数据,HOST2接4X16的内存模组被设置为读取存储数据;也可以上行连接4个主机,下行连接8X16的内存模组,走CXL协议。每个主机分出两组X16 port,接两个X16内存模组。
可选地,如图3所示,在控制芯片的switch层面,可以由多个VCXL组成,可以根据主机的指令要求,映射到指定的物理端口,最多可以支持分配16个逻辑存储空间。每个VCXL都有自己的ID,在主机的控制下用来绑定或者解绑VCXL到物理端口的映射,实现不同端口间的配置。
本实施例中的端口可以基于应用实际需求进行灵活配置,主机的数量和主机的端口资源信息也可以进行灵活配置。从而可以实现灵活调配资源的目的。
在一个示例性实施例中,在K大于1的情况下,一个主机对应多个内存模组,多个内存模组组成一个主机的内存资源池,其中,主机与内存模组之间通过CXL协议传输数据,主机的数量和内存模组的数量呈正比关系。
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到根据上述实施例的方法可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件,但很多情况下前者是更佳的实施方式。基于这样的理解,本申请的技术方案本质上或者说对相关技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个非易失性可读存储介质(如ROM/RAM、磁碟、光盘)中,包括若干指令用以使得一台终端设备(可以是手机,计算机,服务器,或者网络设备等)执行本申请各个实施例的方法。
在本实施例中还提供了一种链路交换的控制系统,链路交换的控制系统包括上述的CXL数据传输板卡该装置被设置为实现上述实施例及可选实施方式,已经进行过说明的不再赘述。
本申请的实施例还提供了一种非易失性可读存储介质,该非易失性可读存储介质中存储有计算机程序,其中,该计算机程序被设置为运行时执行上述任一项方法实施例中 的步骤。
在一个示例性实施例中,上述非易失性可读存储介质可以包括但不限于:U盘、只读存储器(Read-Only Memory,简称为ROM)、随机存取存储器(Random Access Memory,简称为RAM)、移动硬盘、磁碟或者光盘等各种可以存储计算机程序的非易失性可读存储介质。
本申请的实施例还提供了一种电子设备,包括存储器和处理器,该存储器中存储有计算机程序,该处理器被设置为运行计算机程序以执行上述任一项方法实施例中的步骤。
本申请的实施例还提供了一种电子设备,图7是根据本申请实施例的电子设备的示意图,如图7所示,包括存储器和处理器,该存储器中存储有计算机程序,该处理器被设置为运行计算机程序以执行上述任一项方法实施例中的步骤。
在一个示例性实施例中,上述电子设备还可以包括传输设备以及输入输出设备,其中,该传输设备和上述处理器连接,该输入输出设备和上述处理器连接。
本实施例中的示例可以参考上述实施例及示例性实施方式中所描述的示例,本实施例在此不再赘述。
显然,本领域的技术人员应该明白,上述的本申请的各模块或各步骤可以用通用的计算装置来实现,它们可以集中在单个的计算装置上,或者分布在多个计算装置所组成的网络上,它们可以用计算装置可执行的程序代码来实现,从而,可以将它们存储在存储装置中由计算装置来执行,并且在某些情况下,可以以不同于此处的顺序执行所示出或描述的步骤,或者将它们分别制作成各个集成电路模块,或者将它们中的多个模块或步骤制作成单个集成电路模块来实现。这样,本申请不限制于任何特定的硬件和软件结合。
以上仅为本申请的可选实施例而已,并不用于限制本申请,对于本领域的技术人员来说,本申请可以有各种更改和变化。凡在本申请的原则之内,所作的任何修改、等同替换、改进等,均应包含在本申请的保护范围之内。

Claims (29)

  1. 一种CXL数据传输板卡,其特征在于,包括:
    控制芯片,所述控制芯片是支持开放式互连标准CXL协议的芯片,所述控制芯片上部署了上行端口和下行端口;
    所述上行端口与主机连接,其中,所述主机是支持高速串行计算机扩展总线标准PCIe协议的设备或支持所述CXL协议的设备;
    所述下行端口与处理器和/或内存模组连接,所述下行端口与所述上行端口对应连接,所述主机被设置为通过所述上行端口将数据传输至对应的所述下行端口,以通过所述下行端口将所述数据传输至所述处理器和/或所述内存模组,其中,所述处理器和所述内存模组均是支持所述PCIe协议的设备或支持所述CXL协议的设备。
  2. 根据权利要求1所述的CXL数据传输板卡,其特征在于,所述上行端口和所述下行端口均为N个,所述N是大于1的自然数。
  3. 根据权利要求2所述的CXL数据传输板卡,其特征在于,所述控制芯片包括:
    控制单元,所述控制单元被设置为控制所述上行端口和所述下行端口之间的CXL链路的拓扑,其中,所述CXL链路被设置为连接对应的所述上行端口和所述下行端口。
  4. 根据权利要求2所述的CXL数据传输板卡,其特征在于,所述控制芯片包括:
    管理单元,所述管理单元被设置为管理所述主机与所述处理器和/或所述内存模组之间的路由信息。
  5. 根据权利要求2-4任一项所述的CXL数据传输板卡,其特征在于,所述控制芯片包括:
    识别单元,所述识别单元与所述处理器和/或所述内存模组连接,所述识别单元被设置为在所述主机与所述处理器和/或所述内存模组建立连接的情况下,识别所述处理器和/或所述内存模组支持的协议类型。
  6. 根据权利要求2所述的CXL数据传输板卡,其特征在于,所述主机包括M个,所述处理器包括P个,所述内存模组包括K个,其中,所述M、所述P以及所述K均是大于或等于1的自然数。
  7. 根据权利要求6所述的CXL数据传输板卡,其特征在于,在所述M大于1的情况下,M个所述主机共享K个所述内存模组。
  8. 根据权利要求6所述的CXL数据传输板卡,其特征在于,所述N为16,所述M大于或等于2且小于或等于7。
  9. 根据权利要求6所述的CXL数据传输板卡,其特征在于,所述N为16,所述M为 7,所述P为7,所述处理器和所述主机一一对应,或者,一个所述主机对应多个所述处理器,其中,所述主机和所述处理器之间通过所述PCIe协议传输数据。
  10. 根据权利要求6所述的CXL数据传输板卡,其特征在于,在所述K大于1的情况下,一个所述主机对应多个所述内存模组,多个所述内存模组组成一个所述主机的内存资源池,其中,所述主机与所述内存模组之间通过所述CXL协议传输数据,所述主机的数量和所述内存模组的数量呈正比关系。
  11. 根据权利要求1所述的CXL数据传输板卡,其特征在于,还包括:
    连接器,所述连接器连接所述控制芯片和管理器,其中,所述管理器被设置为通过所述连接器管理所述控制芯片的运行。
  12. 根据权利要求1所述的CXL数据传输板卡,其特征在于,还包括:
    时钟发生器,所述时钟发生器与所述控制芯片连接,被设置为产生启动所述控制芯片的第一初始化时钟信号,并被设置为在所述控制芯片启动之后,产生协调所述控制芯片的时钟频率的第一时钟信号。
  13. 根据权利要求12所述的CXL数据传输板卡,其特征在于,所述时钟发生器通过预设接口与所述控制芯片、所述处理器以及所述内存模组均连接,被设置为分别产生启动所述控制芯片、所述处理器以及所述内存模组的第二初始化时钟信号,并被设置为在所述控制芯片、所述处理器以及所述内存模组启动之后,产生协调所述控制芯片、所述处理器以及所述内存模组的时钟频率的第二时钟信号。
  14. 根据权利要求12所述的CXL数据传输板卡,其特征在于,所述CXL数据传输板卡还包括:
    调试设备,所述调试设备与所述控制芯片连接,被设置为对所述控制芯片的CXL性能进行调试。
  15. 根据权利要求1所述的CXL数据传输板卡,其特征在于,还包括:
    所述控制芯片的SWITCH层面包括多个虚拟CXL,所述虚拟CXL用于根据所述主机的指令,映射到所述控制芯片中指定的物理端口,以进行不同的上行端口和下行端口之间的配置,其中,不同的所述上行端口和所述下行端口之间的配置对应不同数量的处理器和内存模组的配置。
  16. 一种控制数据传输的方法,其特征在于,包括:
    通过上行端口接收主机发送的数据请求,其中,所述数据请求中包括待传输的数据,所述主机是支持PCIe协议的设备或支持CXL协议的设备;
    响应所述数据请求,从路由表中查找所述主机的路由信息;
    按照所述路由信息将所述数据传输至与所述上行端口对应的下行端口连接的处理器和/或内存模组中,其中,所述下行端口与所述上行端口对应连接,所述处理器和所述内存模组均是支持所述PCIe协议的设备或支持所述CXL协议的设备。
  17. 根据权利要求16所述的方法,其特征在于,所述上行端口和所述下行端口均为N个,所述N是大于1的自然数。
  18. 根据权利要求16所述的方法,其特征在于,响应所述数据请求,从路由表中查找所述主机的路由信息之前,所述方法还包括:
    设置N个所述上行端口和N个所述下行端口之间的CXL链路的拓扑,其中,所述CXL链路用于连接对应的所述上行端口和所述下行端口;
    基于所述CXL链路的拓扑确定所述主机与所述处理器和/或所述内存模组之间的路由,得到所述路由表。
  19. 根据权利要求16所述的方法,其特征在于,按照所述路由信息将所述数据传输至与所述上行端口对应的下行端口连接的处理器和/或内存模组中,包括:
    通过所述下行端口将所述数据缓存至缓存单元中;
    从所述路由信息中选择目标路由;
    按照所述目标路由将所述缓存单元中的所述数据传输至所述下行端口连接的所述处理器和/或所述内存模组中。
  20. 根据权利要求16所述的方法,其特征在于,按照所述路由信息将所述数据传输至与所述上行端口对应的下行端口连接的处理器和/或内存模组中之前,所述方法还包括:
    在所述主机与所述处理器和/或所述内存模组建立连接的情况下,识别所述处理器和/或所述内存模组支持的协议类型,以按照所述处理器和/或所述内存模组支持的协议类型将所述数据传输至所述处理器和/或所述内存模组支持的协议类型。
  21. 根据权利要求16所述的方法,其特征在于,所述主机包括M个,所述处理器包括P个,所述内存模组包括K个,其中,所述M、所述P以及所述K均是大于或等于1的自然数。
  22. 根据权利要求21所述的方法,其特征在于,在所述M大于1的情况下,M个所述主机共享K个所述内存模组。
  23. 根据权利要求21所述的方法,其特征在于,N为16,所述M大于或等于2且小于或等于7,其中,所述N是所述上行端口和所述下行端口的数量,所述N是大于1的自然数。
  24. 根据权利要求21所述的方法,其特征在于,N为16,所述M为7,所述P为7,所述处理器和所述主机一一对应,或者,一个所述主机对应多个所述处理器,其中,所述主机和所述处理器之间通过所述PCIe协议传输数据,所述N是所述上行端口和所述下行端口的数量,所述N是大于1的自然数。
  25. 根据权利要求21所述的方法,其特征在于,在所述K大于1的情况下,一个所述主机对应多个所述内存模组,多个所述内存模组组成一个所述主机的内存资源池,其中,所述主机与所述内存模组之间通过所述CXL协议传输数据,所述主机的数量和所述内存模组的数量呈正比关系。
  26. 根据权利要求16所述的方法,其特征在于,控制芯片的SWITCH层面包括多个虚拟CXL,所述虚拟CXL用于根据所述主机的指令,映射到所述控制芯片中指定的物理端口,以进行不同的上行端口和下行端口之间的配置,其中,不同的所述上行端口和所述下行端口之间的配置对应不同数量的处理器和内存模组的配置,所述控制芯片是支持开放式互连标准CXL协议的芯片,所述控制芯片上部署了上行端口和下行端口。
  27. 一种链路交换的控制系统,其特征在于,所述链路交换的控制系统包括权利要求1至15任一项所述的CXL数据传输板卡。
  28. 一种非易失性可读存储介质,其特征在于,所述非易失性可读存储介质中存储有计算机程序,其中,所述计算机程序被处理器执行时实现所述权利要求16至26任一项中所述的方法的步骤。
  29. 一种电子设备,包括存储器、处理器以及存储在所述存储器上并可在所述处理器上运行的计算机程序,其特征在于,所述处理器执行所述计算机程序时实现所述权利要求16至26任一项中所述的方法的步骤。
PCT/CN2024/083114 2023-06-28 2024-03-21 Cxl数据传输板卡及控制数据传输的方法 Ceased WO2025001344A1 (zh)

Priority Applications (1)

Application Number Priority Date Filing Date Title
US19/122,952 US12547579B2 (en) 2023-06-28 2024-03-21 Board for CXL data transmission, method for data transmission control and device

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202310776401.5A CN116501681B (zh) 2023-06-28 2023-06-28 Cxl数据传输板卡及控制数据传输的方法
CN202310776401.5 2023-06-28

Publications (1)

Publication Number Publication Date
WO2025001344A1 true WO2025001344A1 (zh) 2025-01-02

Family

ID=87330559

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/083114 Ceased WO2025001344A1 (zh) 2023-06-28 2024-03-21 Cxl数据传输板卡及控制数据传输的方法

Country Status (3)

Country Link
US (1) US12547579B2 (zh)
CN (1) CN116501681B (zh)
WO (1) WO2025001344A1 (zh)

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN119537294A (zh) * 2025-01-22 2025-02-28 苏州元脑智能科技有限公司 一种用于加速计算的控制器和加速计算系统
CN120075324A (zh) * 2025-04-23 2025-05-30 上海芯力基半导体有限公司 支持混用PCIe与CXL协议的装置、交换机及方法
CN120406701A (zh) * 2025-06-30 2025-08-01 苏州元脑智能科技有限公司 资源池及其复位控制方法、主系统的时序控制方法
CN120578624A (zh) * 2025-07-31 2025-09-02 苏州元脑智能科技有限公司 电子设备、数据处理方法、设备、介质及程序产品

Families Citing this family (13)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116501681B (zh) 2023-06-28 2023-09-29 苏州浪潮智能科技有限公司 Cxl数据传输板卡及控制数据传输的方法
CN116701287B (zh) * 2023-08-09 2023-12-08 西安甘鑫科技股份有限公司 一种基于pcie的多设备兼容设备拓展方法
CN116781511B (zh) * 2023-08-22 2023-11-03 苏州浪潮智能科技有限公司 主机系统的配置方法及设备、装置、计算系统、存储介质
CN116886644B (zh) * 2023-09-06 2024-01-26 苏州浪潮智能科技有限公司 交换芯片、内存扩展模组和内存扩展系统
CN119717163A (zh) * 2023-09-28 2025-03-28 上海曦智科技有限公司 光模块、光电集成半导体结构以及光交换系统
CN117033001B (zh) * 2023-10-09 2024-02-20 苏州元脑智能科技有限公司 服务器系统、配置方法、cpu、控制模组与存储介质
CN117555768B (zh) * 2023-11-23 2025-12-02 浪潮(北京)电子信息产业有限公司 计算快速链路设备的测试方法、装置、系统及设备和介质
CN117938849B (zh) * 2023-12-11 2025-09-16 超聚变数字技术有限公司 传输通道管理方法、数据传输方法、管理设备及计算设备
CN118018481A (zh) * 2023-12-26 2024-05-10 超聚变数字技术有限公司 数据传输方法、设备及系统
CN117493026B (zh) * 2023-12-29 2024-03-22 苏州元脑智能科技有限公司 一种多主机与多计算快速链接内存设备系统及其应用设备
CN117493237B (zh) * 2023-12-29 2024-04-09 苏州元脑智能科技有限公司 计算设备、服务器、数据处理方法和存储介质
CN119544640A (zh) 2024-03-22 2025-02-28 北京字跳网络技术有限公司 用于转发数据的设备、方法、装置和存储介质
CN121357134A (zh) * 2024-07-15 2026-01-16 云智能资产控股(新加坡)私人股份有限公司 内存扩展系统、方法、交换机和内存池

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113742256A (zh) * 2020-05-28 2021-12-03 三星电子株式会社 用于可扩展且一致性存储器装置的系统和方法
US20230029026A1 (en) * 2022-09-30 2023-01-26 Intel Corporation Flexible resource sharing in a network
CN116126742A (zh) * 2023-01-30 2023-05-16 苏州浪潮智能科技有限公司 内存访问方法、装置、服务器及存储介质
CN116166434A (zh) * 2023-02-28 2023-05-26 苏州浪潮智能科技有限公司 处理器分配方法及系统、装置、存储介质、电子设备
CN116501681A (zh) * 2023-06-28 2023-07-28 苏州浪潮智能科技有限公司 Cxl数据传输板卡及控制数据传输的方法

Family Cites Families (12)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8923172B2 (en) * 2009-08-24 2014-12-30 Qualcomm Incorporated Deterministic backoff channel access
EP3667974A1 (en) * 2015-07-29 2020-06-17 Huawei Technologies Co., Ltd. Feedback information sending apparatus and method, and feedback information receiving apparatus and method
CN108616341B (zh) * 2016-12-13 2020-05-26 电信科学技术研究院 一种数据传输方法、基站及终端
CN111788786A (zh) * 2019-01-11 2020-10-16 Oppo广东移动通信有限公司 用于传输反馈信息的方法、终端设备和网络设备
US11461263B2 (en) * 2020-04-06 2022-10-04 Samsung Electronics Co., Ltd. Disaggregated memory server
CN112597094B (zh) * 2020-12-24 2024-05-31 联想长风科技(北京)有限公司 一种提高rdma传输效率的装置及方法
CN215769533U (zh) * 2021-07-26 2022-02-08 联想长风科技(北京)有限公司 一种基于cxl加速计算的板卡
US11601377B1 (en) * 2021-10-11 2023-03-07 Cisco Technology, Inc. Unlocking computing resources for decomposable data centers
CN115858146A (zh) * 2022-11-09 2023-03-28 阿里巴巴(中国)有限公司 内存扩展系统和计算节点
CN115904714B (zh) * 2022-11-18 2025-07-08 苏州浪潮智能科技有限公司 一种内存分配系统和服务器
CN115982078A (zh) * 2023-01-19 2023-04-18 北京超弦存储器研究院 一种cxl内存模组及内存存储系统
CN115934366A (zh) * 2023-03-15 2023-04-07 浪潮电子信息产业股份有限公司 服务器存储扩展方法、装置、设备、介质及整机柜系统

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113742256A (zh) * 2020-05-28 2021-12-03 三星电子株式会社 用于可扩展且一致性存储器装置的系统和方法
US20230029026A1 (en) * 2022-09-30 2023-01-26 Intel Corporation Flexible resource sharing in a network
CN116126742A (zh) * 2023-01-30 2023-05-16 苏州浪潮智能科技有限公司 内存访问方法、装置、服务器及存储介质
CN116166434A (zh) * 2023-02-28 2023-05-26 苏州浪潮智能科技有限公司 处理器分配方法及系统、装置、存储介质、电子设备
CN116501681A (zh) * 2023-06-28 2023-07-28 苏州浪潮智能科技有限公司 Cxl数据传输板卡及控制数据传输的方法

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN119537294A (zh) * 2025-01-22 2025-02-28 苏州元脑智能科技有限公司 一种用于加速计算的控制器和加速计算系统
CN120075324A (zh) * 2025-04-23 2025-05-30 上海芯力基半导体有限公司 支持混用PCIe与CXL协议的装置、交换机及方法
CN120075324B (zh) * 2025-04-23 2025-08-19 上海芯力基半导体有限公司 支持混用PCIe与CXL协议的装置、交换机及方法
CN120406701A (zh) * 2025-06-30 2025-08-01 苏州元脑智能科技有限公司 资源池及其复位控制方法、主系统的时序控制方法
CN120578624A (zh) * 2025-07-31 2025-09-02 苏州元脑智能科技有限公司 电子设备、数据处理方法、设备、介质及程序产品

Also Published As

Publication number Publication date
CN116501681B (zh) 2023-09-29
US20260010510A1 (en) 2026-01-08
CN116501681A (zh) 2023-07-28
US12547579B2 (en) 2026-02-10

Similar Documents

Publication Publication Date Title
WO2025001344A1 (zh) Cxl数据传输板卡及控制数据传输的方法
CN110809760B (zh) 资源池的管理方法、装置、资源池控制单元和通信设备
US11080221B2 (en) Switching device, peripheral component interconnect express system, and method for initializing peripheral component interconnect express system
CN110941576B (zh) 具有多模pcie功能的存储控制器的系统、方法和设备
WO2025227986A1 (zh) 服务器系统、服务器系统的资源调度方法、芯片及芯粒
US8904079B2 (en) Tunneling platform management messages through inter-processor interconnects
TW202145025A (zh) 用於管理記憶體資源的系統以及實行遠端直接記憶體存取的方法
CN114546913B (zh) 一种基于pcie接口的多主机之间数据高速交互的方法和装置
US9026687B1 (en) Host based enumeration and configuration for computer expansion bus controllers
CN119201469B (zh) 一种计算系统、方法、设备、介质及程序产品
US20130042019A1 (en) Multi-Server Consolidated Input/Output (IO) Device
CN110362515A (zh) 驱动器至驱动器存储系统、存储驱动器和存储数据的方法
CN106575283B (zh) 使用元胞自动机的群集服务器配置
US12596657B2 (en) Network instantiated peripheral devices
WO2025138695A1 (zh) 一种计算设备、管理控制器及数据处理方法
CN103927233A (zh) 多节点内存互联装置及一种大规模计算机集群
WO2025242114A1 (zh) Cxl交换板卡、cxl内存分配系统、分配方法及装置
US20250209013A1 (en) Dynamic server rebalancing
CN120950441A (zh) 一种PCIe拓扑切换方法、装置、设备及可读存储介质
CN117971135B (zh) 存储设备的访问方法、装置、存储介质和电子设备
CN119025460A (zh) 一种计算高速互联链路cxl设备的适配方法
WO2023186143A1 (zh) 一种数据处理方法、主机及相关设备
WO2023177982A1 (en) Dynamic server rebalancing
CN113709066B (zh) 一种PCIe通信装置及BMC
CN116208557A (zh) Bmc的i2c链路均衡分配方法、系统、终端及存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24829996

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE