EP1813039A2 - Method and apparatus for aligning data in a wide, high-speed, source synchronous parallel link - Google Patents

Method and apparatus for aligning data in a wide, high-speed, source synchronous parallel link

Info

Publication number
EP1813039A2
EP1813039A2 EP05851478A EP05851478A EP1813039A2 EP 1813039 A2 EP1813039 A2 EP 1813039A2 EP 05851478 A EP05851478 A EP 05851478A EP 05851478 A EP05851478 A EP 05851478A EP 1813039 A2 EP1813039 A2 EP 1813039A2
Authority
EP
European Patent Office
Prior art keywords
clock
data
group
signal
groups
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
EP05851478A
Other languages
German (de)
French (fr)
Other versions
EP1813039A4 (en
EP1813039B1 (en
Inventor
Dipankar Bhattacharya
Bangalore Priyadarshan
Jaushin Lee
Francois Gautier-Le Boulch
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Cisco Technology Inc
Original Assignee
Cisco Technology Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Cisco Technology Inc filed Critical Cisco Technology Inc
Publication of EP1813039A2 publication Critical patent/EP1813039A2/en
Publication of EP1813039A4 publication Critical patent/EP1813039A4/en
Application granted granted Critical
Publication of EP1813039B1 publication Critical patent/EP1813039B1/en
Anticipated expiration legal-status Critical
Expired - Lifetime legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F13/00Interconnection of, or transfer of information or other signals between, memories, input/output devices or central processing units
    • G06F13/38Information transfer, e.g. on bus
    • G06F13/42Bus transfer protocol, e.g. handshake; Synchronisation
    • G06F13/4282Bus transfer protocol, e.g. handshake; Synchronisation on a serial bus, e.g. I2C bus, SPI bus
    • G06F13/4291Bus transfer protocol, e.g. handshake; Synchronisation on a serial bus, e.g. I2C bus, SPI bus using a clocked protocol
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04JMULTIPLEX COMMUNICATION
    • H04J3/00Time-division multiplex systems
    • H04J3/02Details
    • H04J3/06Synchronising arrangements
    • H04J3/062Synchronisation of signals having the same nominal but fluctuating bit rates, e.g. using buffers
    • H04J3/0632Synchronisation of packets and cells, e.g. transmission of voice via a packet network, circuit emulation service [CES]
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L25/00Baseband systems
    • H04L25/02Details ; arrangements for supplying electrical power along data transmission lines
    • H04L25/14Channel dividing arrangements, i.e. in which a single bit stream is divided between several baseband channels and reassembled at the receiver
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L7/00Arrangements for synchronising receiver with transmitter
    • H04L7/0008Synchronisation information channels, e.g. clock distribution lines
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L7/00Arrangements for synchronising receiver with transmitter
    • H04L7/04Speed or phase control by synchronisation signals
    • H04L7/041Speed or phase control by synchronisation signals using special codes as synchronising signal

Definitions

  • the source-synchronous bus has been used to increase the speed of buses in many designs. Data and clock are sourced from the same device on the bus. The receiving device uses the clock from the bus to sample the data on the bus. Since the clock and data are driven and distributed similarly, they have similar delays and hence such buses can be run faster than buses using other clocking schemes.
  • DDR double-data rate
  • One embodiment of the invention allows the clock period of the transmit and receive core clocks to get smaller than the skew between the copies of source- synchronous clock.
  • the maximum frequency of operation of the link is thereby increased to the limit reachable for sampling a small number of data pins with a source-synchronous clock received on a pair (clock high and clock low) of clock pins. There is no limit imposed because of skew between multiple copies of source-synchronous clocks.
  • both the transmitting and receiving devices use a PLL (phase-locked loop) to phase-align their internal core-clocks with a common external reference clock. This limits the jitter and wander of the source- synchronous clock with respect to the receiver core-clock and that in turn reduces the depth of the receive-data FIFOs.
  • PLL phase-locked loop
  • the transmitting device may send data in a single clock from one or more logical buses in its core-clock domain over multiple source synchronous links.
  • the receive-data FIFOs and the deskew protocol align the data from the logical bus(es) in the core-clock domain of the receiving device.
  • FIG. 1 is a functional timing diagram of an embodiment of the invention
  • FIG. 2 depicts the transmitter interface model of an embodiment of the invention
  • FIG. 3 depicts the receiver interface model of an embodiment of the invention
  • FIG. 4 depicts the clock distributions and clock domains of an embodiment of the invention
  • FIG. 5 depicts an embodiment of the invention having a link using two clock-groups
  • FIG. 6 depicts a receiving device model for multiple clock copies of an embodiment of the invention
  • Fig. 7 depicts an embodiment of the invention having deskew logic in the receive interface model
  • FIG. 8 depicts a FIFO of an embodiment of the invention
  • Fig. 9 depicts two transmitters coupled in parallel to a receiver
  • Fig. 10 depicts a transmitter and receiver coupled by independent logical buses divided into clock groups.
  • SSPL source-synchronous parallel link
  • the embodiments include features such as multiple clock groups where multiple clock copies are transmitted and a limited number of data pins are associated with each clock signal, a deskew unit that aligns data sampled from different clock groups to a core clock, a clock generation system for forming clean copies of core clocks that have low jitter and noise, etc.
  • FIG. 23 is a functional timing diagram of the SSPL of this embodiment where the set of bits transferred in RO (rising edge of DcIk) and FO (falling edge of DcIk) form a data word.
  • DcIkJI data clock high
  • DcIk 2 L data clock low
  • Da[N:0] is the group of data pins sampled by using clock pair DcIk 3 H. and DcIk 3 L.
  • data can be transferred on rising and falling edge of a single clock, on the rising edges of two complementary clocks, or on the falling edges of two complementary clocks.
  • the group of N+l data pins referred to as Di[N:0] is associated with a pair of clock pins, DclkjH & DclkjL.
  • the letters a, b, and so on are used to replace the subscript "i", e.g., D a [N:0] associated with DcIk 3 H & DcIk 3 L, D b [N: 0] associated with Dclk b H & Dclk b L and so on.
  • the SSPL uses source-synchronous clocking, Le. data-pins and clock-pins are driven from the same source device to minimize skew between the data-pins and clock-pins.
  • Each data-pin of the SSPL transfers 2 bits of information in one cycle-time.
  • the maximum frequency of data-pins is the same as that of the clock-pins.
  • Figs. 2 and 3 depict the transmitter interface model and receiver interface model (SSPLrx).
  • the transmitter and receiver each include a logic and array core that utilizes a transmitter core clock (TCC) and receiver core-clock (RCC) respectively.
  • TCC transmitter core clock
  • RRC receiver core-clock
  • the transmitter clock (TcIk) is derived from TCC and is transmitted along with the data as a source-synchronous clock.
  • the source synchronous clock is received at the RcIk inputs of the receiver interface.
  • the phase relationship between RcIk and RCC is indeterminate.
  • the receiver interface samples data from the SSPL and transfers it to the receiver core-clock domain.
  • the D a In[N:0] input of the SSPLrx module is sampled with RclkaHIn and RcIk 3 LIn clocks.
  • the resultant data is output on D a ROut[N:0] (rising edge) and D a FOut[N:0] (falling edge) outputs in the RCC clock domain.
  • the module continuously samples the input and produces output.
  • the rcvRst input which is activated by an input pin and/or controlled through a programmable register, initializes the module
  • the design of the SSPLrx module does not depend on phase comparison between the source-synchronous clocks and the receiver core-clock. Such phase difference may change during device operation and cause clock-slip. Therefore, the SSPLrx design uses a synchronization technique that does not depend on phase comparison and is described in greater detail below.
  • Fig. 4 shows the SSPL clock distribution technique.
  • the transmitting device uses logic that runs at the SSPL-clock frequency or at double that frequency.
  • the clock used for this logic is referred to as the transmitter core-clock.
  • the transmitter core- clock is synthesized from a clean (low-jitter) system clock input to the transmitter device.
  • the leaf of the clock-tree of the transmitter core-clock may be phase-locked to the system clock input and the SSPL-clock is derived from the transmitter core clock.
  • the receiving device uses logic that runs at the SSPL clock-frequency or at double that frequency.
  • the clock used for this logic is referred to as the receiver core- clock.
  • the receiver core-clock is synthesized from a clean (low-jitter) system clock input to the receiving device.
  • the leaf of the clock tree of the receiver core-clock may be phase- locked to the system clock input.
  • the system clock inputs to the transmitting and receiving devices are of the same frequency as the SSPL clock and are copies of a clock from the same source.
  • the transmitter and receiver core-clocks may be phase-locked to the system clock inputs so that the transmitter and receiver clock-tree delays have no effect on the phase difference between the transmitter- and receiver-core clock.
  • the transmitter and receiver core clocks can have different frequencies.
  • the number of data-pins associated with a pair of clock pins is limited to between 18 and 20.
  • bandwidth requirement of the link requires a large number of data-pins, multiple copies of clocks are used.
  • Each pair of clock-pins and the associated data-pins are referred to as a clock-group.
  • the pins of a clock-group are located physically close to each other in both the transmitting and the receiving device.
  • the transmitter and the traces are carefully designed to minimize skew within a clock group.
  • the clocks carried on these pins are derived from the same source, the skew between the clock copies in the different clock- groups may be substantial at the receiver interface.
  • FIG. 5 depicts an SSPL with two clock groups, referred to as "a” and "b".
  • each clock-group may be skewed relative to each other when they arrive at the receiver interface.
  • a system for removing the skew between the signal groups sampled by different received copies of the transmit clock will now be described.
  • Fig. 6 depicts an embodiment where the receiving device uses a separate copy of the receiver interface module for each clock-group.
  • a deskew logic block is also depicted which aligns the data output from the multiple receiver interfaces and presents it to the receiver core.
  • the timing budget limits the maximum skew between the outputs of the different receiver interface modules to one RCC period. Therefore, data from clock-groups that arrive early may need to be delayed by one RCC period in order to align with data from clock-groups that arrive late.
  • the invention is not limited by this constraint and the skew between clock groups may be less than, equal to, or greater than the clock period.
  • deskew logic may be:
  • a protocol at device initialization is used to align the edges of the different copies of the SSPL-clock at the receiver.
  • One data pin of each SSPL clock group is used for this purpose and is referred to as the SSPL-Mt pin.
  • the transmitting device drives '0' on the SSPL-Mt pin of all clock-groups. This is called the initial-value. Then it drives ' 1 ' on the SSPL-Mt pin of all the clock-groups (e.g. D a [0] and D b [O] in Fig. 5) simultaneously for one SSPL clock period. This value is called the initialization-pattern.
  • the sequence of driving the Mtial-value followed by the initialization-pattern is referred to as the training sequence.
  • the receiving device detects the transition from initial- value to the Mtialization-pattern to deskew the data sampled from different clock-groups as described in more detail below.
  • Fig. 7 is a detailed schematic diagram of an embodiment of the SSPLrx that supports deskewing.
  • Fig. 7 depicts a Deskew State Machine that generates a FIFO Write Restart (WrRst) signal, two FIFOs for buffering data received on the rising and falling edges of RcIk, an M stage synchronizer, and a RdyCtrl block to generate the FIFO Read Restart (RdRst) signal.
  • WrRst FIFO Write Restart
  • RdRst FIFO Read Restart
  • the data is held for deskewing in the FIFOs inside the SSPLrx module.
  • the delay between writing the first data to the Rclk a LIh-clocked FIFO and reading the same data is established during device initialization and by the delay through synchronizer.
  • the synchronizer uses 'M' stages of flops clocked with RCC. For other synchronizer structures and core-clock frequencies, the delay through the synchronizer will be different. Depending on that and the timing budget, it may be necessary to have additional delays and/or flops to generate ready#A.
  • the depth of the FIFOs must be such that the output data is held valid for sufficient time before that entry in the FIFO is overwritten with new data.
  • the maximum time that the write-clock can advance and the maximum synchronizer-delay is factored into deciding the FIFO depth.
  • the deskew state machine drives the WrRst input of the FIFO high to hold the write-pointer of these FIFOs in the initial state.
  • the deskew state machine drives WrRst low and allows the write-pointer to advance.
  • the write-pointer and ready# signals in all the SSPLrx modules are controlled by the transmitter interface through the initialization sequence.
  • Fig. 7 depicts the deskew logic for clock group "a". This logic is repeated for each clock group. As described above, during the initialization signal a logic "0" signal is driven on the DO[O] signal of each clock group and these signals may be skewed relative to each other. Thus, the time of assertion of WrRst signal will vary from clock group to clock group depending on the amount of relative skew between the clock groups.
  • Fig. 8 depicts an embodiment of the FIFO, depicted in Fig. 7, that supports separate read- and write-clocks.
  • the FIFO is deep enough to absorb the skew between the different clock groups.
  • the write-clock is used to write data to the FIFO and advance the write-pointer inside the FIFO.
  • the input incrRdPtr is used to control the increment step of the read-pointer. This input is tied to ' 1 ' where both clocks are the same frequency.
  • the WrRst and RdRst to the FIFO counters are active low.
  • the WrRst input to the FIFO is high, the counter used for the write-pointer inside the FIFO is held in its initial state.
  • the RdRst input to the FIFO is high, the counter used for the read-pointer inside the FIFO is held in its initial state.
  • the FIFO uses edge-triggered D- flops with enable (EN) as storage elements. The data input is sampled by one set of the flops even when WrRst is driven low.
  • the WrRst signal is not driven low until the transition from the initial value to the initialization pattern of the training sequence is detected.
  • the WrRst signal is then input to the synchronizer and is output as the Ready# signal after a fixed delay.
  • the amount of this fixed delay is controlled by a value encoded in the DeskewStateDelay[d:O] signal.
  • the actual implementation of the system on a chip may require additional flops to be added between the output of the M stage Synchronizer and the SSPLrx macros thereby inserting additional delay after FIFO initialization requiring more FIFO depth.
  • the DeskewStateDelay[d:O] signal is used to program the Deskew State machine to delay the assertion of WrRst to the input of the Synchronizer. This delayable WrRst signal is denominated as the RdRstSync signal in Fig. 7.
  • the Deskew Logic of Fig. 6 includes a one-bit deskew state machine (Fig. 7) for clock group "a" and another one-bit deskew state machine for clock group "b".
  • Fig. 7 a one-bit deskew state machine
  • FIG. 8 the FIFO starts writing the received data when WrRst is driven low.
  • the RdRstSync signal is driven low either simultaneously with WrRstA or after a fixed delay.
  • the M stage Synchronizer is driven by the internal clock signal RCC and forms the boundary between the receive clock domain and the internal clock domain.
  • the Ready signal is synchronized to RCC.
  • the signal Ready b will be driven low before the signal Ready a .
  • all FIFO Read Counters receive a RdRst signal which is in the form of the logical OR of all the Readyi signals driven low by the individual one-bit deskew state machines so that no data will be read from the FIFOs until the initial data of all the clock groups has been written to a corresponding FIFO.
  • the RdRst signal can thus be used to keep all the FIFO read pointers "on hold" until all the groups are initialized and ready to read out data.
  • the RdRst signal will be driven low only after both Ready a and Readyb signals are driven low and the first data received on both clock group signals will be read in synchronism from the FIFO when RdRst is driven low and the skew between the clock groups is removed.
  • the RdRst signal is derived from the logical AND of the Ready a and Ready b signals delayed by S clocks, where an interval of S clock delays is greater than the maximum budgeted skew interval between clock groups.
  • the RdRst signal will be driven low if any of the Ready signals are driven low. This removes a possible fault where one of the Ready signals getting stuck could hang up the receiver. However, this option adds a delay since the Ready signal will not be driven low until after the S clock delayed expires.
  • Fig. 9 shows a configuration where multiple transmitting devices are connected in parallel.
  • the receiving device needs to align the data from UIa and UIb received from two different devices.
  • One or more bits on the SSPL must provide framing information for data-frame driven by the transmitting device.
  • the receiving device can use the framing information to align data from different transmitting devices.
  • this embodiment provides a mechanism to align data from two transmitting devices.
  • UIa and UIb drive the initialization pattern in response to a "send-initialization-pattern" command from the receiving device.
  • the receiving device sends this command simultaneously to both Ul a and UI b .
  • the receiving device uses the initialization patterns from the two devices to deskew the data from Ul 3 and UI b (similarly to how a receiving device deskews data from two clock-groups as described with reference to Figs. 7- 8).
  • the skew between the source-synchronous clocks from two devices can be larger than the skew between two clock-groups from the same device.
  • the SSPL is also used for CmdA and • CmdB buses.
  • the synchronizers in the SSPL receiver interface in Ul a and UI b may skew the commands by an additional period. If the latency of Ul 3 and UI b is unequal, the receiving device needs to support additional skew amounting to the latency-difference between Ul a and Ul b .
  • the data in the different SSPLrx modules in the receiving device may be skewed with respect to each other due to:
  • FIG. 10 depicts two chips: Tx (transmitter) and Rx (receiver). There are two busses going from T to R, labeled M and N where M has 3 clock groups Ml , M2 and M3, and N has 2 clock groups Nl andN2.
  • the bus M is self-contained and independent of N, meaning all the necessary signaling is present within M so that the core logic in Tx can transfer data through M to core logic in R.
  • N is self-contained and independent of M. This means that buses M and N can independently carry two "streams" of data from Tx to Rx.
  • M1-M2 and N1-N2 are treated as a single bus in the SSPL domain, so that the total delay on all the groups (Ml, M2, M3, Nl, N2) would be 5 (due to N2), the temporal relationship between the data on M and N is preserved, and the transmitter core and receiver core remain in sync with respect to M and N regardless of physical layer skews to nicely decouple the logic layer protocols from the physical layer protocols and keep the core logic design "clean" and independent of SSPL skews.

Landscapes

  • Engineering & Computer Science (AREA)
  • Signal Processing (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Theoretical Computer Science (AREA)
  • General Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Computer Hardware Design (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Power Engineering (AREA)
  • Synchronisation In Digital Transmission Systems (AREA)
  • Information Transfer Systems (AREA)
  • Time-Division Multiplex Systems (AREA)
  • Communication Control (AREA)

Abstract

A source-synchronous parallel interface divides a wide data bus into clock-groups including a sub-group of the data lines and a clock line carrying a copy of the transmit clock. The traces in a clock-group are located physically close together to minimize skew between the signals carried on the traces of the clock-group. Deskew logic on the receiver compensates for skew between received clock-group signals.

Description

PATENT APPLICATION
METHOD AND APPARATUS FOR ALIGNING DATA TN A WIDE, HIGH-SPEED, SOURCE SYNCHRONOUS PARALLEL LINK
RELATED APPLICATIONS
[01] This application is a continuation in part of the commonly-assigned United States patent application entitled HIGH-SPEED MEMORY FOR USE IN NETWORKING SYSTEMS, filed June 16, 2003, A/N 10/462,866, which is hereby incorporated by reference for all purposes.
BACKGROUND OF THE INVENTION
[02] The source-synchronous bus has been used to increase the speed of buses in many designs. Data and clock are sourced from the same device on the bus. The receiving device uses the clock from the bus to sample the data on the bus. Since the clock and data are driven and distributed similarly, they have similar delays and hence such buses can be run faster than buses using other clocking schemes.
[03] At higher speed, being able to drive a clock becomes challenging especially when the data pins are driven and sampled on both edges of the clock. This is referred to as double-data rate or DDR.
[04] One of the limitations on speed derives from the fact that as the number of data pins gets large, the skew between those pins increases, where Clock Skew is the variation in the transition point of a clock signal due to delay in the propagation path. Since all pins need to be sampled with the same clock, clock skew limits the speed of the bus. In DDR3, SRAMs, and in fast packet forwarding ASICs, this limitation is overcome by limiting the number of data pins associated with a clock pin. For wider data buses, multiple copies of source-synchronous CIoCkS1 are used. But still the skew between copies of clocks has to be limited to much less than the clock period in order to align the data sampled with different copies of clocks.
[05] Accordingly, new parallel interfaces need to be developed that allow high speed data transfer between Devices with a large number of pins.
BRIEF SUMMARY OF THE INVENTION
[06] One embodiment of the invention allows the clock period of the transmit and receive core clocks to get smaller than the skew between the copies of source- synchronous clock. The maximum frequency of operation of the link is thereby increased to the limit reachable for sampling a small number of data pins with a source-synchronous clock received on a pair (clock high and clock low) of clock pins. There is no limit imposed because of skew between multiple copies of source-synchronous clocks.
[07] In another embodiment of the invention, for each copy of source- synchronous clock, data is written into a receive-data FIFO in the receiver and data is read from all these FIFOs using a single core clock. An initialization protocol is used to align data between multiple FIFOs. The initialization protocol and the receive-data FIFOs can also be used to align data coming from multiple devices connected in parallel to the same receiving device.
[08] In another embodiment of the invention, both the transmitting and receiving devices use a PLL (phase-locked loop) to phase-align their internal core-clocks with a common external reference clock. This limits the jitter and wander of the source- synchronous clock with respect to the receiver core-clock and that in turn reduces the depth of the receive-data FIFOs.
[09] In another embodiment of the invention, the transmitting device may send data in a single clock from one or more logical buses in its core-clock domain over multiple source synchronous links. The receive-data FIFOs and the deskew protocol align the data from the logical bus(es) in the core-clock domain of the receiving device.
[10] Other features and advantages of the invention will be apparent in view of the following detailed description and appended figures.
BRIEF DESCRIPTION OF THE DRAWINGS
[11] Fig. 1 is a functional timing diagram of an embodiment of the invention;
[12] Fig. 2 depicts the transmitter interface model of an embodiment of the invention;
[13] Fig. 3 depicts the receiver interface model of an embodiment of the invention;
[14] Fig. 4 depicts the clock distributions and clock domains of an embodiment of the invention;
[15] Fig. 5 depicts an embodiment of the invention having a link using two clock-groups;
[16] Fig. 6 depicts a receiving device model for multiple clock copies of an embodiment of the invention; [17] Fig. 7 depicts an embodiment of the invention having deskew logic in the receive interface model;
[18] Fig. 8 depicts a FIFO of an embodiment of the invention;
[19] Fig. 9 depicts two transmitters coupled in parallel to a receiver; and
[20] Fig. 10 depicts a transmitter and receiver coupled by independent logical buses divided into clock groups.
DETAILED DESCRIPTION OF THE INVENTION [21] Reference will now be made in detail to various embodiments of the invention. Examples of these embodiments are illustrated in the accompanying drawings. While the invention will be described in conjunction with these embodiments, it will be understood that it is not intended to limit the invention to any embodiment. On the contrary, it is intended to cover alternatives, modifications, and equivalents as maybe included within the spirit and scope of the invention as defined by the appended claims. In the following description, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments. However, the present invention may be practiced without some or all of these specific details. In other instances, well known process operations have not been described in detail in order not to unnecessarily obscure the present invention.
[22] Several embodiments will now be described to implement a high-speed source-synchronous parallel link (SSPL) for transferring data at high speed between devices with a large number of pins. The embodiments include features such as multiple clock groups where multiple clock copies are transmitted and a limited number of data pins are associated with each clock signal, a deskew unit that aligns data sampled from different clock groups to a core clock, a clock generation system for forming clean copies of core clocks that have low jitter and noise, etc.
[23] In the following embodiments a synchronous unidirectional parallel interface is described where the interface includes data-pins and clock-pins. All the information is carried on data-pins and the clock-pins toggle at a fixed frequency. This clock is referred to as the SSPL-clock. The receiver uses this clock to sample data received on the data-pins. The set of bits transferred in a clock cycle-time is referred to as a data-word and the link supports the transfer of a continuous stream of data-words, one in each clock cycle- time. [24] Figure 1 is a functional timing diagram of the SSPL of this embodiment where the set of bits transferred in RO (rising edge of DcIk) and FO (falling edge of DcIk) form a data word. DcIkJI (data clock high) and DcIk2L (data clock low) refer to the high and low edges of the SSPL clock pair and Da[N:0] is the group of data pins sampled by using clock pair DcIk3H. and DcIk3L. In various embodiments, data can be transferred on rising and falling edge of a single clock, on the rising edges of two complementary clocks, or on the falling edges of two complementary clocks.
[25] In the following, the group of N+l data pins, referred to as Di[N:0] is associated with a pair of clock pins, DclkjH & DclkjL. To describe multiple sets of such data and clock pins the letters a, b, and so on are used to replace the subscript "i", e.g., Da[N:0] associated with DcIk3H & DcIk3L, Db[N: 0] associated with DclkbH & DclkbL and so on.
[26] Ia this embodiment the following design choices are made: ' • The SSPL uses source-synchronous clocking, Le. data-pins and clock-pins are driven from the same source device to minimize skew between the data-pins and clock-pins.
• Each data-pin of the SSPL transfers 2 bits of information in one cycle-time. Thus the maximum frequency of data-pins is the same as that of the clock-pins.
• Data transition is center-aligned with the clock edges.
• All data- and clock-pins are single-ended.
• Clock pins occur in pairs (of opposite phase).
[27] However, a dual clock signal utilizing either differential or complementary logic, can be utilized as is known in the art.
[28] Figs. 2 and 3 depict the transmitter interface model and receiver interface model (SSPLrx). The transmitter and receiver each include a logic and array core that utilizes a transmitter core clock (TCC) and receiver core-clock (RCC) respectively. The transmitter clock (TcIk) is derived from TCC and is transmitted along with the data as a source-synchronous clock. The source synchronous clock is received at the RcIk inputs of the receiver interface. The phase relationship between RcIk and RCC is indeterminate. The receiver interface samples data from the SSPL and transfers it to the receiver core-clock domain.
[29] The DaIn[N:0] input of the SSPLrx module is sampled with RclkaHIn and RcIk3LIn clocks. The resultant data is output on DaROut[N:0] (rising edge) and DaFOut[N:0] (falling edge) outputs in the RCC clock domain. After initialization, the module continuously samples the input and produces output. The rcvRst input, which is activated by an input pin and/or controlled through a programmable register, initializes the module
[30] The design of the SSPLrx module does not depend on phase comparison between the source-synchronous clocks and the receiver core-clock. Such phase difference may change during device operation and cause clock-slip. Therefore, the SSPLrx design uses a synchronization technique that does not depend on phase comparison and is described in greater detail below.
[31] As depicted in Fig. 3, data from the SSPL is clocked in using RcIk and clocked out using RCC.
[32] A technique for synthesizing clean (low-jitter) transmitter and receiver core clocks, where jitter refers to the uncertainty, or variability, of waveform timing, will now be described with reference to Fig. 4.
[33] Fig. 4 shows the SSPL clock distribution technique. The transmitting device uses logic that runs at the SSPL-clock frequency or at double that frequency. The clock used for this logic is referred to as the transmitter core-clock. The transmitter core- clock is synthesized from a clean (low-jitter) system clock input to the transmitter device. The leaf of the clock-tree of the transmitter core-clock may be phase-locked to the system clock input and the SSPL-clock is derived from the transmitter core clock.
[34] The receiving device uses logic that runs at the SSPL clock-frequency or at double that frequency. The clock used for this logic is referred to as the receiver core- clock. The receiver core-clock is synthesized from a clean (low-jitter) system clock input to the receiving device. The leaf of the clock tree of the receiver core-clock may be phase- locked to the system clock input.
[35] In this embodiment, the system clock inputs to the transmitting and receiving devices are of the same frequency as the SSPL clock and are copies of a clock from the same source. Also, the transmitter and receiver core-clocks may be phase-locked to the system clock inputs so that the transmitter and receiver clock-tree delays have no effect on the phase difference between the transmitter- and receiver-core clock. Alternatively, the transmitter and receiver core clocks can have different frequencies.
[36] An embodiment that utilizes multiple source-synchronous clock groups will now be described with reference to Fig. 5.
[37] In this embodiment, the number of data-pins associated with a pair of clock pins is limited to between 18 and 20. When bandwidth requirement of the link requires a large number of data-pins, multiple copies of clocks are used. Each pair of clock-pins and the associated data-pins are referred to as a clock-group.
[38] The pins of a clock-group are located physically close to each other in both the transmitting and the receiving device. The transmitter and the traces are carefully designed to minimize skew within a clock group. Though the clocks carried on these pins are derived from the same source, the skew between the clock copies in the different clock- groups may be substantial at the receiver interface.
[39] Fig. 5 depicts an SSPL with two clock groups, referred to as "a" and "b". By dividing the wide bus into source-synchronous clock groups the wide bus is effectively divided up into a series of smaller buses to reduce skew and allow for higher clock speeds.
[40] However, as described above, the different copies of the clock, and associated data signals, in each clock-group may be skewed relative to each other when they arrive at the receiver interface. A system for removing the skew between the signal groups sampled by different received copies of the transmit clock will now be described.
[41] Fig. 6 depicts an embodiment where the receiving device uses a separate copy of the receiver interface module for each clock-group. A deskew logic block is also depicted which aligns the data output from the multiple receiver interfaces and presents it to the receiver core.
[42] The output of the S SPLrx modules in the different receiver interfaces may be skewed with respect to each other due to:
• Skew between copies of SSPL-clocks
• Skew between rcvRst inputs to S SPLrx
• Skew between RCC inputs to SSPLrx
[43] In this embodiment, the timing budget limits the maximum skew between the outputs of the different receiver interface modules to one RCC period. Therefore, data from clock-groups that arrive early may need to be delayed by one RCC period in order to align with data from clock-groups that arrive late. However, the invention is not limited by this constraint and the skew between clock groups may be less than, equal to, or greater than the clock period.
[44] Ih different embodiments the deskew logic may be:
• Integrated with receiver interface,
• Integrated with receiver core, • Implemented as a separate module.
[45] As depicted in Figs. 1 and 2, data is clocked on the rising and falling edges of the transmit clock. In this embodiment the period of TcIk and RCC are the same and the first data frame is clocked on TcIkH. Therefore, in the example of two clock groups, "a" and "b", skewed by one RCC clock cycle where clock-group a is delayed relative to clock group "b", the data sampled with RclkaJB could arrive after the data sampled with RclkbL. The deskewing logic correctly aligns the data presented to the receiver core.
[46] A protocol at device initialization is used to align the edges of the different copies of the SSPL-clock at the receiver. One data pin of each SSPL clock group is used for this purpose and is referred to as the SSPL-Mt pin.
[47] Initially the transmitting device drives '0' on the SSPL-Mt pin of all clock-groups. This is called the initial-value. Then it drives ' 1 ' on the SSPL-Mt pin of all the clock-groups (e.g. Da[0] and Db[O] in Fig. 5) simultaneously for one SSPL clock period. This value is called the initialization-pattern. The sequence of driving the Mtial-value followed by the initialization-pattern is referred to as the training sequence. The receiving device detects the transition from initial- value to the Mtialization-pattern to deskew the data sampled from different clock-groups as described in more detail below.
[48] A first embodiment of the deskew integrated with the receiver interface will now be described with reference to Fig. 7 and Fig. 8.
[49] Fig. 7 is a detailed schematic diagram of an embodiment of the SSPLrx that supports deskewing. Fig. 7 depicts a Deskew State Machine that generates a FIFO Write Restart (WrRst) signal, two FIFOs for buffering data received on the rising and falling edges of RcIk, an M stage synchronizer, and a RdyCtrl block to generate the FIFO Read Restart (RdRst) signal. The data is held for deskewing in the FIFOs inside the SSPLrx module. The delay between writing the first data to the RclkaLIh-clocked FIFO and reading the same data is established during device initialization and by the delay through synchronizer. IQ the model represented in Fig. 7, the synchronizer uses 'M' stages of flops clocked with RCC. For other synchronizer structures and core-clock frequencies, the delay through the synchronizer will be different. Depending on that and the timing budget, it may be necessary to have additional delays and/or flops to generate ready#A.
[50] The depth of the FIFOs must be such that the output data is held valid for sufficient time before that entry in the FIFO is overwritten with new data. The maximum time that the write-clock can advance and the maximum synchronizer-delay is factored into deciding the FIFO depth. [51] After device initialization, the deskew state machine drives the WrRst input of the FIFO high to hold the write-pointer of these FIFOs in the initial state. After detecting the initialization sequence, the deskew state machine drives WrRst low and allows the write-pointer to advance. Thus, the write-pointer and ready# signals in all the SSPLrx modules are controlled by the transmitter interface through the initialization sequence.
[52] 3h this embodiment, in order to avoid putting extra load on DaIn[O], the DO[O] input of the deskew state machine is driven from the flop in the FIFO that samples DaIn[0] during device initialization. This flop in the FIFO must not be held under reset in order to allow propagation of the DaIn[0] value when WrRst is high;
[53] The output of the deskew logic (Q) is initialized to LOW when the Rst signal is asserted. It then is driven to HIGH and remains HIGH when the training sequence (D[O] =1) is received.
[54] Fig. 7 depicts the deskew logic for clock group "a". This logic is repeated for each clock group. As described above, during the initialization signal a logic "0" signal is driven on the DO[O] signal of each clock group and these signals may be skewed relative to each other. Thus, the time of assertion of WrRst signal will vary from clock group to clock group depending on the amount of relative skew between the clock groups.
[55] Fig. 8 depicts an embodiment of the FIFO, depicted in Fig. 7, that supports separate read- and write-clocks. The FIFO is deep enough to absorb the skew between the different clock groups. The write-clock is used to write data to the FIFO and advance the write-pointer inside the FIFO. The read-clock is used to sample data from the FIFO and advance the read-pointer inside the FIFO, if (incrRdPtr = 1). In implementations where the frequency of the RdCIk is a multiple of the frequency of WrCIk, the input incrRdPtr is used to control the increment step of the read-pointer. This input is tied to ' 1 ' where both clocks are the same frequency.
[56] The WrRst and RdRst to the FIFO counters are active low. When the WrRst input to the FIFO is high, the counter used for the write-pointer inside the FIFO is held in its initial state. Similarly, when the RdRst input to the FIFO is high, the counter used for the read-pointer inside the FIFO is held in its initial state. The FIFO uses edge-triggered D- flops with enable (EN) as storage elements. The data input is sampled by one set of the flops even when WrRst is driven low.
[57] The use of the deskew logic to deskew data between multiple clock groups will now be described in more detail with reference to Figs. 6, 7 and 8. As described above with reference to Fig. 7, the WrRst signal is not driven low until the transition from the initial value to the initialization pattern of the training sequence is detected. The WrRst signal is then input to the synchronizer and is output as the Ready# signal after a fixed delay.
[58] The amount of this fixed delay is controlled by a value encoded in the DeskewStateDelay[d:O] signal. The actual implementation of the system on a chip may require additional flops to be added between the output of the M stage Synchronizer and the SSPLrx macros thereby inserting additional delay after FIFO initialization requiring more FIFO depth. In this embodiment the DeskewStateDelay[d:O] signal is used to program the Deskew State machine to delay the assertion of WrRst to the input of the Synchronizer. This delayable WrRst signal is denominated as the RdRstSync signal in Fig. 7.
[59] In this embodiment, the Deskew Logic of Fig. 6 includes a one-bit deskew state machine (Fig. 7) for clock group "a" and another one-bit deskew state machine for clock group "b". In this example it is assumed that the data of clock group "a" are delayed relative to the data of clock group "b". As described above, WrRst will be driven low by each one-bit state machine when the transition of the initial value of training sequence is detected. Thus, referring to Fig. 8, the FIFO starts writing the received data when WrRst is driven low. In this case, because of the skew between clock groups "a" and "b", the signal WrRstb will be driven low before the signal WrRsta and data from clock group "b" will be read into the FIFO prior to data from group "a".
[60] The RdRstSync signal is driven low either simultaneously with WrRstA or after a fixed delay. The M stage Synchronizer is driven by the internal clock signal RCC and forms the boundary between the receive clock domain and the internal clock domain. The Ready signal is synchronized to RCC.
[61] In the example currently being described, the signal Readyb will be driven low before the signal Readya. However, in this case all FIFO Read Counters receive a RdRst signal which is in the form of the logical OR of all the Readyi signals driven low by the individual one-bit deskew state machines so that no data will be read from the FIFOs until the initial data of all the clock groups has been written to a corresponding FIFO. The RdRst signal can thus be used to keep all the FIFO read pointers "on hold" until all the groups are initialized and ready to read out data. Accordingly, the RdRst signal will be driven low only after both Readya and Readyb signals are driven low and the first data received on both clock group signals will be read in synchronism from the FIFO when RdRst is driven low and the skew between the clock groups is removed.
[62] In an alternative embodiment, the RdRst signal is derived from the logical AND of the Readya and Readyb signals delayed by S clocks, where an interval of S clock delays is greater than the maximum budgeted skew interval between clock groups. In this" case, the RdRst signal will be driven low if any of the Ready signals are driven low. This removes a possible fault where one of the Ready signals getting stuck could hang up the receiver. However, this option adds a delay since the Ready signal will not be driven low until after the S clock delayed expires.
[63] Fig. 9 shows a configuration where multiple transmitting devices are connected in parallel. Li response to identical commands sent by the receiving device on Cmda and Cmdb bus, UIa and UIb drives different parts of the same data-frame on Da[N:0] and Db[N:0]. The receiving device needs to align the data from UIa and UIb received from two different devices. One or more bits on the SSPL must provide framing information for data-frame driven by the transmitting device. The receiving device can use the framing information to align data from different transmitting devices.
[64] In case the transmitting device has the same latency for all commands, this embodiment provides a mechanism to align data from two transmitting devices. UIa and UIb drive the initialization pattern in response to a "send-initialization-pattern" command from the receiving device. During initialization, the receiving device sends this command simultaneously to both Ul a and UIb. The receiving device then uses the initialization patterns from the two devices to deskew the data from Ul3 and UIb (similarly to how a receiving device deskews data from two clock-groups as described with reference to Figs. 7- 8).
[65] Due to skew between the core-clocks of Ul a and Ulb, the skew between the source-synchronous clocks from two devices can be larger than the skew between two clock-groups from the same device. The SSPL is also used for CmdA and CmdB buses. The synchronizers in the SSPL receiver interface in Ula and UIb may skew the commands by an additional period. If the latency of Ul3 and UIb is unequal, the receiving device needs to support additional skew amounting to the latency-difference between Ul a and Ulb.
[66] The data in the different SSPLrx modules in the receiving device may be skewed with respect to each other due to:
• Skew between Cmda and Cmdb bus clocks
• Skew between TCCa and TCCb
• Phase-error and jitter of frequency-synthesizer inside Ula and Ulb • Delay difference between transmitter cores in Ul a and Ulb • Delay difference between source-synchronous clocks from Ul3 and UIb
• Skew between Rstln inputs to SSPLrx modules in the receiving device
• Jitter and skew of RCC in receiving device
[67] 3h another embodiment the deskew logic deskews independent buses while maintaining the temporal relationship between the data on the buses. For example, Fig. 10 depicts two chips: Tx (transmitter) and Rx (receiver). There are two busses going from T to R, labeled M and N where M has 3 clock groups Ml , M2 and M3, and N has 2 clock groups Nl andN2.
[68] The bus M is self-contained and independent of N, meaning all the necessary signaling is present within M so that the core logic in Tx can transfer data through M to core logic in R. Likewise, N is self-contained and independent of M. This means that buses M and N can independently carry two "streams" of data from Tx to Rx. However, there are applications where there is a "temporal" relationship between the data on M and N. For example, an element of data on M (like a packet) may precede an element of data on N (for example, some information related to the previous packet) by a fixed number of core clock periods. The following example illustrates this:
• M:XXXXXXXX1234XX56XX7XXXX....
• N:XXXXXXXXXXABCDXXEFXGXXX....
[69] The data "ABCD" on N follows the data "1234" on M by two clocks (in the transmitter core logic domain). The following is an example of what could happen when these busses go through the SSPL. Assuming Ml has a zero skew, and with respect to Ml,
• skew(Ml, M2) = 2
• skew(Ml, M3) = 4
• skew(Ml, Nl) = 4
• skew(Ml,-N2) = 5
[70] If M1-M3 and N1-N2 are treated as two busses and grouped separately, then:
• the "M" set deskews Ml , M2 and M3, and the total delay on M bus on the receiver side would be 4 (due to M3), and
• the "N" set deskews Nl and N2, and the total delay on N bus on the receiver side would be 5 (due to N2)
• this means that the receiver gets data on the N bus later with respect to data on M. [71] So, the SSPL skews in the "physical layer" (board, 10, etc.) have altered the temporal relationship between data on M and N and the receiver core logic has to have additional logic to handle this.
[72] Instead, in this embodiment M1-M2 and N1-N2 are treated as a single bus in the SSPL domain, so that the total delay on all the groups (Ml, M2, M3, Nl, N2) would be 5 (due to N2), the temporal relationship between the data on M and N is preserved, and the transmitter core and receiver core remain in sync with respect to M and N regardless of physical layer skews to nicely decouple the logic layer protocols from the physical layer protocols and keep the core logic design "clean" and independent of SSPL skews.
[73] The invention has now been described with reference to the preferred embodiments. Alternatives and substitutions will now be apparent to persons of skill in the art. For example, the logic levels described above are arbitrary and may be varied as is known in the art. Further, the number of data lines in a clock group depends on system design and timing budgets. Accordingly, it is not intended to limit the invention except as provided by the appended claims.

Claims

WHAT IS CLAIMED IS:
1. A source-synchronous parallel link system coupling a receiver device and a transmitter device, said system comprising: a parallel bus, coupling the transmitter and receiver, including N+l data lines and K clock lines, with N and K, such that K > 2 and N > K , where the N+l data lines are grouped into K clock-groups, each clock group including a distinct set of data lines and an associated clock line, with the data lines and associated clock line in a clock group located physically close to one another to minimize skew between the data signals and clock signal carried on the lines of the clock-group; with the transmitter comprising:
Tx core logic running at a transmitter core clock (TCC) frequency;
K transmitter interface modules, each coupled to an associated clock-group of the parallel bus, for clocking data onto the data lines of the associated clock-group in synchronism with a transmit clock copy derived from TCC; with the receiver comprising:
Rx core logic running at a receiver core clock (RCC) frequency;
K receiver interface modules, each coupled to an associated clock-group of the parallel bus, for sampling data from the data lines of the associated clock-group in synchronism with a receive clock copy received on the clock line of the associated clock- group; and deskew logic, coupled to receive the receive clock copy, RCC, and data sampled by the receiver interface modules, that removes skew between data signals received from different clock-groups prior to presenting the data signals to the core logic of the receiver.
2. The system of claim lwith the deskew logic further comprising a read- write FIFO with received data from a clock group written to the read-write FIFO using the receive clock copy and read from the FIFO using the RCC.
3. The system of claim 1 where data is clocked on the rising and falling edges of the transmit clock copy and a positive and negative transmit clock copy are carried on two clock lines of each clock-group.
4. The system of claim 1 where the transmitter core clock and receiver core clock are derived from a common clock source signal to reduce jitter.
5. The system of claim 1 where the Tx core logic transmits a training sequence, comprising a transition to an initial value on a data line of each clock group, and the deskew logic detects a transition to the initial value in the training sequence to deskew data sampled from different clock-groups.
6. The system of claim 5 where: the deskew logic asserts a FIFO write signal for each clock group when a transition to the initial value is detected on a data line of the clock group and where the deskew logic asserts a ready signal for each clock group delayed by a selected interval after the transition to the initial value is detected on the data line of the clock group; and where the deskew logic asserts a FIFO read signal after the ready signals from all clock groups have been asserted to deskew data received on the data lines of different clock groups.
7. The system of claim 1 further comprising: a second transmitter coupled to one of the clock groups of the parallel bus.
8. The system of claim 1 where TCC and RCC are not equal.
9. A method, performed at a receiver, for deskewing first and second groups of received data signals, with each of the first and second group of data signals accompanied, respectively, by a first and a second receive clock signal, with conductors for transmitting each group of data signals and the associated clock signal designed to minimize skew between received data and clock signals in the group, and with the receiver including first and second FIFOs coupled to receive, respectively, said first and second groups of data signals, said method comprising: sampling, in synchronism with the first receive clock signal, the first group of received data signals into the first FIFO after the transition of a selected received data signal, in the first group, to an initialization pattern value is detected; sampling, in synchronism with the second receive clock signal, the second group of received data signals into the second FIFO after the transition of a selected received data signal, in the second group, to an initialization pattern value is detected; asserting a first ready delayed by a fixed interval from when the transition to the initialization pattern value of the selected data signal, in the first group, is detected; asserting a second ready delayed by a fixed interval from when the transition to the initialization pattern value of the selected data signal, in the second group, is detected; and reading data from the first and second FIFOs after assertion of a read signal that is asserted subsequent to the assertion of both the first and second ready signals to remove relative skew between the data signals of the first and second data groups.
10. A system for deskewing first and second groups of received data signals, with each of the first and second group of data signals accompanied, respectively, by a first and second receive clock signal, with conductors for transmitting each group of data signals and the associated clock signal designed to minimize skew between received data and clock signals in the group, and with the receiver including first and second FIFOs coupled to receive, respectively, said first and second groups of data signals, said system comprising: means for sampling, in synchronism, with the first receive clock signal, the first group of received data signals into the first FIFO after the transition of a selected received data signal, in the first group, to an initialization pattern value is detected; means for sampling, in synchronism with the second receive clock signal, the second group of received data signals into the second FIFO after the transition of a selected received data signal, in the second group, to an initialization pattern value is detected; means for asserting a first ready delayed by a fixed interval from when the transition to the initialization pattern value of the selected data signal, in the first group, is detected; means for asserting a second ready delayed by a fixed interval from when the transition to the initialization pattern value of the selected data signal, in the second group, is detected; and means for reading data from the first and second FIFOs after assertion of a read signal that is asserted subsequent to the assertion of both the first and second ready signals to remove relative skew between the data signals of the first and second data groups.
11. A method, performed at a transmitter, for transmitting data on a wide parallel bus at high, speed, with the wide parallel bus including a plurality of signal groups being conductors designed to minimize skew between signals transmitted on a signal group, said method comprising: sampling, in synchronism with a transmit reference clock signal having a transmit frequency, a first subset of data signals to be transmitted on data lines of a first signal group of the wide parallel bus; generating a first clock signal, having a frequency equal to the transmit frequency, to be transmitted on a first clock line included in the first signal group of the wide parallel bus; sampling, in synchronism with a transmit reference clock signal having a transmit frequency, a second subset of data signals to be transmitted on data lines of a second signal group of the wide parallel bus; generating a second clock signal, having a frequency equal to the transmit frequency, to be transmitted on a second clock line included in the second signal group of the wide parallel bus; and generating a training sequence, prior to transmitting said data signals, transmitted on a selected data line in the first and second data groups.
12. A system, performed at a transmitter, for transmitting data on a wide parallel bus at high speed, with the wide parallel bus including a plurality of signal groups being conductors designed to minimize skew between signals transmitted on a signal group, said method comprising: means for sampling, in synchronism with a transmit reference clock signal having a transmit frequency, a first subset of data signals to be transmitted on data lines of a first signal group of the wide parallel bus; means for generating a first clock signal, having a frequency equal to the transmit frequency, to be transmitted on a first clock line included in the first signal group of the wide parallel bus; means for sampling, in synchronism with a transmit reference clock signal having a transmit frequency, a second subset of data signals to be transmitted on data lines of a second signal group of the wide parallel bus; means for generating a second clock signal, having a frequency equal to the transmit frequency, to be transmitted on a second clock line included in the second signal group of the wide parallel bus; and means for generating a training sequence, prior to transmitting said data signals, transmitted on a selected data line in the first and second data groups.
13. A system for transmitting data at high-speed over a wide, source synchronous, parallel link between a receiver, having a receiver core clock, and a transmitter comprising: a transmitter interface that divides the serial links into a set of clock groups, each clock group having a set of data lines and a clock line, with the clock line transmitting an associated clock signal utilized to sample data signals onto the data lines, and that generates an initialization signal sampled on a data line of each clock group at the same time; a receive FIFO for each clock group that absorbs skew between data signals in different clock groups; a receiver interface for each clock group that asserts a write start signal subsequent to the receipt of the initialization signal and that writes data into a corresponding FIFO in synchronism with the clock signal transmitted on the clock line subsequent to the assertion of the write signal; a synchronizer for each clock group that synchronizes the asserted write start signal to the receiver core clock and asserts a read ready signal after a fixed delay; and ready control block, coupled to the FIFO for each clock group and coupled to receive the ready signal for each clock group, for asserting a read start signal subsequent to the assertion of the ready signal for each clock group to synchronize the reading of data from all FIFOs and to remove skew between the clock groups.
EP05851478A 2004-11-15 2005-11-09 Method and apparatus for aligning data in a wide, high-speed, source synchronous parallel link Expired - Lifetime EP1813039B1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US10/989,703 US7720107B2 (en) 2003-06-16 2004-11-15 Aligning data in a wide, high-speed, source synchronous parallel link
PCT/US2005/040629 WO2006055374A2 (en) 2004-11-15 2005-11-09 Method and apparatus for aligning data in a wide, high-speed, source synchronous parallel link

Publications (3)

Publication Number Publication Date
EP1813039A2 true EP1813039A2 (en) 2007-08-01
EP1813039A4 EP1813039A4 (en) 2010-06-02
EP1813039B1 EP1813039B1 (en) 2012-01-04

Family

ID=36407632

Family Applications (1)

Application Number Title Priority Date Filing Date
EP05851478A Expired - Lifetime EP1813039B1 (en) 2004-11-15 2005-11-09 Method and apparatus for aligning data in a wide, high-speed, source synchronous parallel link

Country Status (4)

Country Link
US (1) US7720107B2 (en)
EP (1) EP1813039B1 (en)
AT (1) ATE540494T1 (en)
WO (1) WO2006055374A2 (en)

Families Citing this family (20)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
FI20041680L (en) * 2004-04-27 2005-10-28 Imbera Electronics Oy Electronic module and method for manufacturing the same
TWI306343B (en) * 2005-09-01 2009-02-11 Via Tech Inc Bus receiver and method of deskewing bus signals
US7415569B2 (en) * 2006-03-22 2008-08-19 Infineon Technologies Ag Memory including a write training block
US7454559B2 (en) * 2006-03-22 2008-11-18 Infineon Technologies Ag Filtering bit position in a memory
US7860202B2 (en) * 2006-04-13 2010-12-28 Etron Technology, Inc. Method and circuit for transferring data stream across multiple clock domains
KR100915387B1 (en) 2006-06-22 2009-09-03 삼성전자주식회사 Method and Apparatus for compensating skew between data signal and clock signal in parallel interface
US7555668B2 (en) * 2006-07-18 2009-06-30 Integrated Device Technology, Inc. DRAM interface circuits that support fast deskew calibration and methods of operating same
US7835479B2 (en) * 2006-10-16 2010-11-16 Advantest Corporation Jitter injection apparatus, jitter injection method, testing apparatus, and communication chip
US8185854B1 (en) * 2006-11-22 2012-05-22 Altera Corporation Method and apparatus for performing parallel slack computation within a shared netlist region
US8640066B1 (en) * 2007-01-10 2014-01-28 Cadence Design Systems, Inc. Multi-phase models for timing closure of integrated circuit designs
US7926011B1 (en) * 2007-01-10 2011-04-12 Cadence Design Systems, Inc. System and method of generating hierarchical block-level timing constraints from chip-level timing constraints
US8977995B1 (en) * 2007-01-10 2015-03-10 Cadence Design Systems, Inc. Timing budgeting of nested partitions for hierarchical integrated circuit designs
US8365113B1 (en) * 2007-01-10 2013-01-29 Cadence Design Systems, Inc. Flow methodology for single pass parallel hierarchical timing closure of integrated circuit designs
JP5056110B2 (en) * 2007-03-28 2012-10-24 ソニー株式会社 Integrated circuit generation apparatus and method
WO2009072038A2 (en) * 2007-12-05 2009-06-11 Nxp B.V. Source-synchronous data link for system-on-chip design
US8611178B2 (en) * 2011-11-11 2013-12-17 Qualcomm Incorporated Device and method to perform memory operations at a clock domain crossing
KR102140057B1 (en) * 2014-01-20 2020-07-31 삼성전자 주식회사 Data interface method having de-skew function and Apparatus there-of
US9832006B1 (en) * 2016-05-24 2017-11-28 Intel Corporation Method, apparatus and system for deskewing parallel interface links
US10129166B2 (en) 2016-06-21 2018-11-13 Intel Corporation Low latency re-timer
US10983944B2 (en) 2019-01-17 2021-04-20 Oracle International Corporation Method for training multichannel data receiver timing

Family Cites Families (14)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH04176232A (en) * 1990-11-09 1992-06-23 Hitachi Ltd Packet communication system and packet communication equipment
JP2694807B2 (en) * 1993-12-16 1997-12-24 日本電気株式会社 Data transmission method
US5734685A (en) * 1996-01-03 1998-03-31 Credence Systems Corporation Clock signal deskewing system
US6536025B2 (en) * 2001-05-14 2003-03-18 Intel Corporation Receiver deskewing of multiple source synchronous bits from a parallel bus
US6839862B2 (en) * 2001-05-31 2005-01-04 Koninklijke Philips Electronics N.V. Parallel data communication having skew intolerant data groups
US6920576B2 (en) * 2001-05-31 2005-07-19 Koninklijke Philips Electronics N.V. Parallel data communication having multiple sync codes
US7085950B2 (en) * 2001-09-28 2006-08-01 Koninklijke Philips Electronics N.V. Parallel data communication realignment of data sent in multiple groups
US6665218B2 (en) * 2001-12-05 2003-12-16 Agilent Technologies, Inc. Self calibrating register for source synchronous clocking systems
US20030112827A1 (en) * 2001-12-13 2003-06-19 International Business Machines Corporation Method and apparatus for deskewing parallel serial data channels using asynchronous elastic buffers
US6996738B2 (en) * 2002-04-15 2006-02-07 Broadcom Corporation Robust and scalable de-skew method for data path skew control
KR100440585B1 (en) * 2002-05-24 2004-07-19 한국전자통신연구원 The method and apparatus of de-skew for the transmission of high speed data among multiple lanes
KR100468733B1 (en) * 2002-06-07 2005-01-29 삼성전자주식회사 Skewed bus driving method and circuit
US20040249964A1 (en) * 2003-03-06 2004-12-09 Thibault Mougel Method of data transfer and apparatus therefor
US7209531B1 (en) * 2003-03-26 2007-04-24 Cavium Networks, Inc. Apparatus and method for data deskew

Also Published As

Publication number Publication date
US7720107B2 (en) 2010-05-18
EP1813039A4 (en) 2010-06-02
WO2006055374A2 (en) 2006-05-26
WO2006055374A3 (en) 2007-01-18
EP1813039B1 (en) 2012-01-04
US20050066142A1 (en) 2005-03-24
ATE540494T1 (en) 2012-01-15

Similar Documents

Publication Publication Date Title
US7720107B2 (en) Aligning data in a wide, high-speed, source synchronous parallel link
US11062743B2 (en) System and method for providing a configurable timing control for a memory system
US8209562B2 (en) Double data rate converter circuit includes a delay locked loop for providing the plurality of clock phase signals
EP1573516B1 (en) Clock skew compensation apparatus and method
US6930932B2 (en) Data signal reception latch control using clock aligned relative to strobe signal
US8073090B2 (en) Synchronous de-skew with programmable latency for multi-lane high speed serial interface
US6334163B1 (en) Elastic interface apparatus and method therefor
US7796652B2 (en) Programmable asynchronous first-in-first-out (FIFO) structure with merging capability
EP3155529B1 (en) Independent synchronization of output data via distributed clock synchronization
IL144674A (en) Dynamic wave-pipelined interface apparatus and methods therefor
US20040084537A1 (en) Method and apparatus for data acquisition
EP1510930A2 (en) Memory interface system and method
US7512201B2 (en) Multi-channel synchronization architecture
JP2023171864A (en) Deskew method for physical layer interface on multi-chip module
CN101253724B (en) Bit-deskewing IO method and system
US8498370B2 (en) Method and apparatus for deskewing data transmissions
JP3562416B2 (en) Inter-LSI data transfer system and source synchronous data transfer method used therefor

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

17P Request for examination filed

Effective date: 20070425

AK Designated contracting states

Kind code of ref document: A2

Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IS IT LI LT LU LV MC NL PL PT RO SE SI SK TR

AX Request for extension of the european patent

Extension state: AL BA HR MK YU

DAX Request for extension of the european patent (deleted)
A4 Supplementary search report drawn up and despatched

Effective date: 20100503

RIC1 Information provided on ipc code assigned before grant

Ipc: G06F 1/12 20060101ALI20100426BHEP

Ipc: H04J 3/06 20060101AFI20070517BHEP

Ipc: H04L 25/14 20060101ALI20100426BHEP

Ipc: H04L 7/04 20060101ALI20100426BHEP

Ipc: H04L 7/00 20060101ALI20100426BHEP

RIC1 Information provided on ipc code assigned before grant

Ipc: H04L 7/00 20060101ALI20110411BHEP

Ipc: H04L 25/14 20060101ALI20110411BHEP

Ipc: H04J 3/06 20060101AFI20110411BHEP

Ipc: H04L 7/04 20060101ALI20110411BHEP

Ipc: G06F 1/12 20060101ALI20110411BHEP

GRAP Despatch of communication of intention to grant a patent

Free format text: ORIGINAL CODE: EPIDOSNIGR1

GRAS Grant fee paid

Free format text: ORIGINAL CODE: EPIDOSNIGR3

GRAA (expected) grant

Free format text: ORIGINAL CODE: 0009210

AK Designated contracting states

Kind code of ref document: B1

Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IS IT LI LT LU LV MC NL PL PT RO SE SI SK TR

REG Reference to a national code

Ref country code: GB

Ref legal event code: FG4D

REG Reference to a national code

Ref country code: CH

Ref legal event code: EP

REG Reference to a national code

Ref country code: AT

Ref legal event code: REF

Ref document number: 540494

Country of ref document: AT

Kind code of ref document: T

Effective date: 20120115

REG Reference to a national code

Ref country code: IE

Ref legal event code: FG4D

REG Reference to a national code

Ref country code: DE

Ref legal event code: R096

Ref document number: 602005032028

Country of ref document: DE

Effective date: 20120308

REG Reference to a national code

Ref country code: NL

Ref legal event code: VDEP

Effective date: 20120104

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: SI

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

LTIE Lt: invalidation of european patent or patent extension

Effective date: 20120104

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: LT

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

Ref country code: BE

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

Ref country code: NL

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

Ref country code: BG

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120404

Ref country code: IS

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120504

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: PL

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

Ref country code: LV

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

Ref country code: PT

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120504

Ref country code: FI

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

Ref country code: GR

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120405

REG Reference to a national code

Ref country code: AT

Ref legal event code: MK05

Ref document number: 540494

Country of ref document: AT

Kind code of ref document: T

Effective date: 20120104

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: CY

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: EE

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

Ref country code: DK

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

Ref country code: RO

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

Ref country code: SE

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

Ref country code: CZ

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

PLBE No opposition filed within time limit

Free format text: ORIGINAL CODE: 0009261

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: NO OPPOSITION FILED WITHIN TIME LIMIT

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: IT

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

Ref country code: SK

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

26N No opposition filed

Effective date: 20121005

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: AT

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

REG Reference to a national code

Ref country code: DE

Ref legal event code: R097

Ref document number: 602005032028

Country of ref document: DE

Effective date: 20121005

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: ES

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120415

REG Reference to a national code

Ref country code: CH

Ref legal event code: PL

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: LI

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20121130

Ref country code: CH

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20121130

REG Reference to a national code

Ref country code: IE

Ref legal event code: MM4A

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: IE

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20121109

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: TR

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20120104

Ref country code: MC

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20121130

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: LU

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20121109

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: HU

Free format text: LAPSE BECAUSE OF FAILURE TO SUBMIT A TRANSLATION OF THE DESCRIPTION OR TO PAY THE FEE WITHIN THE PRESCRIBED TIME-LIMIT

Effective date: 20051109

REG Reference to a national code

Ref country code: FR

Ref legal event code: PLFP

Year of fee payment: 11

REG Reference to a national code

Ref country code: FR

Ref legal event code: PLFP

Year of fee payment: 12

REG Reference to a national code

Ref country code: FR

Ref legal event code: PLFP

Year of fee payment: 13

PGFP Annual fee paid to national office [announced via postgrant information from national office to epo]

Ref country code: DE

Payment date: 20191127

Year of fee payment: 15

PGFP Annual fee paid to national office [announced via postgrant information from national office to epo]

Ref country code: FR

Payment date: 20191125

Year of fee payment: 15

PGFP Annual fee paid to national office [announced via postgrant information from national office to epo]

Ref country code: GB

Payment date: 20191127

Year of fee payment: 15

REG Reference to a national code

Ref country code: DE

Ref legal event code: R119

Ref document number: 602005032028

Country of ref document: DE

GBPC Gb: european patent ceased through non-payment of renewal fee

Effective date: 20201109

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: FR

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20201130

PG25 Lapsed in a contracting state [announced via postgrant information from national office to epo]

Ref country code: GB

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20201109

Ref country code: DE

Free format text: LAPSE BECAUSE OF NON-PAYMENT OF DUE FEES

Effective date: 20210601

P01 Opt-out of the competence of the unified patent court (upc) registered

Effective date: 20230525