EP4720853A1 - Extension processing circuitry start-up - Google Patents
Extension processing circuitry start-upInfo
- Publication number
- EP4720853A1 EP4720853A1 EP24706783.8A EP24706783A EP4720853A1 EP 4720853 A1 EP4720853 A1 EP 4720853A1 EP 24706783 A EP24706783 A EP 24706783A EP 4720853 A1 EP4720853 A1 EP 4720853A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- extension
- data processing
- processing circuitry
- circuitry
- instruction
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
- G06F9/3836—Instruction issuing, e.g. dynamic instruction scheduling or out of order instruction execution
- G06F9/3842—Speculative instruction execution
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30003—Arrangements for executing specific machine instructions
- G06F9/30076—Arrangements for executing specific machine instructions to perform miscellaneous control operations, e.g. NOP
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30003—Arrangements for executing specific machine instructions
- G06F9/30076—Arrangements for executing specific machine instructions to perform miscellaneous control operations, e.g. NOP
- G06F9/3009—Thread control instructions
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30181—Instruction operation extension or modification
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
- G06F9/3818—Decoding for concurrent execution
- G06F9/382—Pipelined decoding, e.g. using predecoding
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
- G06F9/3824—Operand accessing
- G06F9/3834—Maintaining memory consistency
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
- G06F9/3836—Instruction issuing, e.g. dynamic instruction scheduling or out of order instruction execution
- G06F9/3842—Speculative instruction execution
- G06F9/3844—Speculative instruction execution using dynamic branch prediction, e.g. using branch history tables
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
- G06F9/3854—Instruction completion, e.g. retiring, committing or graduating
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
- G06F9/3854—Instruction completion, e.g. retiring, committing or graduating
- G06F9/3858—Result writeback, i.e. updating the architectural state or memory
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
- G06F9/3861—Recovery, e.g. branch miss-prediction, exception handling
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
- G06F9/3877—Concurrent instruction execution, e.g. pipeline or look ahead using a secondary processor, e.g. coprocessor
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/46—Multiprogramming arrangements
Landscapes
- Engineering & Computer Science (AREA)
- Software Systems (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Advance Control (AREA)
Abstract
Apparatuses, methods of data processing, computer programs, and computer- readable media are disclosed. A data processing pipeline performs data processing operations defined by a received sequence of instructions. Extension processing circuitry associated with the data processing pipeline performs a delegated task in response to a delegation signal received from the data processing pipeline, performing the delegated task asynchronously to the data processing pipeline. The data processing pipeline performs speculative instruction execution. In response to an extension start instruction, the data processing pipeline issues the delegation signal to the extension processing circuitry to delegate the delegated task and. the extension processing circuitry responds by commencing the delegated task before a speculation confirmation is generated for the extension start instruction. The extension processing circuitry ensures that no results generated by the delegated task are visible outside the extension processing circuitry until the speculation confirmation is generated for the extension start instruction.
Description
EXTENSION PROCESSING CIRCUITRY START-UP
The present techniques relate to an apparatus, a method of operating an apparatus, a computer program, and a computer-readable medium.
An apparatus may comprise a data processing pipeline configured to perform data processing operations in dependence on a received sequence of instructions.
At least some examples provide an apparatus for data processing, comprising: a data processing pipeline configured to perform data processing operations in dependence on a received sequence of instructions, wherein the data processing pipeline comprises decoding circuitry configured to decode the received sequence of instructions and to generate control signals to control the data processing pipeline to perform the data processing operations; and extension processing circuitry associated with the data processing pipeline and configured to perform a delegated task in response to a delegation signal received from the data processing pipeline, wherein the extension processing circuitry is configured to perform the delegated task asynchronously to the data processing operations performed by data processing pipeline, wherein the data processing pipeline is configured to perform speculative instruction execution, whereby modifications of state of the apparatus resulting from execution of a speculatively executed instruction are not committed until a speculation confirmation is generated indicating that execution of the speculatively executed instruction was correct, wherein the decoding circuitry is responsive to an extension start instruction specifying the delegated task to generate the control signals to control the data processing pipeline to issue the delegation signal to the extension processing circuitry to delegate the delegated task, and wherein the extension processing circuitry is responsive to the delegation signal to commence the delegated task before the speculation confirmation is generated for the extension start instruction,
whereby the extension processing circuitry is configured to ensure that no results generated by the delegated task are visible outside the extension processing circuitry until the speculation confirmation is generated for the extension start instruction.
At least some examples provide a non-transitory computer-readable medium to store computer-readable code for fabrication of the apparatus.
At least some examples provide a method of operating an apparatus comprising: performing data processing operations in a data processing pipeline in dependence on a received sequence of instructions; decoding the received sequence of instructions in decoding circuitry of the data processing pipeline and generating control signals to control the data processing pipeline to perform the data processing operations; performing a delegated task in extension processing circuitry associated with the data processing pipeline in response to a delegation signal received from the data processing pipeline, wherein the delegated task is performed asynchronously to the data processing operations performed by data processing pipeline; performing speculative instruction execution in the data processing pipeline, whereby modifications of state of the apparatus resulting from execution of a speculatively executed instruction are not committed until a speculation confirmation is generated indicating that execution of the speculatively executed instruction was correct; in response to an extension start instruction specifying the delegated task, generating the control signals to control the data processing pipeline to issue the delegation signal to the extension processing circuitry to delegate the delegated task; in response to the delegation signal, commencing the delegated task in the extension processing circuitry before the speculation confirmation is generated for the extension start instruction; and ensuring that no results generated by the delegated task are visible to the data processing pipeline until the speculation confirmation is generated for the extension start instruction.
At least some examples provide a computer program for controlling a host data processing apparatus to provide an instruction execution environment, the computer program comprising: data processing pipeline logic configured to perform data processing operations in dependence on a received sequence of instructions, wherein the data processing pipeline logic comprises decoding logic configured to decode the received sequence of instructions and to generate control signals to control the data processing pipeline logic to perform the data processing operations; and extension processing logic associated with the data processing pipeline logic and configured to perform a delegated task in response to a delegation signal received from the data processing pipeline logic, wherein the extension processing logic is configured to perform the delegated task asynchronously to the data processing operations performed by data processing pipeline logic, wherein the data processing pipeline logic is configured to perform speculative instruction execution, whereby modifications of state of the instruction execution environment resulting from execution of a speculatively executed instruction are not committed until a speculation confirmation is generated indicating that execution of the speculatively executed instruction was correct, wherein the decoding logic is responsive to an extension start instruction specifying the delegated task to generate the control signals to control the data processing pipeline logic to issue the delegation signal to the extension processing logic to delegate the delegated task, and wherein the extension processing logic is responsive to the delegation signal to commence the delegated task before the speculation confirmation is generated for the extension start instruction, whereby the extension processing logic is configured to ensure that no results generated by the delegated task are visible to the data processing pipeline logic until the speculation confirmation is generated for the extension start instruction.
The present techniques will be described further, by way of example only, with reference to embodiments thereof as illustrated in the accompanying drawings, to be read in conjunction with the following description, in which:
Figure 1 schematically illustrates a data processing apparatus which may embody various examples of the present techniques;
Figure 2 schematically illustrates a data processing apparatus which may embody various examples of the present techniques;
Figure 3 schematically illustrates a data processing apparatus which may embody various examples of the present techniques;
Figure 4 is a state diagram illustrating an example set of states between which extension processing circuitry of the present techniques may transition;
Figure 5 schematically illustrates an extension start instruction being present on a predicted branch path during speculative instruction execution in accordance with some examples;
Figure 6 schematically illustrates a data processing apparatus comprising confidence calibration circuitry which may embody various examples of the present techniques;
Figure 7 schematically illustrates confidence calibration circuitry modifying or substituting branch predictions generated in a data processing pipeline for use by extension processing circuitry in accordance with some examples;
Figure 8 schematically illustrates an extension setup instruction causing extension processing circuitry to perform one or more preparatory steps in accordance with some examples;
Figure 9 schematically illustrates an extension setup instruction causing extension processing circuitry to perform one or more preparatory steps in accordance with some examples;
Figure 10 schematically illustrates an extension setup instruction causing extension processing circuitry to perform one or more preparatory steps in accordance with some examples;
Figure 11 is a flow diagram showing a sequence of steps that are taken in the method of some examples; and
Figure 12 schematically illustrates a simulator implementation that may be used.
In one example herein there is an apparatus for data processing, comprising: a data processing pipeline configured to perform data processing operations in dependence on a received sequence of instructions, wherein the data processing pipeline comprises decoding circuitry configured to decode the received sequence of instructions and to generate control signals to control the data processing pipeline to perform the data processing operations; and extension processing circuitry associated with the data processing pipeline and configured to perform a delegated task in response to a delegation signal received from the data processing pipeline, wherein the extension processing circuitry is configured to perform the delegated task asynchronously to the data processing operations performed by data processing pipeline, wherein the data processing pipeline is configured to perform speculative instruction execution, whereby modifications of state of the apparatus resulting from execution of a speculatively executed instruction are not committed until a speculation confirmation is generated indicating that execution of the speculatively executed instruction was correct, wherein the decoding circuitry is responsive to an extension start instruction specifying the delegated task to generate the control signals to control the data processing pipeline to issue the delegation signal to the extension processing circuitry to delegate the delegated task, and wherein the extension processing circuitry is responsive to the delegation signal to commence the delegated task before the speculation confirmation is generated for the extension start instruction, whereby the extension processing circuitry is configured to ensure that no results generated by the delegated task are visible outside the extension processing circuitry until the speculation confirmation is generated for the extension start instruction.
An apparatus comprising a data processing pipeline can be required to perform a limitless variety of data processing operations as defined by the sequence of instructions provided to it. In order efficiently to perform those data processing
operations, the data processing pipeline may be configured with a variety of functional units, each with a given specialised type of data processing ability, such as arithmetic logic units (ALUs), floating point (FP) units, load/store units, and so on. Yet even with such specialised functional units being provided as part of the data processing pipeline, the inventors of the present techniques have established that in some types of data processing, that is in certain programs (i.e. sequences of instructions), there can be particular functions which are frequently executed and which require an amount of processing, such that the provision of custom hardware dedicated to supporting these functions is worthwhile, since it could significantly impact the overall performance of the apparatus. In identifying such functions, two key properties were deemed to be relevant: a function’s ubiquity (i.e. it can also be found in the many other use-cases) and a function’s impact (i.e. the proportion of time spent executing such a function is a significant percentage of the overall runtime, such that improvements in its execution made a significant difference to the overall use-case). Such impactful, ubiquitous functions have been found to include tasks or functions such as memcpy, memset, compression, encryption, and string processing, although the present techniques are not limited to these particular examples. The present techniques provide extension processing circuitry that is associated with the data processing pipeline and is configured to perform such a function (a delegated task) in response to a delegation signal received from the data processing pipeline. Such extension processing circuitry may also be referred to as a threadlet extension (TE) herein. The sequence of operations it carries out to perform the defined function may also be referred to as a threadlet herein. The extension processing circuitry, although closely associated (tightly coupled) with the data processing pipeline, is configured to perform the delegated task asynchronously to the data processing operations performed by data processing pipeline. The data processing pipeline may also be referred to as the CPU herein. Threadlets are functions or collections of operations that can be executed asynchronously relative to other CPU activity once launched. The asynchronous operation of the extension processing circuitry with respect to the data processing pipeline is possible because, unlike some prior art techniques, the extension processing circuitry receives a directive or command from the thread currently executing on the CPU and performs the required operations independently, that is without requiring a stream of instructions from the CPU that
directly control or influence its internal operation. The CPU is therefore free to continue executing other code and potentially reduce overall runtime by overlapping the execution of the instruction stream after the directive or command is sent to the extension processing circuitry with the operation of the extension processing circuitry. Whilst the ability for the extension processing circuitry to operate asynchronously with respect to the data processing pipeline is thus advantageous, the inventors of the present techniques have observed that when the data processing pipeline is configured to perform speculative instruction execution, it can be disruptive to the efficient usage of the data processing pipeline if that speculative execution needs curtailing, when a task is delegated to the extension processing circuitry from the data processing pipeline. Nevertheless such curtailment of the extent of speculative execution is indeed a standard approach in the prior art when tasks are offloaded to special-purpose hardware that is coupled to a processor. For example, commonly a dispatch barrier may be deployed or the need to drain the reorder buffer (ROB) in the vicinity of special offload instructions is enforced, which can have an adverse effect on performance.
Such approaches of disabling speculation or launching a threadlet only after a speculation barrier might be detrimental to performance in the case where the granularity of the task performed by the threadlet is low. The reason for this is that out-of-order processors develop considerable “momentum” during their operation, that is, they build up a lot of relevant state within their pipelines that typically takes many cycles to regenerate when exceptional events such as speculation barriers or pipeline flushes occur. The impact of this effect will depend on a number of factors such as the size of the reorder buffer (instruction window), the number of issue slots, the number of functional units, the degree of speculation enabled by the branch prediction circuitry, the rate of collision between younger loads and older stores, etc.
The present techniques therefore take an alternative approach, which enables the start of the threadlet’ s execution to be speculative. That is, when an extension start instruction is encountered in the sequence of instructions, a requirement for the extension start instruction to be committed before the extension processing circuitry commences the delegated task is not made and instead the extension processing circuitry
can speculatively start the delegated task. The countermeasure provided to ensure that this approach does cause problems, is that the extension processing circuitry is configured to ensure that no results generated by the delegated task are visible outside the extension processing circuitry until the speculation confirmation is generated for the extension start instruction. Any new state generated by the delegated task is confined to the extension processing circuitry.
If the extension start instruction gets committed, then this fact is signalled to the extension processing circuitry and it can then start flushing parts of its LI cache or other internal buffers. However, if the extension start instruction gets cancelled, this is also signalled to the extension processing circuitry, and thus in such cases the data processing pipeline is configured to generate a speculation cancellation when it is determined that execution of the speculatively executed instruction was incorrect, and wherein the extension processing circuitry is responsive to the speculation cancellation generated for the extension start instruction to invalidate any result generated by the delegated task. The extension processing circuitry thus performs a rollback by invalidating any modifications made, such as in its LI cache, deleting any information gathered during setup operations such as TLB entries, and reverting to an idle state.
The extension processing circuitry may hold such results internally in a variety of ways, but in some examples the extension processing circuitry comprises a private data store configured to hold results generated by the delegated task. Note that generally speaking, the term “results” is used herein when applied to the product of data processing tasks performed by the extension processing circuitry and which according to the present techniques are confined to the extension processing circuitry and not visible to the remainder of the apparatus until the speculation confirmation is generated for the extension start instruction. Conversely, generally speaking the term “output(s)” is used in the case of the product of data processing tasks performed by the extension processing circuitry which are released from or made accessible outside the extension processing circuitry.
The capacity of the extension processing circuitry to hold new state which it generates is necessarily finite, and thus in some examples the extension processing circuitry is responsive to results generated by the delegated task reaching a result storage capacity of the private data store to stall performance of the delegated task. The extension processing circuitry then remains stalled, until the extension start instruction gets committed or cancelled.
For a data processing pipeline which performs speculative instruction execution, the decision of which branch outcome to speculatively execute is typically based on branch confidence information generated by branch prediction circuitry on the basis of previous execution history. Moreover the branch confidence information may also be tuned based on a cost-benefit analysis from the perspective of the data processing pipeline. However the present techniques recognise that such branch confidence information and such a cost-benefit analysis may not be as well suited to extension processing circuitry. In short, the cost of speculation for the extension processing circuitry may be different to that of the rest of the apparatus. In view of this, in some examples the data processing pipeline comprises branch prediction circuitry configured to generate branch confidence information indicative of a predicted likelihood of a branch direction being taken, wherein the data processing pipeline is configured to perform the speculative instruction execution in dependence on the branch confidence information, and wherein the apparatus further comprises confidence calibration circuitry configured to generate extension steering confidence information associated with the extension start instruction, wherein the extension processing circuitry is configured to determine whether to commence the delegated task before the speculation confirmation is generated for the extension start instruction in dependence on the extension steering confidence information.
This enables the decision of whether to launch the extension processing circuitry speculatively to allow for the particular cost and benefit with respect to the extension processing circuitry. This approach recognises that when the extension processing
circuitry is launched speculatively, when it should not have, there is potentially a higher penalty in terms of wasted system bandwidth, power, and so on than there would be for regular CPU instructions. On the other hand, when the extension processing circuitry is not launched speculatively at the appropriate time when it should have been, then there is potentially lost performance opportunity because it can take many cycles to complete the setup of the extension processing circuitry. Thus, the generated extension steering confidence information can provide a more tailored prediction of the permissible depth of speculation than the branch predictor might typically provide. For instance, the confidence calculation could be used to throttle prefetchers, or to determine how far to progress within the extension processing circuitry’s execution before stalling in order to await further confirmation from the data processing pipeline of the speculation result.
The confidence calibration circuitry may be arranged to generate the extension steering confidence information in dependence on a number of predefined parameters, observed runtime metrics, and/or other characteristics or factors relating to the extension processing circuitry. Accordingly in some examples, the confidence calibration circuitry is configured to generate the extension steering confidence information in dependence on the branch confidence information and in dependence on at least one extension processing circuitry specific factor.
In some examples, the confidence calibration circuitry configured to generate the extension steering confidence information in dependence on a relative size of the delegated task.
In some examples the confidence calibration circuitry is configured to generate the extension steering confidence information in dependence on an execution history of an instruction sequence portion comprising the extension start instruction. Execution path statistics around the call site where one or more extension start instructions occur may be gathered, and the extension steering confidence information generated at least in part based on those statistics.
In some examples the confidence calibration circuitry is configured to generate the extension steering confidence information in dependence on a stalling history of the extension processing circuitry. This may be a recent stalling history or a relevant history (e.g. of executing the same portion of code). Thus the confidence calibration circuitry can monitor the frequency of stalls in the extension processing circuitry, where one reason for such stalls can be the storage limits of the extension processing circuitry for holding speculatively generated state. If such limits are too frequently reached, causing corresponding stalls, the threshold for triggering speculative execution by the extension processing circuitry can be increased (by adjustment of the extension steering confidence information).
Launching the extension processing circuitry may require a number of initialisation actions to be carried out. Such initialisation actions can be triggered by the confidence calibration circuitry, when appropriate. Hence in some examples, the confidence calibration circuitry is configured to cause, in dependence on the extension steering confidence information, the extension processing circuitry to perform at least one preparatory step to configure the extension processing circuitry for performance of the delegated task.
Various preparatory steps may be initiated by the confidence calibration circuitry and the confidence calibration circuitry can also cause a selected instance of extension processing circuitry to be prepared for action, when the apparatus comprises multiple instances of extension processing circuitry. Thus in some examples there are multiple instances of extension processing circuitry, wherein the confidence calibration circuitry is configured to specify a selected instance of extension processing circuitry to perform the at least one preparatory step.
One example of a preparatory step is for internal caches of the extension processing circuitry to be warmed up, i.e. for particular content to be brought into such a cache or caches, in preparation for an expected workload. Various mechanisms may be employed to do this, but in some examples the confidence calibration circuitry is configured specify an address identifier, wherein the address identifier is indicative of
cache content which is to be brought into a private cache of the extension processing circuitry as at least part of the at least one preparatory step.
Alternatively, or in addition to, the ability of the confidence calibration circuitry to initiate various preparatory steps for the extension processing circuitry, the present techniques further propose the provision of an instruction, forming part of the instruction set of the data processing pipeline (CPU), which can also be used to initiate various preparatory steps. Hence in some examples, the decoding circuitry is responsive to an extension setup instruction to generate the control signals to control the extension processing circuitry to perform at least one preparatory step to configure the extension processing circuitry for performance of the delegated task.
As in the case of the preparatory steps initiated by the confidence calibration circuitry, the preparatory steps initiated by the extension setup instruction can take a variety of forms. In some examples, the apparatus comprises multiple instances of extension processing circuitry, and wherein the extension setup instruction specifies a selected instance of extension processing circuitry to perform the at least one preparatory step.
In some examples, the extension setup instruction specifies an address identifier, wherein the address identifier is indicative of cache content which is to be brought into a private cache of the extension processing circuitry as the at least one preparatory step.
In some examples, the extension processing circuitry comprises a private address translation buffer and the at least one preparatory step comprises copying address translation information from a main address translation buffer of the data processing pipeline into the private address translation buffer.
The extension processing circuitry may be arranged in various ways to handle the saving of its state when it is caused to perform a context switch from a current execution context to a further execution context. In particular some approaches to state saving on a context switch are labelled “eager”, when essentially all extension
processing circuitry state is caused to be saved immediately when a context switch is triggered. Other approaches are labelled “lazy”, when the saving of extension processing circuitry state is delayed, only to be enforced when it is established that that state would otherwise be overwritten, e.g. by the actions of the incoming context. The present techniques further propose that, when the extension processing circuitry is configured for such lazy context saving, and an extension context save instruction has been executed on the data processing pipeline (CPU), the context saving may be brought forward and proactively performed as a preparation step for launching the extension processing circuitry. Hence in some examples, the decoding circuitry is responsive to an extension context save instruction comprising a storage location identifier, to generate the control signals to trigger the extension processing circuitry to perform a context switch from a current execution context to a further execution context, wherein the extension processing circuitry is configured to defer storing the extension state information to a location identified by the storage location identifier until a point at which it is determined that the further execution context requires the extension processing circuitry to modify the extension state information, and wherein the at least one preparatory step triggered in response to the extension setup instruction comprises storing the extension state information to the location identified by the storage location identifier.
The extension processing circuitry may be arranged to be clock-gated and/or power-gated when in an idle state, and preparing the extension processing circuitry for active processing may involve exiting the clock-gated and/or power-gated state. Hence in some examples, the extension processing circuitry is configured to be in at least one of a clock-gated state and/or a power-gated state when in an idle state, and wherein the at least one preparatory step triggered in response to the extension setup instruction comprises causing the extension processing circuitry to exit at least one of the clock-gated state and/or the power-gated state.
Where the extension processing circuitry may consume a non-trivial amount of power (relative to the remainder of the apparatus), it is proposed that in some examples exiting the clock-gated state and/or the power-gated state is performed incrementally.
For example, when powering up the extension processing circuitry as a preparation step for launch, sections of the extension processing circuitry which are each in the clockgated state and/or power-gated state can be caused to be exited to an active state sequentially. One advantage of doing so is that power supply stability for the apparatus may thereby be improved, since a large rate of change of current consumption by the extension processing circuitry is avoided, which could otherwise cause a drop in power supply voltage, adversely affecting other components of the apparatus.
In one example herein there is a (non-transitory) computer-readable medium to store computer-readable code for fabrication of the apparatus of any of the above examples.
In one example herein there is a method of operating an apparatus comprising: performing data processing operations in a data processing pipeline in dependence on a received sequence of instructions; decoding the received sequence of instructions in decoding circuitry of the data processing pipeline and generating control signals to control the data processing pipeline to perform the data processing operations; performing a delegated task in extension processing circuitry associated with the data processing pipeline in response to a delegation signal received from the data processing pipeline, wherein the delegated task is performed asynchronously to the data processing operations performed by data processing pipeline; performing speculative instruction execution in the data processing pipeline, whereby modifications of state of the apparatus resulting from execution of a speculatively executed instruction are not committed until a speculation confirmation is generated indicating that execution of the speculatively executed instruction was correct; in response to an extension start instruction specifying the delegated task, generating the control signals to control the data processing pipeline to issue the delegation signal to the extension processing circuitry to delegate the delegated task; in response to the delegation signal, commencing the delegated task in the extension processing circuitry before the speculation confirmation is generated for the extension start instruction; and
ensuring that no results generated by the delegated task are visible to the data processing pipeline until the speculation confirmation is generated for the extension start instruction.
In one example herein there is a computer program for controlling a host data processing apparatus to provide an instruction execution environment, the computer program comprising: data processing pipeline logic configured to perform data processing operations in dependence on a received sequence of instructions, wherein the data processing pipeline logic comprises decoding logic configured to decode the received sequence of instructions and to generate control signals to control the data processing pipeline logic to perform the data processing operations; and extension processing logic associated with the data processing pipeline logic and configured to perform a delegated task in response to a delegation signal received from the data processing pipeline logic, wherein the extension processing logic is configured to perform the delegated task asynchronously to the data processing operations performed by data processing pipeline logic, wherein the data processing pipeline logic is configured to perform speculative instruction execution, whereby modifications of state of the instruction execution environment resulting from execution of a speculatively executed instruction are not committed until a speculation confirmation is generated indicating that execution of the speculatively executed instruction was correct, wherein the decoding logic is responsive to an extension start instruction specifying the delegated task to generate the control signals to control the data processing pipeline logic to issue the delegation signal to the extension processing logic to delegate the delegated task, and wherein the extension processing logic is responsive to the delegation signal to commence the delegated task before the speculation confirmation is generated for the extension start instruction, whereby the extension processing logic is configured to ensure that no results generated by the delegated task are visible to the data processing pipeline logic until the speculation confirmation is generated for the extension start instruction.
Some particular embodiments are now described with reference to the figures.
Figure 1 schematically illustrates a data processing apparatus 10 according to some examples. The data processing apparatus 10 is schematically shown to have a pipelined configuration, which for the purposes of brevity and clarity is shown in a conceptual representation here. The illustrated pipeline stages comprise an instruction cache 11, a fetch stage 12, a decode stage 13, a micro-op cache 14, an issue stage 15, and a register access stage 16. A sequence of instructions is retrieved from memory (not shown) and cached in the instruction cache 11. The fetch stage 12 controls which instructions are retrieved as the sequence of instructions and these instructions are then decoded in the decode stage 13. This decoding essentially identifies the type of each instruction, as well as any further operands specified by the instruction, and generates control signals to control the remainder of the apparatus to perform the data processing operation(s) defined by the instruction. Decoding the instructions may comprise splitting an instruction into one or more micro-ops, and these micro-ops can be cached in the micro-op cache 14. The final stage of the pipeline before execution is the issue stage 15, where instructions (or micro-ops) are queued pending the availability of the register values they specify as operands and the corresponding functional unit of the data processing pipeline which will carry out the defined operation. Generally the data processing operation(s) defined by the instructions are carried out by the functional units that form part of the data processing pipeline, namely the load/store unit 17, the execute unit 18, and the execute unit 19. These latter execute units may for example be arithmetic logic units (ALUs), floating point units (FPUs), and so on. The functional units that form part of the data processing pipeline perform their data processing operations on data values which are provided from a set of registers (conceptually represented by the register access stage 16 in the figure) and result values of those data processing operations are returned to the set of registers. The load/store unit 17 is provided for the purpose of storing values from the set of registers to the memory system, of which only a level 1 cache 21 and a level 2 cache 22 are shown in the figure. The LI cache 21 is private to the data processing apparatus 10 and the L2 cache 22 may be shared with another data processing apparatus, when part of a wider data processing
system. The data processing apparatus 10 is also shown to comprise a branch unit 20, which monitors execution flow of the sequence of instructions and seeks to predict, based on previous execution history, whether a given branch will be taken or not. The predictions from the branch unit 20 inform the sequence of instructions caused to be fetched by the fetch stage 12.
The data processing apparatus 10 further comprises extension processing circuitry 23, which is provided to support efficient performance of one or more defined functions, which have been established to be impactful and ubiquitous for the data processing operations which this data processing apparatus 10 carries out. Example functions of this type have been found to include tasks or functions such as memcpy, memset, compression, encryption, and string processing, although the present techniques are not limited to these particular examples. The extension processing circuitry is closely associated with the data processing pipeline and is configured to perform the defined function (also referred to herein as a delegated task) in response to a delegation signal received from the data processing pipeline. The extension processing circuitry 23 is an example of a threadlet extension (TE) according to the present techniques. The sequence of operations it carries out to perform the defined function is referred to as a threadlet herein. The extension processing circuitry 23, although closely associated with the data processing pipeline, is configured to perform the delegated task asynchronously to the data processing operations performed by data processing pipeline. The data processing pipeline may also be referred to as the CPU herein. Threadlets are functions or collections of operations that can be executed asynchronously relative to other CPU activity once launched. The directive or command sent to the extension processing circuitry 23 to initiate the delegated task is generated in response to an extension start instruction defined for this purpose in the instruction set of the data processing pipeline. Thus, an extension start instruction progresses along the data processing pipeline in the manner that any other CPU instruction would, but when the decoding circuitry 13 identifies the extension start instruction it can signal directly to the extension processing circuitry 23. The close integration of the extension processing circuitry 23 with data processing pipeline is illustrated by the fact that the extension processing circuitry 23 has direct access to the load/ store unit 17, and thus it shares the
data processing pipeline’s path to memory. The extension processing circuitry 23 also has access to the set of registers 16, such that for example, the extension start instruction can specify one or more registers as operands, and the values from these registers are then passed directly to the extension processing circuitry 23 in association with the command sent to initiate the delegated task. Upon completion of the task, results of the delegated task can be provided as outputs and returned to the register values via an extension synchronisation instruction.
Figure 2 schematically illustrates a data processing apparatus 30 according to some examples. It will be noted that the arrangement of components of the data processing apparatus 30 is similar to that of the components of the data processing apparatus 10 shown in Figure 1. One difference is that whilst the data processing apparatus 10 of Figure 1 is intended to represent an in-order processor, the data processing apparatus 30 is an out-of-order processor. As one consequence of this the data processing pipeline of the data processing apparatus 30 comprises a rename stage 35, allowing the data processing apparatus 30 to vary the order in which it executes instructions of the sequence of instructions, such that they can be executed in an order dictated by when their operands become available, and the availability of functional units, rather than the order in which they appear in the sequence. The illustrated pipeline stages comprise an instruction cache 31, a fetch stage 32, a decode stage 33, a micro-op cache 34, the rename stage 35, an issue stage 35, and a register access stage 37. A sequence of instructions is retrieved from memory (not shown) and cached in the instruction cache 31. Instructions pass through the data processing pipeline in the manner described above with reference to the data processing apparatus 10 of Figure 1, with the further register renaming that is performed by the rename stage 35. The functional units of the data processing pipeline in this example are the load unit 38, the store unit 39, the FPU 41, the integer ALU 42, and the vector unit 43. The throughput of the FPU 41, the integer ALU 42, and the vector unit 43 is sufficient that a result cache 44 is provided an intermediary before results of their data processing are returned to the registers 37. A branch prediction unit 45 is also provided and its predictions inform the operation of the fetch stage 32.
The data processing apparatus 30 further comprises extension processing circuitry (“threadlet extension”) 49, which is provided to support efficient performance of one or more defined functions, which have been established to be impactful and ubiquitous for the data processing operations which this data processing apparatus 30 carries out. The extension processing circuitry 49 is closely associated with the data processing pipeline and is configured to perform the defined function in response to a delegation signal received from the data processing pipeline. In the example of Figure 2, this delegation signal is shown emanating from the issue queue stage 36. Notably, this is after the rename stage 35, such that the extension processing circuitry 49 can operate with respect to the physical registers of the set of registers 37 according to the same mapping of architectural registers used for the rest of the apparatus. As in the example of Figure 1, the data processing pipeline (instruction cache 31 through to the register read stage 37, the load / store units 38 and 39, and the functional units 41-45) may also be referred to as the CPU. The threadlet extension 49 operates asynchronously relative to other CPU activity once launched. The directive or command sent to the extension processing circuitry 49 to initiate the delegated task is generated in response to an extension start instruction defined for this purpose in the instruction set of the data processing pipeline. The close integration of the extension processing circuitry 49 with data processing pipeline also apparent in this example by the fact that the extension processing circuitry 49 has direct access to the load unit 38 and the store buffer 40, and thus it shares the data processing pipeline’s path to memory. The extension processing circuitry 49 also has access to the set of registers 37, such that for example, the extension start instruction can specify one or more registers as operands, and the values from these registers are then passed directly to the extension processing circuitry 49 in association with the command sent to initiate the delegated task. Note that the output of the branch prediction unit 45 is also provided to the extension processing circuitry 49. Upon completion of the task, results of the delegated task can be provided as outputs and returned to the register values via an extension synchronisation instruction.
Figure 3 schematically illustrates a data processing apparatus 50 according to some examples. This example provides a comparison to the examples of Figure 1 and Figure 2, in which examples the extension processing circuitry was closely embedded
with the data processing pipeline, to the extent that those instances of extension processing circuitry may be considered to be within the CPU. In the example apparatus 50 of Figure 3, the CPU 51 and the extension processing circuitry (threadlet extension) 52 are not as closely integrated. For example this is illustrated by the fact that each has its own path to memory, with an LI cache 53 private to the CPU 51 and an LI cache 54 private to the threadlet extension 52. They share the L2 cache 55. Nevertheless, the threadlet extension 52 remains tightly coupled to the CPU 51, and can be launched quickly when an extension start instruction is encountered in the CPU pipeline specifying the function this threadlet extension 52 performs. The threadlet extension 52 can get data directly from CPU registers at the start of its execution. Upon completion, it can return values via an extension synchronisation instruction. Figure 3 also shows the threadlet extension 52 as having its own private TLB 56, in which it can cache currently used address translations. As a preparatory step before or associated with the delegation signal, content from the TLB 57 in the CPU 51 can be copied into the private TLB 56 in order to pre-warm this cache before the threadlet begins operation.
Figure 4 is a state diagram illustrating an example set of states between which an extension processing circuitry (TE) transitions in some examples. Initially the TE is in an IDLE state 60. When an extension start (XSTART) instruction is encountered by the data processing pipeline, a delegation signal can cause the TE switches to the SETUP state 61. This may also require a signal indicating that the XSTART instruction has been committed to be asserted. In the SETUP state 61, certain actions necessary for preparing the TE can be performed, for example, in examples in which the TE has a separate path to memory (as in the case of Figure 3), one setup task is the transfer of relevant entries currently in the CPU’s TLB to a private TLB within the TE. This enables the TE to perform translations independently at a faster rate than if it were to rely entirely on the existing translation mechanism within the CPU. If the TE has been in a clock-gated or power-gated condition when in the IDLE state 60, the SETUP state 61 may also comprise the task of exiting the TE from that clock-gated or power-gated condition. Once the SETUP state 61 is complete the TE can switch to the RUNNING state 62. If the TE encounters a memory fault during its processing, it asserts a signal which will raise an interrupt within the CPU, causing it to stop executing the main thread and switch
to a handler. The TE switches to the INTERRUPTED state 63. The address generating the fault is placed in a special syndrome system register and a bit in the Program Status Register (PSR) will be set enabling the handler to quickly determine the source of the fault. Setting a bit in the PSR makes communicating the resumption of the threadlet straightforward, because the handler can reset the relevant bit in the SPSR and when the CP SR is restored from the SPSR during exception return, the TE can detect the resetting of this bit and resume executing. The TE will also switch to the INTERRUPTED state 63 if the main thread gets switched out, e.g. during a context-switch initiated by the operating system. In the INTERRUPTED state 63, the TE may be clock-gated or powergated, unless some other thread launches a new command directed at it or the associated thread returns resumes execution or the handler returns. The TE returns from the INTERRUPTED state 63 to the RUNNING state 62 via the RELOAD state 64 in which any context or state relevant to its execution, which was previously saved to memory, can be restored. This might be the case if another thread made use of a TE which was previously interrupted. Finally, when the extension reaches the end of the offloaded granule of computation (the delegated task) it moves to the IDLE state 60. The TE will advertise completion of the task, so that an extension synchronisation instruction (XSYNC) can pick up that “done” signal and, if required, provide a return value to a specified register. If the TE has any lingering data in its private caches it might also need to flush these entries upon completion.
An example of using threadlets is now set out. The programmer or compiler identifies functions whose execution in custom hardware (extension processing circuitry) satisfies the cost-benefit thresholds in their use-case. An instruction (such as XSTART) is used to launches a command within the designated CPU extension. An example use written in pseudo-code (for such an identified function “funcX”) is as follows: funcA () {
XSTART {x0 - x3}, #imm_op //funcX(a, b, c, d);
II
12
13
14
XSYNC xO, #imm_op
}
Thus, within the function funcA, the XSTART instruction initializes the CPU extension and transfers to the extension processing circuitry the parameters (a, b, c, d) for funcX, which are in registers xO, xl, x2, x3 respectively. The XSTART instruction in this example also specifies the immediate value #imm_op, which defines the specific function to be carried out. For example, whilst there might only be one instance of extension processing circuitry, it may be capable of performing more than one function, or at least more than one variant of a function, and the immediate value #imm_op can select the desired variant and/or function. In other examples there may be more than one instance of extension processing circuitry and the immediate value #imm_op can select between them. Depending on the setup, the extension could also automatically get a copy of relevant entries in the TLB. The extension processing circuitry then carries out the task required (funcX) and during its execution, the CPU is free to carry on executing other instructions II, 12, 13, 14, etc. At some point in the future, the CPU executes an extension synchronisation instruction (XSYNC) which automatically checks whether the extension has completed or not. If it has not, for some variants of the extension synchronisation instruction, the CPU will wait for the delegated task to complete. Other variants of the extension synchronisation instruction (e.g. the XSYNCS variant) allow the CPU can carry on executing other code (if there are alternative routines available or stop executing and wait for completion of the extension (typically if there is nothing else to execute in the interim).
Figure 5 schematically illustrates an extension start instruction being present on a predicted branch path during speculative instruction execution in accordance with some examples. The instruction flow shown in the upper part of the figure, wherein in
a sequence of instructions 100 a first conditional branch instruction (BRANCHI) is to be found. Accordingly, in dependence on the condition on which this branch instruction is conditional, further instruction execution will either follow the taken (T) path or the not taken (NT) path. In this example, the taken path leads to a further sequence of instructions 101 at the conclusion of which a further non-conditional branch instruction (BRANCH2) causes the instruction flow to jump back to the instruction which sequentially follows the first conditional branch instruction (BRANCHI). Thus it can be seen that the further sequence of instructions 101 is an additional function which is sometimes executed during the execution of the first sequence of instructions 100. Moreover it is to be noted that the further sequence of instructions 101 comprises an XSTART instruction, i.e. an extension start instruction for launching a threadlet. The data processing apparatus which executes the instructions shown in the upper part of Figure 5 is configured for speculative instruction execution, wherein one example of such speculation arises in the execution of conditional branch instructions. That is, the data processing apparatus does not wait until the condition on which the first conditional branch instruction (BRANCHI) depends to be definitively resolved, before speculatively continuing instruction execution along either the not-taken (NT) or the taken (T) path. This is done on the basis of a history of execution of this sequence of instructions, i.e. how frequently in the past, when BRANCHI is encountered at this point, is the outcome taken (T) or not-taken (NT). When the data processing apparatus has previously mainly taken one path rather than the other, the speculative execution will assume that the same will occur on this iteration. The data processing apparatus 110 is shown in the lower part of Figure 5, comprising a data processing pipeline 110, which is partly controlled by the predictions generated by the branch prediction unit 111, that is, the data processing pipeline 110 is arranged to speculatively execute instructions on the basis of those predictions. Figure 5 shows that the first conditional branch instruction (BRANCHI) is predicted as taken (T) by the branch prediction unit 111 and therefore on this basis the data processing pipeline 110 presses ahead with execution of the further sequence of instructions 101 before this speculation prediction is confirmed as correct. Where the further sequence of instructions 101 comprises an XSTART instruction, the data processing pipeline 110 speculatively executes this instruction, and according to the present techniques, this causes a delegation signal to be sent to the
extension processing circuitry 112, which commences the delegated task without waiting for a signal from the commit stage 113, indicating that the speculation was correct and that the XSTART instruction is or will be committed. The extension processing circuitry 112 also receives the speculation prediction generated by the branch prediction unit 111 and thus knows that (at least initially) the XSTART has been executed speculatively and therefore the extension processing circuitry ensures that no results 114 generated by the delegated task are visible outside the extension processing circuitry (as “output”) until the speculation confirmation is generated for the extension start instruction. Such sandboxed results may be held in a variety of storage structures in the extension processing circuitry 112, but the results 114 may for example represent one or more of a private data cache, data memory, registers and so on.
Figure 6 schematically illustrates a data processing apparatus 120 comprising confidence calibration circuitry which may embody various examples of the present techniques. Here the data processing pipeline is shown as the CPU 121, associated with which there is extension processing circuitry (threadlet extension) 122. These two processing components have their own memory path, and the CPU 121 has a private LI cache 123, whilst the extension processing circuitry 122 has a private LI cache 124. They share access to the memory system via the shared L2 cache 125. Each also has its own TLB, namely TLB 126 in the CPU 121 and TLB 130 in the threadlet extension 122. The CPU 121 is configured to perform speculative instruction execution on the basis of branch predictions generated by its branch prediction unit 127. As explained above with reference to Figure 5, by extension the threadlet extension 122 also performs some speculative function performance. That is, when the CPU 121 encounters an extension start instruction, a delegation signal is sent to the threadlet extension 122. However, whether and when the threadlet extension 122 commences performance of the delegated task further depends on the confidence calibration circuitry (predictor confidence calibration unit) 128. The PCCU 128 can modify or even entirely substitute the branch predictions generated by its branch prediction unit 127. The PCCU 128 thus signals a modified or substituted branch prediction (extension steering confidence information) to the threadlet extension 122. The PCCU 128 can base its signalling to the extension processing circuitry (threadlet extension) 122 on a range of factors. These may comprise
any of: factors specific to the extension processing circuitry; a relative size of the delegated task; an execution history of an instruction sequence portion comprising the extension start instruction; and/or a recent stalling history of the extension processing circuitry. Further, the PCCU 128 is also arranged to signal to the extension processing circuitry 122 that it should to perform at least one preparatory step to configure the extension processing circuitry for performance of the delegated task. That is, the PCCU 128 may not immediately cause the extension processing circuitry 122 to launch, but can cause the extension processing circuitry 122 to prepare for launch e.g. by exiting an idle state, powering up components, or prepopulating caches with data. For the latter purpose the PCCU 128 is also communicatively coupled to a prefetcher 129 and can thus trigger the retrieval of predefined data associated with a delegated task (e.g. using an address identifier). When there is more than one instance of extension processing circuitry in the apparatus, the PCCU 128 can specify a selected instance of extension processing circuitry to perform the at least one preparatory step. When the extension processing circuitry 122 begins speculative performance of a task, it does so ensuring that no results generated by the delegated task are visible outside the extension processing circuitry until speculation confirmation (spec result ok) signal is received from the CPU 121. In the event that the speculative execution of the XSTART instruction was incorrect, and a speculation cancellation signal (spec result cancelled) is received from the CPU 121, the extension processing circuitry 122 rolls back the results it has held internally, invalidating modified cache content and so on.
Figure 7 schematically illustrates confidence calibration circuitry modifying or substituting branch predictions generated in a data processing pipeline for use by extension processing circuitry in accordance with some examples. The data processing pipeline 150 comprises branch prediction 151. Predictions generated by the branch prediction 151 are received by the confidence calibration circuitry 152. The confidence calibration circuitry 152 modifies or substitutes the branch predictions for use of the extension processing circuitry 153. Figure 7 schematically illustrates three types of information on which the confidence calibration circuitry 152 can base its calculations. A first is a set of task information 154, which it holds. This may for example be an indication of the granularity (size) of a number of tasks which the extension processing
circuitry 153 can carry out, thus for any given task which delegated to the extension processing circuitry 153 and for which performance may begin speculatively, the confidence calibration circuitry 152 can modify the likelihood of that speculative start by the extension processing circuitry 153 in dependence on the task size. A second type of information held by the confidence calibration circuitry 152 is the path statistics 155. These are based on information received from the data processing pipeline 150, indicative of the frequency with which execution paths include an extension start instruction are followed. This may also form the basis for a modification of the likelihood of a speculative start to a delegated task triggered by a given extension start instruction. A third type of information held by the confidence calibration circuitry 152 is the stall history 156. This represents a history of how frequently the extension processing circuitry 153 has stalled, due to the limits of its internal storage being reached before the speculation confirmation signal was received. When the frequency of such stalls is too high, the likelihood of a speculative start to a delegated task can be reduced.
Figure 8 schematically illustrates an extension setup instruction causing extension processing circuitry to perform one or more preparatory steps in accordance with some examples. Here, the XSETUP instruction takes the form: XSETUP xO, #imm. Thus an XSETUP instruction 160 of this form, when decoded by the CPU’s decoder 161, causes the extension processing circuitry 162 to perform one or more preparatory steps in the expectation of soon being delegated a task by the CPU. The content of register xO and the immediate value #imm can be used to steer aspects of these preparatory steps. When there is more than one instance of extension processing circuitry in the apparatus, either the content of the register xO or the immediate value #imm can be used to indicate a particular instance of the extension processing circuitry which should carry out the preparatory steps. Alternatively or in addition, the content of the register xO or the immediate value #imm can be received by the extension processing circuitry 162 to steer the particular preparatory steps it carries out. As example preparatory steps the extension processing circuitry 162 can cause content of its TLB 164 to be copied in from the CPU’s TLB and/or can cause specified data to be prepopulated in its LI cache 165.
Figure 9 schematically illustrates an extension setup instruction causing extension processing circuitry to perform one or more preparatory steps in accordance with some examples. Here, the XSETUP instruction may additionally specify a register and or an immediate value as operands, although this is not essential. An XSETUP instruction 170 of this form, when decoded by the CPU’s decoder 171, causes the extension processing circuitry 172 to perform one or more preparatory steps in the expectation of soon being delegated a task by the CPU. Here the example is given of the extension processing circuitry 172 having context state 173 which is residual from a previous context in which the extension processing circuitry 172 executed. This context state has not yet been saved, since the extension processing circuitry 172 is configured in a lazy-save mode, whereby the context state 173 will only (usually) be saved when a new incoming context is about to modify the context state 173. However as a preparatory step triggered by the XSETUP instruction, the extension processing circuitry 172 is caused to save the context state 173 to a specified location 174 in memory 175. This preparatory saving of context state 173 can also be triggered by the PCCU 176.
Figure 10 schematically illustrates an extension setup instruction causing extension processing circuitry to perform one or more preparatory steps in accordance with some examples. Here, the XSETUP instruction may additionally specify a register and or an immediate value as operands, although this is not essential. An XSETUP instruction 180 of this form, when decoded by the CPU’s decoder 181, causes the extension processing circuitry 182 to perform one or more preparatory steps in the expectation of soon being delegated a task by the CPU. Here the example is given of the extension processing circuitry 182 being in an idle state in which it is either or both of clock-gated and power-gated. This is controlled by the clock control 183 and the power control 184. The powering up of the extension processing circuitry may be incremental, for example, the four components 185, 186, 187, and 188 in this case are not powered up simultaneously, but rather in sequence. To take just one example these components could be four memory banks, which have been powered down in the idle state. This exiting of a clock-gated and/or power-gated state can also be triggered by the PCCU 189.
Figure 11 is a flow diagram showing a sequence of steps that are taken in the method of some examples. The flow can be considered to begin at step 200, at which an instruction sequence is being executed in the data processing pipeline. Step 201 indicates that the instruction execution includes speculative instruction execution and at step 202 it is determined whether an extension start instruction has been encountered whilst speculatively executing. When this is not the case the flow simply loops through steps 200-202. However, when an extension start instruction is speculatively executed the flow proceeds to step 203 at which the extension processing circuitry commences the delegated task before the extension start instruction is committed, executing this delegated task asynchronously to the data processing pipeline. It is then determined at step 204 whether the extension start instruction (XSTART) has been committed. When it has, the flow proceeds to step 205, where it is determined if the delegated task is complete. When the delegated task is complete then at step 206 its results are available externally as outputs (i.e. by the remainder of the data processing apparatus other than the extension processing circuitry). An XSYNC instruction may then gather these outputs for use by the data processing pipeline in other data processing operations. The flow returns to step 200. Otherwise, if at step 205 it is determined that the task is not yet complete, the flow proceeds via step 206 and the delegated task continues, yet the partial results of the delegated task can be allowed to be visible externally as outputs (since the XSTART instruction has committed and therefore the delegated task is no longer speculative). Returning to a consideration of step 204, when the XSTART instruction has not committed, the flow proceeds to step 208, at which is determined whether the delegated task is complete. When the delegated task is not yet complete, the flow proceeds to step 209 and the delegated task continues, but its results are not visible externally to the extension processing circuitry. The flow then loops back via step 204 to determine if the XSTART instruction has now been committed. When the task is determined to be complete at step 208, the flow proceeds to step 210 and it is determined if the XSTART instruction has been cancelled. If it has not, the flow returns to step 204 to determine whether the XSTART has now been committed. However if at step 210 it is determined that the XSTART instruction has been cancelled, the flow
proceeds to step 211, at which the results of the delegated task are invalidated (and thus never visible outside the extension processing circuitry) and the flow returns to step 200.
Figure 12 schematically illustrates a simulator implementation that may be used. Whilst the earlier described embodiments implement the present invention in terms of apparatus and methods for operating specific processing hardware supporting the techniques concerned, it is also possible to provide an instruction execution environment in accordance with the embodiments described herein which is implemented through the use of a computer program. Such computer programs are often referred to as simulators, insofar as they provide a software based implementation of a hardware architecture. Varieties of simulator computer programs include emulators, virtual machines, models, and binary translators, including dynamic binary translators. Typically, a simulator implementation may run on a host processor 515, optionally running a host operating system 510, supporting the simulator program 505. In some arrangements, there may be multiple layers of simulation between the hardware and the provided instruction execution environment, and/or multiple distinct instruction execution environments provided on the same host processor. Historically, powerful processors have been required to provide simulator implementations which execute at a reasonable speed, but such an approach may be justified in certain circumstances, such as when there is a desire to run code native to another processor for compatibility or re-use reasons. For example, the simulator implementation may provide an instruction execution environment with additional functionality which is not supported by the host processor hardware, or provide an instruction execution environment typically associated with a different hardware architecture. An overview of simulation is given in “Some Efficient Architecture Simulation Techniques”, Robert Bedichek, Winter 1990 USENIX Conference, Pages 53 - 63.
To the extent that embodiments have previously been described with reference to particular hardware constructs or features, in a simulated embodiment, equivalent functionality may be provided by suitable software constructs or features. For example, particular circuitry may be implemented in a simulated embodiment as computer program logic. Similarly, memory hardware, such as a register or cache, may be
implemented in a simulated embodiment as a software data structure. In arrangements where one or more of the hardware elements referenced in the previously described embodiments are present on the host hardware (for example, host processor 515), some simulated embodiments may make use of the host hardware, where suitable.
The simulator program 505 may be stored on a computer-readable storage medium (which may be a non-transitory medium), and provides a program interface (instruction execution environment) to the target code 500 (which may include applications, operating systems and a hypervisor) which is the same as the interface of the hardware architecture being modelled by the simulator program 505. Thus, the program instructions of the target code 500 may be executed from within the instruction execution environment using the simulator program 505, so that a host computer 515 which does not actually have the hardware features of the apparatuses discussed above can emulate these features.
Concepts described herein may be embodied in computer-readable code for fabrication of an apparatus that embodies the described concepts. For example, the computer-readable code can be used at one or more stages of a semiconductor design and fabrication process, including an electronic design automation (EDA) stage, to fabricate an integrated circuit comprising the apparatus embodying the concepts. The above computer-readable code may additionally or alternatively enable the definition, modelling, simulation, verification and/or testing of an apparatus embodying the concepts described herein.
For example, the computer-readable code for fabrication of an apparatus embodying the concepts described herein can be embodied in code defining a hardware description language (HDL) representation of the concepts. For example, the code may define a register-transfer-level (RTL) abstraction of one or more logic circuits for defining an apparatus embodying the concepts. The code may define a HDL representation of the one or more logic circuits embodying the apparatus in Verilog, SystemVerilog, Chisel, or VHDL (Very High-Speed Integrated Circuit Hardware Description Language) as well as intermediate representations such as FIRRTL.
Computer-readable code may provide definitions embodying the concept using systemlevel modelling languages such as SystemC and SystemVerilog or other behavioural representations of the concepts that can be interpreted by a computer to enable simulation, functional and/or formal verification, and testing of the concepts.
Additionally or alternatively, the computer-readable code may define a low-level description of integrated circuit components that embody concepts described herein, such as one or more netlists or integrated circuit layout definitions, including representations such as GDSII. The one or more netlists or other computer-readable representation of integrated circuit components may be generated by applying one or more logic synthesis processes to an RTL representation to generate definitions for use in fabrication of an apparatus embodying the invention. Alternatively or additionally, the one or more logic synthesis processes can generate from the computer-readable code a bitstream to be loaded into a field programmable gate array (FPGA) to configure the FPGA to embody the described concepts. The FPGA may be deployed for the purposes of verification and test of the concepts prior to fabrication in an integrated circuit or the FPGA may be deployed in a product directly.
The computer-readable code may comprise a mix of code representations for fabrication of an apparatus, for example including a mix of one or more of an RTL representation, a netlist representation, or another computer-readable definition to be used in a semiconductor design and fabrication process to fabricate an apparatus embodying the invention. Alternatively or additionally, the concept may be defined in a combination of a computer-readable definition to be used in a semiconductor design and fabrication process to fabricate an apparatus and computer-readable code defining instructions which are to be executed by the defined apparatus once fabricated.
Such computer-readable code can be disposed in any known transitory computer-readable medium (such as wired or wireless transmission of code over a network) or non-transitory computer-readable medium such as semiconductor, magnetic disk, or optical disc. An integrated circuit fabricated using the computer-readable code may comprise components such as one or more of a central processing unit, graphics
processing unit, neural processing unit, digital signal processor or other components that individually or collectively embody the concept.
In brief overall summary, apparatuses, methods of data processing, computer programs, and computer-readable media are disclosed. A data processing pipeline performs data processing operations defined by a received sequence of instructions. Extension processing circuitry associated with the data processing pipeline performs a delegated task in response to a delegation signal received from the data processing pipeline, performing the delegated task asynchronously to the data processing pipeline. The data processing pipeline performs speculative instruction execution. In response to an extension start instruction, the data processing pipeline issues the delegation signal to the extension processing circuitry to delegate the delegated task and. the extension processing circuitry responds by commencing the delegated task before a speculation confirmation is generated for the extension start instruction. The extension processing circuitry ensures that no results generated by the delegated task are visible outside the extension processing circuitry until the speculation confirmation is generated for the extension start instruction.
In the present application, the words “configured to. . . ” are used to mean that an element of an apparatus has a configuration able to carry out the defined operation. In this context, a “configuration” means an arrangement or manner of interconnection of hardware or software. For example, the apparatus may have dedicated hardware which provides the defined operation, or a processor or other processing device may be programmed to perform the function. “Configured to” does not imply that the apparatus element needs to be changed in any way in order to provide the defined operation.
Although illustrative embodiments have been described in detail herein with reference to the accompanying drawings, it is to be understood that the invention is not limited to those precise embodiments, and that various changes, additions and modifications can be effected therein by one skilled in the art without departing from the scope of the invention as defined by the appended claims. For example, various
combinations of the features of the dependent claims could be made with the features of the independent claims without departing from the scope of the present invention.
Claims
1. Apparatus for data processing, comprising: a data processing pipeline configured to perform data processing operations in dependence on a received sequence of instructions, wherein the data processing pipeline comprises decoding circuitry configured to decode the received sequence of instructions and to generate control signals to control the data processing pipeline to perform the data processing operations; and extension processing circuitry associated with the data processing pipeline and configured to perform a delegated task in response to a delegation signal received from the data processing pipeline, wherein the extension processing circuitry is configured to perform the delegated task asynchronously to the data processing operations performed by the data processing pipeline, wherein the data processing pipeline is configured to perform speculative instruction execution, whereby modifications of state of the apparatus resulting from execution of a speculatively executed instruction are not committed until a speculation confirmation is generated indicating that execution of the speculatively executed instruction was correct, wherein the decoding circuitry is responsive to an extension start instruction specifying the delegated task to generate the control signals to control the data processing pipeline to issue the delegation signal to the extension processing circuitry to delegate the delegated task, and wherein the extension processing circuitry is responsive to the delegation signal to commence the delegated task before the speculation confirmation is generated for the extension start instruction, whereby the extension processing circuitry is configured to ensure that no results generated by the delegated task are visible outside the extension processing circuitry until the speculation confirmation is generated for the extension start instruction.
2. The apparatus as claimed in claim 1, wherein the data processing pipeline is configured to generate a speculation cancellation when it is determined that execution of the speculatively executed instruction was incorrect,
and wherein the extension processing circuitry is responsive to the speculation cancellation generated for the extension start instruction to invalidate any result generated by the delegated task.
3. The apparatus as claimed in claim 1 or claim 2, wherein the extension processing circuitry comprises a private data store configured to hold results generated by the delegated task.
4. The apparatus as claimed in claim 3, wherein the extension processing circuitry is responsive to results generated by the delegated task reaching a result storage capacity of the private data store to stall performance of the delegated task.
5. The apparatus as claimed in any preceding claim, wherein the data processing pipeline comprises branch prediction circuitry configured to generate branch confidence information indicative of a predicted likelihood of a branch direction being taken, wherein the data processing pipeline is configured to perform the speculative instruction execution in dependence on the branch confidence information, and wherein the apparatus further comprises confidence calibration circuitry configured to generate extension steering confidence information associated with the extension start instruction, wherein the extension processing circuitry is configured to determine whether to commence the delegated task before the speculation confirmation is generated for the extension start instruction in dependence on the extension steering confidence information.
6. The apparatus as claimed in claim 5, wherein the confidence calibration circuitry is configured to generate the extension steering confidence information in dependence on the branch confidence information and in dependence on at least one extension processing circuitry specific factor.
7. The apparatus as claimed in claim 5 or claim 6, wherein the confidence calibration circuitry configured to generate the extension steering confidence information in dependence on a relative size of the delegated task.
8. The apparatus as claimed in any of claims 5-7, wherein the confidence calibration circuitry is configured to generate the extension steering confidence information in dependence on an execution history of an instruction sequence portion comprising the extension start instruction.
9. The apparatus as claimed in any of claims 5-8, wherein the confidence calibration circuitry is configured to generate the extension steering confidence information in dependence on a stalling history of the extension processing circuitry.
10. The apparatus as claimed in any of claims 5-9, wherein the confidence calibration circuitry is configured to cause, in dependence on the extension steering confidence information, the extension processing circuitry to perform at least one preparatory step to configure the extension processing circuitry for performance of the delegated task.
11. The apparatus as claimed in claim 10, comprising multiple instances of extension processing circuitry, wherein the confidence calibration circuitry is configured to specify a selected instance of extension processing circuitry to perform the at least one preparatory step.
12. The apparatus as claimed in claim 10 or claim 11, wherein the confidence calibration circuitry is configured specify an address identifier, wherein the address identifier is indicative of cache content which is to be brought into a private cache of the extension processing circuitry as at least a part of the at least one preparatory step.
13. The apparatus as claimed in any of claims 5-10, wherein the decoding circuitry is responsive to an extension setup instruction to generate the control signals to control
the extension processing circuitry to perform at least one preparatory step to configure the extension processing circuitry for performance of the delegated task.
14. The apparatus as claimed in claim 13, comprising multiple instances of extension processing circuitry, and wherein the extension setup instruction specifies a selected instance of extension processing circuitry to perform the at least one preparatory step.
15. The apparatus as claimed in claim 13 or claim 14, wherein the extension setup instruction specifies an address identifier, wherein the address identifier is indicative of cache content which is to be brought into a private cache of the extension processing circuitry as the at least one preparatory step.
16. The apparatus as claimed in any of claims 10-15, wherein the extension processing circuitry comprises a private address translation buffer and the at least one preparatory step comprises copying address translation information from a main address translation buffer of the data processing pipeline into the private address translation buffer.
17. The apparatus as claimed in any of claims 10-16, wherein the decoding circuitry is responsive to an extension context save instruction comprising a storage location identifier, to generate the control signals to trigger the extension processing circuitry to perform a context switch from a current execution context to a further execution context, wherein the extension processing circuitry is configured to defer storing the extension state information to a location identified by the storage location identifier until a point at which it is determined that the further execution context requires the extension processing circuitry to modify the extension state information, and wherein the at least one preparatory step triggered in response to the extension setup instruction comprises storing the extension state information to the location identified by the storage location identifier.
18. The apparatus as claimed in any of claims 10-17, wherein the extension processing circuitry is configured to be in at least one of a clock-gated state and/or a power-gated state when in an idle state, and wherein the at least one preparatory step triggered in response to the extension setup instruction comprises causing the extension processing circuitry to exit at least one of the clock-gated state and/or the power-gated state.
19. The apparatus as claimed in claim 18, wherein exiting the clock-gated state and/or the power-gated state is performed incrementally.
20. A non-transitory computer-readable medium to store computer-readable code for fabrication of the apparatus of any of claims 1 to 19.
21. A method of operating an apparatus comprising: performing data processing operations in a data processing pipeline in dependence on a received sequence of instructions; decoding the received sequence of instructions in decoding circuitry of the data processing pipeline and generating control signals to control the data processing pipeline to perform the data processing operations; performing a delegated task in extension processing circuitry associated with the data processing pipeline in response to a delegation signal received from the data processing pipeline, wherein the delegated task is performed asynchronously to the data processing operations performed by data processing pipeline; performing speculative instruction execution in the data processing pipeline, whereby modifications of state of the apparatus resulting from execution of a speculatively executed instruction are not committed until a speculation confirmation is generated indicating that execution of the speculatively executed instruction was correct; in response to an extension start instruction specifying the delegated task, generating the control signals to control the data processing pipeline to issue the delegation signal to the extension processing circuitry to delegate the delegated task;
in response to the delegation signal, commencing the delegated task in the extension processing circuitry before the speculation confirmation is generated for the extension start instruction; and ensuring that no results generated by the delegated task are visible to the data processing pipeline until the speculation confirmation is generated for the extension start instruction.
22. A computer program for controlling a host data processing apparatus to provide an instruction execution environment, the computer program comprising: data processing pipeline logic configured to perform data processing operations in dependence on a received sequence of instructions, wherein the data processing pipeline logic comprises decoding logic configured to decode the received sequence of instructions and to generate control signals to control the data processing pipeline logic to perform the data processing operations; and extension processing logic associated with the data processing pipeline logic and configured to perform a delegated task in response to a delegation signal received from the data processing pipeline logic, wherein the extension processing logic is configured to perform the delegated task asynchronously to the data processing operations performed by data processing pipeline logic, wherein the data processing pipeline logic is configured to perform speculative instruction execution, whereby modifications of state of the instruction execution environment resulting from execution of a speculatively executed instruction are not committed until a speculation confirmation is generated indicating that execution of the speculatively executed instruction was correct, wherein the decoding logic is responsive to an extension start instruction specifying the delegated task to generate the control signals to control the data processing pipeline logic to issue the delegation signal to the extension processing logic to delegate the delegated task, and wherein the extension processing logic is responsive to the delegation signal to commence the delegated task before the speculation confirmation is generated for the extension start instruction,
whereby the extension processing logic is configured to ensure that no results generated by the delegated task are visible to the data processing pipeline logic until the speculation confirmation is generated for the extension start instruction.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GB2308384.3A GB2630754B (en) | 2023-06-05 | 2023-06-05 | Extension processing circuitry start-up |
| PCT/GB2024/050352 WO2024252114A1 (en) | 2023-06-05 | 2024-02-09 | Extension processing circuitry start-up |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4720853A1 true EP4720853A1 (en) | 2026-04-08 |
Family
ID=87156834
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24706783.8A Pending EP4720853A1 (en) | 2023-06-05 | 2024-02-09 | Extension processing circuitry start-up |
Country Status (7)
| Country | Link |
|---|---|
| EP (1) | EP4720853A1 (en) |
| KR (1) | KR20260018061A (en) |
| CN (1) | CN121241331A (en) |
| GB (1) | GB2630754B (en) |
| IL (1) | IL324757A (en) |
| TW (1) | TW202449595A (en) |
| WO (1) | WO2024252114A1 (en) |
Family Cites Families (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8566568B2 (en) * | 2006-08-16 | 2013-10-22 | Qualcomm Incorporated | Method and apparatus for executing processor instructions based on a dynamically alterable delay |
| US8521961B2 (en) * | 2009-08-20 | 2013-08-27 | International Business Machines Corporation | Checkpointing in speculative versioning caches |
| CN104583957B (en) * | 2012-06-15 | 2018-08-10 | 英特尔公司 | With the speculative instructions sequence without the rearrangement for disambiguating out of order load store queue |
| US9495159B2 (en) * | 2013-09-27 | 2016-11-15 | Intel Corporation | Two level re-order buffer |
| GB2519108A (en) * | 2013-10-09 | 2015-04-15 | Advanced Risc Mach Ltd | A data processing apparatus and method for controlling performance of speculative vector operations |
| GB2570110B (en) * | 2018-01-10 | 2020-04-15 | Advanced Risc Mach Ltd | Speculative cache storage region |
| US10942738B2 (en) * | 2019-03-29 | 2021-03-09 | Intel Corporation | Accelerator systems and methods for matrix operations |
| US12073251B2 (en) * | 2020-12-29 | 2024-08-27 | Advanced Micro Devices, Inc. | Offloading computations from a processor to remote execution logic |
| US11550585B2 (en) * | 2021-03-23 | 2023-01-10 | Arm Limited | Accelerator interface mechanism for data processing system |
-
2023
- 2023-06-05 GB GB2308384.3A patent/GB2630754B/en active Active
-
2024
- 2024-02-09 EP EP24706783.8A patent/EP4720853A1/en active Pending
- 2024-02-09 CN CN202480036905.6A patent/CN121241331A/en active Pending
- 2024-02-09 WO PCT/GB2024/050352 patent/WO2024252114A1/en not_active Ceased
- 2024-02-09 KR KR1020257042356A patent/KR20260018061A/en active Pending
- 2024-02-09 IL IL324757A patent/IL324757A/en unknown
- 2024-03-08 TW TW113108576A patent/TW202449595A/en unknown
Also Published As
| Publication number | Publication date |
|---|---|
| TW202449595A (en) | 2024-12-16 |
| GB2630754A (en) | 2024-12-11 |
| WO2024252114A1 (en) | 2024-12-12 |
| GB202308384D0 (en) | 2023-07-19 |
| GB2630754B (en) | 2025-09-24 |
| KR20260018061A (en) | 2026-02-06 |
| CN121241331A (en) | 2025-12-30 |
| IL324757A (en) | 2026-01-01 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11972264B2 (en) | Micro-operation supply rate variation | |
| US20230120596A1 (en) | Responding to branch misprediction for predicated-loop-terminating branch instruction | |
| WO2024252114A1 (en) | Extension processing circuitry start-up | |
| WO2024252115A1 (en) | Task delegation | |
| EP4720851A1 (en) | Linking delegated tasks | |
| EP4720847A1 (en) | Maintaining state information | |
| WO2024252111A1 (en) | Triggering execution of an alternative function | |
| GB2630749A (en) | Hazard-checking in task delegation | |
| KR20260017396A (en) | Memory handling with delegated tasks | |
| GB2639993A (en) | Validation of integrity confirmation information |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251222 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |