EP3286647A1 - Placement d'une tâche de calcul sur un processeur fonctionnellement asymetrique - Google Patents
Placement d'une tâche de calcul sur un processeur fonctionnellement asymetriqueInfo
- Publication number
- EP3286647A1 EP3286647A1 EP16711185.5A EP16711185A EP3286647A1 EP 3286647 A1 EP3286647 A1 EP 3286647A1 EP 16711185 A EP16711185 A EP 16711185A EP 3286647 A1 EP3286647 A1 EP 3286647A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- processor
- instructions
- core
- hardware
- task
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30003—Arrangements for executing specific machine instructions
- G06F9/30007—Arrangements for executing specific machine instructions to perform operations on data operands
- G06F9/3001—Arithmetic instructions
- G06F9/30014—Arithmetic instructions with variable precision
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/46—Multiprogramming arrangements
- G06F9/48—Program initiating; Program switching, e.g. by interrupt
- G06F9/4806—Task transfer initiation or dispatching
- G06F9/4843—Task transfer initiation or dispatching by program, e.g. task dispatcher, supervisor, operating system
- G06F9/485—Task life-cycle, e.g. stopping, restarting, resuming execution
- G06F9/4856—Task life-cycle, e.g. stopping, restarting, resuming execution resumption being on a different machine, e.g. task migration, virtual machine migration
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F15/00—Digital computers in general; Data processing equipment in general
- G06F15/76—Architectures of general purpose stored program computers
- G06F15/80—Architectures of general purpose stored program computers comprising an array of processing units with common control, e.g. single instruction multiple data processors
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30003—Arrangements for executing specific machine instructions
- G06F9/30007—Arrangements for executing specific machine instructions to perform operations on data operands
- G06F9/30021—Compare instructions, e.g. Greater-Than, Equal-To, MINMAX
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30003—Arrangements for executing specific machine instructions
- G06F9/30076—Arrangements for executing specific machine instructions to perform miscellaneous control operations, e.g. NOP
- G06F9/30083—Power or thermal control instructions
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30098—Register arrangements
- G06F9/30101—Special purpose registers
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/3017—Runtime instruction translation, e.g. macros
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/46—Multiprogramming arrangements
- G06F9/48—Program initiating; Program switching, e.g. by interrupt
- G06F9/4806—Task transfer initiation or dispatching
- G06F9/4843—Task transfer initiation or dispatching by program, e.g. task dispatcher, supervisor, operating system
- G06F9/4881—Scheduling strategies for dispatcher, e.g. round robin, multi-level priority queues
- G06F9/4893—Scheduling strategies for dispatcher, e.g. round robin, multi-level priority queues taking into account power or heat criteria
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/46—Multiprogramming arrangements
- G06F9/50—Allocation of resources, e.g. of the central processing unit [CPU]
- G06F9/5005—Allocation of resources, e.g. of the central processing unit [CPU] to service a request
- G06F9/5027—Allocation of resources, e.g. of the central processing unit [CPU] to service a request the resource being a machine, e.g. CPUs, Servers, Terminals
- G06F9/5044—Allocation of resources, e.g. of the central processing unit [CPU] to service a request the resource being a machine, e.g. CPUs, Servers, Terminals considering hardware capabilities
-
- Y—GENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
- Y02—TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
- Y02D—CLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
- Y02D10/00—Energy efficient computing, e.g. low power processors, power management or thermal management
Definitions
- the invention relates to multi-core processors in general and the management of the placement of calculation tasks on multi-core processors that are functionally asymmetrical in particular.
- a multi-core processor may include one or more hardware extensions, intended to accelerate parts of very specific software codes that are difficult to parallelize.
- these hardware extensions may include circuits for floating point computation or vector computation.
- a multi-core processor is said to be "functionally asymmetric" when certain extensions fail some processor cores, i.e. when at least one core of the multi-core processor does not have the hardware extension required for execution of a given instruction set.
- a functionally asymmetric processor can be characterized by unequal distribution (or association) of extensions to the cores of processors. Managing a functionally asymmetric multi-core processor poses several technical problems. One of these technical problems is to effectively manage the placement of computing tasks on different processor cores.
- Patent document WO2013101139 entitled "PROVIDING AN ASYMMETRIC MULTICORE PROCESSOR SYSTEM TRANSPARENTLY TO AN OPERATING SYSTEM” discloses a system comprising a multi-core processor with several groups of cores.
- the second group may be of an instruction set architecture (ISA) different from the first group, or of the same ISA architecture to be defined but with a different level of power and performance of media.
- the processor further includes a migration unit that processes the migration requests for a number of different scenarios and causes a context switch to dynamically migrate a process from a second core to a first core of the first group. This change of material context dynamic can be transparent to the operating system.
- the present invention relates to a method for managing a computing task on a functionally asymmetric multi-core processor, wherein at least one core of said processor is associated with one or more hardware extensions, the method comprising the steps of receiving a task of calculation associated with instructions executable by a hardware extension; receive calibration data associated with the hardware extension; and determining an opportunity cost of performing the computing task based on the calibration data.
- Developments describe the determination of the calibration data, in particular by counting or by calculation (online and / or offline) of the classes of instructions executed, the execution of a predefined set of instructions representative of the execution space of extension, consideration of energy and temperature aspects, translation or emulation of instructions or placement of calculation tasks on different cores. System and software aspects are described.
- the invention is implemented at the hardware level (hardware) and the operating system.
- the invention retrieves and analyzes certain information of the last quanta of execution (time allocated by the scheduler to the task on a heart) and thus estimates the cost future task placement on different types of hearts.
- the task placement is flexible and transparent for the user.
- the method according to the invention performs the placement of the calculation tasks according to more diversified and global objectives than the known solutions of the state of the art which stick to the instructions associated with the calculation tasks.
- the placement of computation tasks involving the switching on or off of one or more cores is achieved by considering only the "strict" use of the extension.
- a core that does not include extensions can not execute instructions for one or more particular extensions.
- current solutions examine whether an extension is used (e.g. application placement on an extended core) or not used (e.g. placement on a core without an extension).
- the known approaches consider the presence of call of such extensions within the source code but proceed to the placement of the tasks exclusively according to the code instructions and ignore in particular any other criterion, including energy.
- Other known approaches are based on an analysis of the source code and the planned use of hardware extensions, proceed to code mutations, but without estimating criteria including, for example, degradation of performance or energy.
- the invention can meet the so-called "multi-objective" needs to be met by a task scheduler such as a) performance, b) energy efficiency and c) thermal stresses ("dark-silicon").
- a task scheduler such as a) performance, b) energy efficiency and c) thermal stresses ("dark-silicon”).
- Certain embodiments make it possible to increase the number of cores of a multi-core processor, while keeping a performance / energy / surface efficiency. The additional programming effort remains small.
- the prediction of the costs and the performance gains of a given task on each of the different cores of a heterogeneous system can be performed dynamically.
- the method according to the invention takes into account dynamic criteria related to the execution of a program.
- a hardware extension by a software code is dictated either explicitly by the programmer of the code, or automatically performed by the compiler.
- the compilation of a software being most often done offline, the programmer or the compiler, can not base his choice to exploit or not an extension only on criteria which are linked to the software itself, and not on dynamic criteria for program execution.
- Without information about the execution environment workload, instant task scheduling, availability of resources), it is usually impossible to determine in advance whether or not to use a given hardware extension.
- the method according to the invention makes it possible to predict and / or estimate the use of one or more hardware extensions.
- a dynamic scheduling and task placement strategy removing various constraints such as (i) predictability and interoperability in the use of constrained heterogeneous systems (functional asymmetry with common base ) and (ii) optimization against global objectives of the system (eg performance, consumption and surface), enabled by a fast, dynamic and transparent prediction of the execution of a given code on a given core .
- certain embodiments of the invention include prediction steps as to the use of the extensions, which can advantageously optimize the energy efficiency and optimize the computing performance. These embodiments may in particular leave the scheduler a certain freedom as to the choice of cores on which to execute the different program phases.
- the method according to the invention makes it possible to optimize the energy consumption, including in the case of a weak use of an extension.
- the method according to the invention makes it possible to place a task on a processor, whether this processor is associated with a hardware extension or not, depending on the implementation of the task (thus with or without exploitation of the extension). As a result, a scheduler can dynamically optimize energy or performance without worrying about the initial implementation of the task.
- the invention makes it possible to predict or dynamically determine the advantage of using one or more hardware extensions and to place the calculation tasks (ie to allocate these tasks to the different processors and or core processors) according to of this prediction.
- the process according to the invention makes it possible to delay the whether or not to use a run-time extension, gives the programmer a higher level of abstraction, and gives the system software increased decision-making freedom for scheduling optimization. Once this flexibility is achieved, the quality of system software decisions becomes solely dependent on the execution environment. To do this, the system software measures the relevant variables of the runtime environment.
- the method according to the invention makes it possible to have task migration freedoms and accumulated knowledge concerning the execution environment in order to optimize the placement of the calculation tasks on one or more asymmetrically functional multicore processors, and this according to the overall objectives of the scheduler.
- the method according to the invention confers freedom of action with regard to the placement of calculation tasks.
- Such freedom of action allows the system to migrate tasks (computation) on any processor core, and this despite the different extensions present in each of these cores.
- the software applications running on the system are associated with a more continuous and flexible exploration space to achieve the objectives (or multi-objectives) in terms of power (eg thermal performance).
- embodiments of the invention may be implemented for "System-on-Chip" (SoC) for consumer-type or on-board electronic applications (eg phones, embedded components, desktop, Internet of Things, etc.) .
- SoC System-on-Chip
- electronic applications eg phones, embedded components, desktop, Internet of Things, etc.
- the use of heterogeneous systems is common to optimize computational efficiency for a workload specific. Since the latter systems are becoming increasingly complex (eg increasing the number of cores and workloads), certain embodiments of the invention make it possible to reduce the impact of this complexification, disregarding the heterogeneity of the systems.
- the method according to the invention allows a dynamic optimization of the performance and the energy thanks to the estimator of the degradation and the unit of prediction.
- the scalability of multi-core processors is improved.
- the asymmetry management of the platform is done transparently for the user, that is to say does not increase the effort or constraints of software development.
- an optimization of the scheduling makes it possible to better answer the technical problems which increase with the number of cores such as the reduction of the surface and the energy consumed, the priority of the tasks, the dark-silicon (temperature).
- the placement / scheduling of tasks according to the invention is flexible.
- the scheduling is done in an optimized and transparent manner for the user.
- the method according to the invention provides for the use of prediction units ("predictor") and / or interest estimators, allowing optimized task placement.
- the compiler can use an extension by planning that it will accelerate the execution, but because of the additional memory displacement (eg between the extension and the register of the "basic" cores), the performance will actually be depreciated.
- the choice of job placement according to the method is flexible.
- the objective of the operating system is not always to optimize the performance, the objective can be a minimization of the energy consumed or a placement optimizing the temperature of the cores ("dark silicon") . Due to dependency on the instruction set, the scheduler may be forced to execute an application according to its resource requirements and not considering a overall objective.
- the placement is independent of other parameters supervised by the scheduler.
- a scheduler typically allocates a quanta of time to each compute task on a processor. If the calculation of a task is not completed, another quanta will be associated with it. The size of these quanta is variable to allow a fair and optimized sharing of different tasks between different processors (typically between 0.1 and 100ms).
- the quantum dimensioning involves a more or less fine detection of the basic type phases. There may be edge effects (eg, ceaseless migrations). Sticking to these edge effects ultimately reduces investment flexibility and optimizations. Description of figures
- FIG. 1 illustrates examples of processor architectures
- FIGS. 2A and 2B illustrate some aspects of the invention in terms of energy efficiency, depending on whether hardware extensions are used or not;
- Figure 3 illustrates examples of architectures and job placement
- Figure 4 provides examples of steps of the method according to the invention. Detailed description of the invention
- the invention generally allows optimized placement of computational tasks on a functionally asymmetric multi-core processor.
- a functionally asymmetric multi-core processor comprises programmable elements or processor cores, using more or less extensive functionalities.
- An “extended” core ie, with one or more hardware extensions
- a “hardware extension” or “hardware extension” is a circuit such as an FPU floating data computing unit, a vector computing unit, a SIMD, a cryptographic processing unit, a communication unit. signal processing, etc.
- a hardware extension introduces a dedicated hardware circuit accessible or connected to a processor core, which circuit provides high performance for specific computing tasks. These specific circuits improve the performance and energy efficiency of a core for particular calculations, but their intensive use can lead to reduced performance in terms of Watt per unit area.
- These hardware extensions associated with the core processor are provided with an instruction set that extends the standard or default (ISA) set. Hardware extensions are usually well integrated into the core pipeline, allowing efficient access to functions through instructions added to the "basic" set (by comparison, a specialized "coprocessor” usually requires protocol-like instructions.
- a calculation task (or "thread” in English) includes instructions, which may be grouped into instruction sequences (temporally) or games or classes of instructions (by nature).
- the term “computational task” (or “task”) refers to a “thread” or “thread”.
- Other denominations include phrases such as “light process”, “processing unit”, “execution unit”, “instruction thread”, “lightened process” or “exetron”.
- the expression designates the execution of a set of instructions of the machine language of a processor. From the user's point of view, these executions seem to run in parallel. However, where each process has its own virtual memory, threads in the same process share virtual memory. However, all threads have their own call stack.
- a computational task does not necessarily use a hardware extension: in the most common case, a compute task is executed using instructions common to all processor cores.
- a computation task can (optionally) be executed by a hardware extension (which nevertheless requires the associated core to "decode" the instructions).
- the function of a hardware extension is to speed up the processing of a specific instruction set: a hardware extension may not speed up the processing of another type of instruction (eg floating point versus integer).
- a processor core can be associated with none (ie zero) or one or more hardware extensions. These hardware extensions are then "exclusive" (a given hardware extension can not be accessed from a third core).
- the processor core includes the hardware extension (s).
- the physical circuits of a hardware extension may include a processor core.
- a processor core is a a set of physical circuits capable of executing programs "autonomously".
- a hardware extension is able to execute a part of the program but is not “autonomous" (it requires the association with at least one processor core).
- a processor core may have all the hardware extensions required to execute the instructions contained in a given task.
- a processor core may also not necessarily have all the hardware extensions required to execute the instructions included in the computational task: the technical problem of moving and / or emulating (ie, functionality or instruction ) - and associated costs - arise.
- Hardware extensions may be of a different nature. An extension can process a set of instructions specific to it.
- Hardware expansion is usually expensive in circuit area and static energy.
- An asymmetric platform compared to a symmetric processor, reduces surface and energy costs by reducing the number of extensions and by turning on or off cores containing extensions only when necessary or advantageous.
- a processor core accesses exclusively one or more hardware extensions dedicated thereto.
- an extension may be "shared" between multiple processor cores.
- the computational load should be distributed as well as possible. For example, it is advantageous not to overload an extension that would be shared.
- the operating system through the scheduler provides quanta of time, each quantum of time being allocated to a given software application.
- the scheduler (or "scheduler" in English) is preemptive type (ie the quanta of time are imposed on software applications).
- a method for managing a computing task on a functionally asymmetric multi-core processor, wherein at least one core of said processor is associated with one or more hardware extensions, the method comprising the steps receiving a computing task, said computing task being associated with executable instructions by a hardware extension associated with the multi-core processor; receive calibration data associated with said hardware extension; and determining an opportunity cost of performing the computing task based on the calibration data.
- One or more opportunity costs of execution can be determined. At least one processor core among the plurality of cores is thus characterized by an opportunity cost of executing the received computation task.
- the method according to the invention considers various "opportunities" of placement, i.e. possibilities or potentialities, which are analyzed.
- the computation task comprises instructions associated with one or more predefined classes of instructions and the hardware extension is associated with one or more predefined classes of instructions, said classes being executable by said extension.
- the calibration data includes coefficients indicative of a unit execution cost per instruction class, said coefficients being determined by comparison between the execution of a predefined set of instructions representative of the execution space of said extension on said hardware extension of on the one hand and the execution of said predefined set of instructions on a processor core without hardware extension on the other hand.
- the "predefined set of instructions representative of the execution space” aims to represent the various possibilities of execution of the instructions, in the most complete i.e. way as exhaustive as possible. With regard to the nature of the instructions (i.e. the classes or types of instructions), completeness can be achieved. On the other hand, the combinatorics of the sequences of the different instructions being virtually infinite, the representation is necessarily imperfect. It can nevertheless be approached asymptotically.
- the principle of executing a set of instructions makes it possible to determine "control points" of the method according to the invention (e.g., effective calibration data). Experimentally, a hundred “programs” were conducted, each containing "tests" (unitary executions) and the execution of software applications in real conditions.
- the unit tests aimed to execute a small number of instruction types, mainly floating point, control and / or memory. Different benchmarks have been used (eg MiBench, SDVBS, WCET, fbench, Polybench).
- Execution of actual software applications has aimed to represent the main sequences in the execution of instructions.
- software applications used different data sets.
- each software application can be compiled for a processor core with hardware extension and also compiled for a core without extension. Both versions of binaries are then executed on the different cores. The difference in execution time is determined as well as the number and nature of the classes of instructions executed. For a fairly large set of applications, it is possible to determine a scatter plot representing the total execution space. As a result, it is possible to correlate the number of instructions associated with each class of instructions and the difference in execution time between cores with or without extension.
- the number of software applications can be increased.
- this representativeness can be determined (use of random instructions, thresholds, confidence intervals, cloud distribution, etc.).
- a minimum (or optimal) number of software applications to be executed under real conditions can be determined. Comparing the execution of instructions on a core with extension and on a core without extension, some results have highlighted deviations of the order of 10 to 20% between the performances as estimated according to the process and the performances actually measured. This order of magnitude reflects the possibility of operational implementation of the invention (and this without additional optimization).
- the method further comprises a step of determining a number of uses of each instruction class associated with the compute task by said hardware extension.
- the step of determining the number of uses of each instruction class includes a step counting the number of uses of each class of instructions.
- the history ie the past is used to estimate or evaluate the future.
- the step of determining the number of uses of each instruction class includes a step of estimating the number of uses of each instruction class, including from uses counted in the past. It is also possible to estimate the number of uses from other methods. It is also possible to combine counting and estimating the number of uses of instruction classes.
- the opportunity cost of execution is determined by summation indexed by instruction class of the coefficients per class of instructions multiplied by the number of uses per instruction class.
- the opportunity cost of execution is also sometimes referred to as “degradation” (from the perspective of a loss of performance). Symmetrically, it can also be “performance gains”("accelerations” or “benefits” or “improvements”).
- the opportunity cost of execution can be determined from the number of uses that are actually counted and / or estimated based on the count history.
- the upstream determination of the "opportunity cost of execution" allows efficient optimizations conducted downstream (as described below).
- the "opportunity cost of execution" therefore corresponds to a "reading grid" specific to the processor and to the management of calculation tasks, that is to say to the definition of a "result". intermediary '(decision support) allowing subsequent specific management and dependent on this characterization.
- This perspective is particularly relevant insofar as it makes it possible to control the processor efficiently (intermediate aggregates make it possible to improve the controllability of the system).
- taking into account the classes of instructions and the number of uses of these classes of instructions allow the determination of said "opportunity cost of execution".
- the coefficients are determined offline. It is possible to determine the calibration data once and for all ("offline").
- the calibration data can be provided by the processor manufacturer.
- the calibration data can be present in a configuration file.
- the coefficients are determined online. The coefficients can be determined during the execution of a program, for example at the "reset" of the platform. In an open system (for example cluster or "Cloud”), whose topology is unknown a priori, it is possible to calibrate each extension at startup and determine the topology globally.
- an open system for example cluster or "Cloud”
- the coefficients are determined by multivariate statistical analysis. Different multivariate statistical analysis techniques can be used, possibly combined.
- the regression (linear) is advantageously fast.
- Principal component analysis (PCA) advantageously makes it possible to reduce the number of coefficients.
- Other techniques that can be used include factorial analysis of correspondence (AFC), so-called factorial analysis, data partitioning (clustering), multidimensional scaling (MDS), the analysis of similarities between variables. , multiple regression analysis, analysis of the ANOVA variance (bivariate), and its multivariate generalization (multivariate analysis of variance), discriminant analysis, canonical correlation analysis, logistic regression (logit model), artificial neural networks, decision trees , models of structural equations, joint analysis etc.
- the received compute task is associated with a predetermined processor core and the opportunity cost of performing the compute task is associated with a processor core other than the predetermined processor core. It is generally considered (but not required) what would be the cost of execution on the predetermined processor (if any), ie what would be the cost of "chasing” execution. In other cases, the other "candidate" cores are taken into consideration (one, several or all of the addressable cores).
- the opportunity cost of performing the computational task is determined for at least one processor core other than the predetermined processor core.
- all processor cores are considered each in turn and the placement optimization is to minimize the opportunity cost of execution (ie, for example, to select the processor core associated with the processor. lowest or lowest opportunity cost of execution).
- such a "minimum" function can be used.
- specific algorithms sequence of steps
- / or analytic functions and / or heuristics may be used (for example, the candidate cores may be compared in pairs and / or sampled according to various modalities. for example, to speed up investment decision-making).
- other criteria can be taken into account to optimize the placement of the calculation task. These criteria may be taken into account additionally (but in some cases may be substituted for the opportunity cost of execution criterion).
- These generally complementary criteria can include, in particular, parameters relating to the execution time of the calculation task and / or to the energy cost associated with the execution of the calculation task and / or the temperature (ie to the local consequence of the calculation task). execution considered).
- the various costs can be compared with each other and various arbitration logic can select a particular heart in consideration of these different criteria.
- combinatorial optimization or multi-objective optimization (some objectives may be antagonistic)
- different mathematical techniques are applicable.
- the weighting of these different criteria can in particular be variable and / or configurable.
- the respective weights allocated to the different placement optimization criteria can for example be configured online or offline. They can be "static” or “dynamic”. For example, the priority and / or weight of these different criteria may be variable over time.
- Analytical functions or algorithms can regulate the different allocations or arbitrations or trade-offs or priorities between optimization criteria for the placement of calculation tasks.
- the energy cost determination includes one or more of the steps of receiving initial indications of use of one or more predefined hardware extensions and / or receiving consumption states. of energy (eg DVFS) per processor core and / or receiving performance asymmetry information and a step of determining a power-gating and / or clock-gating energy optimization.
- energy eg DVFS
- the method further comprises a step of determining an adaptation cost of the instructions associated with the computing task, said step comprising one or more of the steps of translating one or more instructions and / or selecting a or multiple instruction versions and / or emulate one or more instructions and / or execute one or more instructions in a virtual machine.
- the adaptation of the instructions becomes necessary if it is determined, following the previous steps, a processor core not having the required hardware extension (heart "not equipped")
- the method further comprises a step of receiving a parameter and / or a logical scheduling and / or placement rule.
- Logical rules Boolean expressions, fuzzy logic, business rules, etc.
- factual threshold values such as maximum temperatures, time ranges, etc.
- the method further comprises a step of moving the compute task from the predetermined processor core to the determined processor core.
- the method further includes a step of disabling or shutting down one or more processor cores. Disable or “cut off the clock signal” or “put the heart in a reduced consumption status "(eg decreasing the clock frequency) or” darkening "(" dark silicon ")
- the functionally asymmetric multi-core processor is a physical processor or a virtual processor.
- the processor is a tangible or physical processor.
- the processor may also be a virtual processor, i.e. logically defined.
- the scope can be defined by the operating system.
- a processor can also be determined by a hypervisor.
- a computer program product comprising code instructions for performing one or more steps of the method, when said program is run on a computer.
- the system comprises a functionally asymmetric multi-core processor, at least one core of said processor being associated with one or more hardware extensions, the system comprising receiving means for receiving a computing task, said computing task being associated executable instructions by a hardware extension associated with the multi-core processor; receiving means for receiving calibration data; and means for determining an opportunity cost of executing the calculation task based on the calibration data.
- the system further comprises means selected from placement means for placing one or more computing tasks on one or more cores of the processor; means for counting the use of instruction classes by a hardware extension, said means including software and / or hardware counters; means or registers for saving the execution context of a calculation task; means for determining the migration cost and / or the adaptation cost and / or energy cost associated with continuing the execution of a calculation task on a predefined processor core; means for receiving one or more parameters and / or scheduling rules; means for determining and / or selecting a processor core; means for executing on a processor core without associated hardware extension a computational task initially intended to execute on a processor comprising one or more hardware extensions; means for moving a compute task from one processor core to another processor core; and means for disabling or shutting down one or more processor cores.
- Figure 1 illustrates examples of processor architectures.
- a functionally asymmetric multi-core processor FAMP 120 is compared to a symmetrical multi-core processor SMP 110 where each core contains all the hardware extensions and compared to a SMP 130 symmetrical multi-core processor containing only basic cores.
- the architecture may be shared memory or distributed. Hearts can sometimes be connected by a NoC (Network-on-Chip).
- a FAMP 120 architecture is close to that of a SMP 130 symmetrical multi-core processor.
- the memory architecture remains homogeneous, but the features of the cores can be heterogeneous.
- a FAMP 120 architecture may comprise four different types of cores (with the same base). The Processor size can therefore be reduced. Because of the ability to turn on or off the cores, and therefore to choose which to turn on the software application as needed, various optimizations can be made (energy gain, reduction of core temperature, etc.). .
- FIGS. 2A and 2B illustrate some aspects of the invention in terms of energy efficiency, depending on whether hardware extensions are used or not.
- FIG. 2A refers to a processor "Full” (or “extended”, ie with one or more hardware extensions) and “Basic” (without hardware extension).
- the figure illustrates the energy consumed by a processor (surfaces 121 and 122) as a function of the power consumed ("Power" on the ordinate 111) during an execution time (time on the abscissa 112) for the same application or thread on the two types of processors.
- a calculation task executed during a time tf consumes an energy E FU II 121.
- a “Basic” processor with an energy power P b a calculation task executed during a time t b consumes E BaS ic 122.
- the power Pf is greater than the power Pb because the hardware extension requires more power. It is common for the execution time t b to be greater than the execution time t f since it does not have an extension, a basic processor performs the calculation task less quickly.
- a general objective is to minimize the energy consumed, ie the area of the surface (shortest possible execution time with minimal power).
- the Full processor is more efficient.
- the ratio t b / t f indicates "acceleration" of the computation due to the extension hardware.
- FIG. 2B illustrates the variation of the energy gain (E BaS ic / E FU II) as a function of the acceleration tb t f , which indicates the opportunity to use an extended processor with hardware extension (s) rather than without. If the E Ba sic / EFuii ratio is less than 1, this indicates that there is less energy consumed with a Basic processor than with an Extended processor (if so, this operation domain 141 is called "Slighlty extended" which means that the hardware extension is not used enough).
- the ratio E Ba sic / EFuii is less than 1 and the acceleration ratio t b t f is less than 1, this shows a gain in both energy consumed and computing time performance to the advantage of the processor without hardware extension (this can happen in cases where the extension is misused).
- the ratio E Ba sic / EFuii is greater than 1, the extended processor consumes less energy (section 143 "highly extended"): the extension accelerates the code (ie the instructions) and the gains are obtained both in terms of energy (it is less) than in terms of acceleration (faster execution).
- the operating domain 144 corresponds to an acceleration of the calculations simultaneously with a lower energy consumption, to the advantage of an extended processor with hardware extension.
- a hardware expansion can reduce power consumption by reducing the number of steps required to perform specific calculations.
- the acceleration of the calculation on the extension can compensate for the additional power required by the extension itself.
- a hardware extension may be more efficient because of the strong coupling between extension and core.
- Hardware extensions frequently share a significant portion of the basic core and take advantage of direct access to L1 and L2 cache types. This coupling notably makes it possible to reduce data transfers and synchronization complexity.
- Hardware expansion usually involves additional cost in terms of area and static power consumption.
- An extension often includes registers of significant size, which may require a larger bandwidth as well as larger data transfers than those required by the core circuit. If the hardware expansion usage is too low, its clean energy consumption is likely to exceed the expected energy savings due to hardware acceleration. Worse, an underutilized hardware extension can lead to a reduction in the performance of a software application. For example, in the case of intensive data transfers between the hardware extension and the core registers, the overhead of the memory transfer may cancel the benefit of the acceleration.
- a hardware extension is usually not "transparent" to the application developer.
- the developer is frequently forced to explicitly handle extended instruction sets.
- the use of hardware extensions can be expensive compared to the single core circuit.
- the ARM processor NEON / VFP extension accounts for about 30% of the surface area of a Cortex A9 processor (not including caches, and 10% counting L1 cache).
- the method according to the invention makes it possible to (a) asymmetrically distribute the extensions in a multi-core processor, (b) to know when the use of an extension is "useful" from a performance and energy point of view and (c) moving the application over time on different cores, regardless of their starting requirement (which is chosen at compile time and therefore not necessarily adequate with the execution context)
- Figure 3 illustrates examples of architectures and job placement. The figure shows examples of scheduling tasks on an SMP, and two types of scheduling on a FAMP. The FPU adds FP-like instructions to an ISA circuit.
- each quanta of the task is executed on a processor having a Floating Point Unit (FPU). It turns out that the first and the last quanta do not contain FP instructions.
- FPU Floating Point Unit
- the first and the last quanta can run on a single core (having no FPU), which reduces the consumption of energy during execution.
- the 3rd quanta contains FP extensions, but not enough to "value" the FP unit.
- the method according to the invention allows the third quanta to be run on a "basic" type of processor, which optimizes the energy consumption to perform this calculation task.
- Figure 4 provides examples of steps of the method according to the invention.
- step 400 which takes place online and in a loop, the data relating to the use of the hardware extensions are collected (for example by the scheduler).
- the use of extensions is monitored ("monitoring" in English).
- the method for recovering data during execution focuses only on instructions for hardware or hardware extensions. According to this embodiment, the method is independent of the process scheduling.
- the code is executed on a processor with extension (ie with one or more extensions)
- the monitoring can rely on hardware counters.
- the extended instructions are no longer executed natively.
- routines for example at the time of adaptation of the code or in the emulation function if applicable
- routines for example at the time of adaptation of the code or in the emulation function if applicable
- the method according to the invention can use software counters and / or hardware counters associated with said software counters.
- a monitoring system can be very heavy at runtime.
- the data collected are linked to the instructions only involving extensions, and in fact the slowdown is negligible because it can be done in parallel with the execution.
- the filling of the counters can be done sequentially during the emulation of the accelerated functionality by the other cores, which can add an additional cost in terms of performance. To reduce this extra cost, it is possible to collect the data periodically, and, if necessary, to interpolate.
- step 410 an event of the scheduler (e.g., the end of a quanta) is determined, which triggers step 320.
- step 420 the collected data is read.
- the scheduler retrieves the data collected by the monitor (A.0) (for example by reading the hardware or software registers).
- a prediction is made. Thanks to the data retrieved, the prediction system studies the behavior of the application with respect to the use of extended instruction classes. By analyzing this behavior, he predicts the future behavior of the application.
- the prediction is said to be simple, ie it is reactive, assuming that the behavior of the application in the next quanta will be similar to the last quanta executed.
- the prediction is said to be "complex" and comprises steps of determining the different types of program phases. For example, these types of phases can be determined by saving a behavior history table and analyzing that history. Signatures can be used.
- the prediction unit may associate signatures with behaviors subsequently executed. When a signature is detected, the prediction unit can indicate the future behavior ie predict the future operations known and associated with the signature.
- step 430 an estimate is made of "degradation”.
- Base versus "Full”
- a given task is not tested on each teacher but it is estimated, at the time of scheduling, the cost that the task would have on different processors.
- the entire code is simulated, taking into account the complete features of the processor.
- This type of approach is generally not usable dynamically (online). Considering the particular case in which the processor cores all have the same base, all the instructions are executed in the same way in the different cores (for example two cores). Some edge effects may appear between cores, for example due to caching, but these effects remain residual and can be taken into account when learning the model.
- the estimation of performance degradation can be done in different ways.
- the estimation of the degradation may in fact be based on a model making it possible to calculate the said degradation as a function of the percentage of each instruction class by the extension. This approach is not necessarily appropriate, however, since different types of instructions can have very different accelerations for the same hardware extension.
- an estimate of the performance degradation can take into account the classes of instructions and assume a form according to:
- the values of the weighting coefficients ⁇ can be calculated dynamically online by performing a learning process performed by executing several codes on different types of processors. These values can also be calculated offline (and stored in a table readable by the scheduler for example).
- the calibration data determined from the execution of real calculation tasks make it possible in particular to take into account the cost of emulation and the complex evolutions of the beta coefficients.
- the beta values can be distinguished for a real platform and for a simulated platform.
- the beta coefficients are used by the scheduler to estimate performance degradation.
- the calibration must be repeated for each implementation of a hardware extension.
- the steps of prediction 430 and estimation of the degradation 440 are chronologically independent.
- the degradation can be estimated (440) from the predicted data (430).
- the degradation 440 can be calculated from the data recovered in step 400 and 420, before the prediction 430 is based on the degradation 440 thus calculated.
- the future degradation is predicted (and not the future instructions of the extensions). The data changes but the prediction process remains the same.
- step 430 the objective is to "predict" the number of instructions of each class that will be executed at t + 1 as a function of the past (of t, t-1 ... t-n).
- step 440 the objective is to "estimate" the execution cost at t + 1 on the different types of cores.
- the steps 430 and 440 can be performed in any order because it is possible independently to estimate the cost at time t on the recovered data and to predict the cost at time t + 1 from the cost at time t.
- a "prediction” approach consists of a relative determination of the future based on the data available in the present.
- An “estimate” allows a determination of the present using data observed in the past.
- an “estimate” allows the determination of the past in a new context from observations of the past in another context.
- a monitoring module captures the calibration data and an analysis module estimates the acceleration associated with the use of an extension.
- the input data must be as close as possible to the time of the extended instructions to be executed at the next quanta (assuming a continuity assumption, eg stable use of a floating point unit).
- a prediction module can then be used. This assumption is realistic when programs have distinct phases of stable behavior over sufficiently long time intervals. In the case of more erratic behavior, more advanced prediction modules can be implemented,
- a decision step is performed.
- a logical decision unit may take into account the information of the predicted degradation estimate and / or the overall objectives assigned to the platform, the objectives including, for example, objectives in terms of energy reduction, performance, priority and availability of hearts.
- the degradation can also take into account parameters such as the migration cost and / or the cost of adaptation of the code.
- the migration cost and the cost of the station may be static (e.g. offline learning) or dynamic (e.g. e-learning or online time monitoring).
- step 460 it is possible to proceed from the migration of one core to another ("core-switching").
- Migration techniques implemented can be performed with context saving (for example with saving the extension registers in software registers accessible from emulation).
- the code is optionally adapted.
- code for example, to execute code with extensions on a "basic" processor, it is possible to use binary adaptation techniques (eg by code translation), multi-code version, or even 'emulation.
- the present invention can be implemented from hardware and / or software elements. It may be available as a computer program product on a computer readable medium.
- the support can be electronic, magnetic, optical or electromagnetic.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Software Systems (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Computer Hardware Design (AREA)
- Computing Systems (AREA)
- Computational Mathematics (AREA)
- Mathematical Analysis (AREA)
- Mathematical Optimization (AREA)
- Pure & Applied Mathematics (AREA)
- Devices For Executing Special Programs (AREA)
- Debugging And Monitoring (AREA)
Abstract
Description
Claims
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| FR1553478A FR3035243B1 (fr) | 2015-04-20 | 2015-04-20 | Placement d'une tache de calcul sur un processeur fonctionnellement asymetrique |
| PCT/EP2016/055401 WO2016173766A1 (fr) | 2015-04-20 | 2016-03-14 | Placement d'une tâche de calcul sur un processeur fonctionnellement asymetrique |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP3286647A1 true EP3286647A1 (fr) | 2018-02-28 |
Family
ID=54291365
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP16711185.5A Withdrawn EP3286647A1 (fr) | 2015-04-20 | 2016-03-14 | Placement d'une tâche de calcul sur un processeur fonctionnellement asymetrique |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US20180095751A1 (fr) |
| EP (1) | EP3286647A1 (fr) |
| FR (1) | FR3035243B1 (fr) |
| WO (1) | WO2016173766A1 (fr) |
Families Citing this family (21)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR102285481B1 (ko) * | 2015-04-09 | 2021-08-02 | 에스케이하이닉스 주식회사 | NoC 반도체 장치의 태스크 매핑 방법 |
| US10700968B2 (en) * | 2016-10-19 | 2020-06-30 | Rex Computing, Inc. | Optimized function assignment in a multi-core processor |
| US10355975B2 (en) | 2016-10-19 | 2019-07-16 | Rex Computing, Inc. | Latency guaranteed network on chip |
| JP2018092311A (ja) * | 2016-12-01 | 2018-06-14 | キヤノン株式会社 | 情報処理装置、その制御方法、及びプログラム |
| ES2933675T3 (es) | 2016-12-31 | 2023-02-13 | Intel Corp | Sistemas, métodos y aparatos para informática heterogénea |
| CA3108151C (fr) | 2017-02-23 | 2024-02-20 | Cerebras Systems Inc. | Apprentissage profond accelere |
| US10795836B2 (en) * | 2017-04-17 | 2020-10-06 | Microsoft Technology Licensing, Llc | Data processing performance enhancement for neural networks using a virtualized data iterator |
| US11488004B2 (en) | 2017-04-17 | 2022-11-01 | Cerebras Systems Inc. | Neuron smearing for accelerated deep learning |
| WO2018193380A1 (fr) | 2017-04-17 | 2018-10-25 | Cerebras Systems Inc. | Vecteurs de tissu pour accélération d'apprentissage profond |
| WO2018193352A1 (fr) | 2017-04-17 | 2018-10-25 | Cerebras Systems Inc. | Tâches déclenchées par un flux de données pour apprentissage profond accéléré |
| WO2020044152A1 (fr) | 2018-08-28 | 2020-03-05 | Cerebras Systems Inc. | Textile informatique mis à l'échelle pour apprentissage profond accéléré |
| US11321087B2 (en) | 2018-08-29 | 2022-05-03 | Cerebras Systems Inc. | ISA enhancements for accelerated deep learning |
| US11328208B2 (en) | 2018-08-29 | 2022-05-10 | Cerebras Systems Inc. | Processor element redundancy for accelerated deep learning |
| US11188617B2 (en) * | 2019-01-10 | 2021-11-30 | Nokia Technologies Oy | Method and network node for internet-of-things (IoT) feature selection for storage and computation |
| WO2021074867A1 (fr) | 2019-10-16 | 2021-04-22 | Cerebras Systems Inc. | Filtrage d'ondelettes avancé pour apprentissage profond accéléré |
| WO2021074795A1 (fr) | 2019-10-16 | 2021-04-22 | Cerebras Systems Inc. | Routage dynamique pour apprentissage profond accéléré |
| US12008398B2 (en) * | 2019-12-28 | 2024-06-11 | Intel Corporation | Performance monitoring in heterogeneous systems |
| US11775298B2 (en) * | 2020-04-24 | 2023-10-03 | Intel Corporation | Frequency scaling for per-core accelerator assignments |
| CN113726474A (zh) * | 2020-05-26 | 2021-11-30 | 索尼公司 | 物联网中的操作电子设备、管理电子设备和通信方法 |
| US11989591B2 (en) * | 2020-09-30 | 2024-05-21 | Advanced Micro Devices, Inc. | Dynamically configurable overprovisioned microprocessor |
| EP4411541B1 (fr) * | 2023-02-01 | 2025-12-17 | Airbus S.A.S. | Procédé de planification de tâches mis en uvre par ordinateur |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP4784827B2 (ja) * | 2006-06-06 | 2011-10-05 | 学校法人早稲田大学 | ヘテロジニアスマルチプロセッサ向けグローバルコンパイラ |
| US8782645B2 (en) * | 2011-05-11 | 2014-07-15 | Advanced Micro Devices, Inc. | Automatic load balancing for heterogeneous cores |
| US9720730B2 (en) | 2011-12-30 | 2017-08-01 | Intel Corporation | Providing an asymmetric multicore processor system transparently to an operating system |
| KR20130115574A (ko) * | 2012-04-12 | 2013-10-22 | 삼성전자주식회사 | 단말기에서 태스크 스케줄링을 수행하는 방법 및 장치 |
| US20150355700A1 (en) * | 2014-06-10 | 2015-12-10 | Qualcomm Incorporated | Systems and methods of managing processor device power consumption |
-
2015
- 2015-04-20 FR FR1553478A patent/FR3035243B1/fr not_active Expired - Fee Related
-
2016
- 2016-03-14 EP EP16711185.5A patent/EP3286647A1/fr not_active Withdrawn
- 2016-03-14 WO PCT/EP2016/055401 patent/WO2016173766A1/fr not_active Ceased
- 2016-03-14 US US15/567,067 patent/US20180095751A1/en not_active Abandoned
Also Published As
| Publication number | Publication date |
|---|---|
| WO2016173766A1 (fr) | 2016-11-03 |
| FR3035243B1 (fr) | 2018-06-29 |
| US20180095751A1 (en) | 2018-04-05 |
| FR3035243A1 (fr) | 2016-10-21 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| EP3286647A1 (fr) | Placement d'une tâche de calcul sur un processeur fonctionnellement asymetrique | |
| Zhou et al. | Aquatope: Qos-and-uncertainty-aware resource management for multi-stage serverless workflows | |
| Um et al. | Fastflow: Accelerating deep learning model training with smart offloading of input data pipeline | |
| US10908884B2 (en) | Methods and apparatus for runtime multi-scheduling of software executing on a heterogeneous system | |
| EP2257876B1 (fr) | Methode de prechargement dans une hierarchie de memoires des configurations d'un systeme heterogene reconfigurable de traitement de l'information | |
| US20230229514A1 (en) | Intelligent orchestration of classic-quantum computational graphs | |
| Wu et al. | Irina: Accelerating dnn inference with efficient online scheduling | |
| Tarsa et al. | Post-silicon cpu adaptation made practical using machine learning | |
| EP3502895A1 (fr) | Commande de la consommation énergétique d'une grappe de serveurs | |
| Pandey et al. | Funcmem: reducing cold start latency in serverless computing through memory prediction and adaptive task execution | |
| Ahmed et al. | RALB‐HC: A resource‐aware load balancer for heterogeneous cluster | |
| Ogden et al. | Layercake: Efficient inference serving with cloud and mobile resources | |
| CN112559053B (zh) | 可重构处理器数据同步处理方法及装置 | |
| Agarwal et al. | On-demand cold start frequency reduction with off-policy reinforcement learning in serverless computing | |
| Lozano et al. | Learning-based phase-aware multi-core CPU workload forecasting | |
| US20230350485A1 (en) | Compiler directed fine grained power management | |
| Saravana Kumar et al. | Cold start prediction and provisioning optimization in serverless computing using deep learning | |
| Yu et al. | Nitro: Boosting Distributed Reinforcement Learning with Serverless Computing | |
| FR3056786B1 (fr) | Procede de gestion des taches de calcul sur un processeur multi-cœurs fonctionnellement asymetrique | |
| EP2252933B1 (fr) | Architecture de traitement informatique accelere | |
| TW202530983A (zh) | 分散式應用之動態適應任務執行並行性 | |
| Gao et al. | Ymir: A scheduler for foundation model fine-tuning workloads in datacenters | |
| Hu et al. | Patchwork: A Unified Framework for RAG Serving | |
| Yuhas et al. | Toward state-aware scheduling of machine-learning workloads | |
| Thethi et al. | Power optimization of a single-core processor using LSTM based encoder–decoder model for online DVFS |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20171010 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| AX | Request for extension of the european patent |
Extension state: BA ME |
|
| DAV | Request for validation of the european patent (deleted) | ||
| DAX | Request for extension of the european patent (deleted) | ||
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: EXAMINATION IS IN PROGRESS |
|
| 17Q | First examination report despatched |
Effective date: 20190618 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION HAS BEEN WITHDRAWN |
|
| 18W | Application withdrawn |
Effective date: 20210106 |