Showing posts with label Advanced computer architecture. Show all posts
Showing posts with label Advanced computer architecture. Show all posts

Thursday, October 1, 2015

IA-64 Architecture



IA-64 is a load/store architecture, with 64-bit memory addresses and registers. There are 128 64-bit general-purpose registers, 128 82-bit floating point registers, 64 1-bit predicate registers  and 8 64-bit branch registers (used to hold branch destination addresses). There are also assorted control registers that we will not concern outselves with. Intel have chosen to use a RISC 2-style register stack with some modifications. The first 32 registers are global, and the remainder are allocated to procedures/functions. The size of the block of registers allocated to a procedure/function is variable. Also, instead of overlapping blocks of registers, the calling procedure's first local register is used as a pointer to a block of the called procedure's local registers for passing parameters.

-64 divides instructions up into groups and bundles. A group is a sequence of instructions that could in principle be executed in parallel provided:
  • Sufficient Hardware is Available — Clearly, it is expected that later processors will have more hardware and will be able to execute a larger number of instructions simultaneously.
  • No Memory Dependencies — The compiler will guarantee that there are no data dependencies between registers that prevent execution, but the hardware must check dependencies in memory. This comes back to the problem of aliasing of memory addresses.

  • I — ALU and non-ALU Integer — These include the usual integer arithmetic/logical operations, together with various moves, tests and shifts.
  • M — ALU and Memory — These also include interger arithmetic/logical operations, together with memory access.
  • F — Floating point — Obviously, floating point operations.
  • B — Branches — Branches and procedure calls. The branches contain static prediction information which is used until the branch prediction algorithm builds up enough history information.
  • L+X — Extended — Various other things.

Wednesday, September 30, 2015

VLIW and EPIC processors.



Instruction-level parallelism in EPIC

Static Multiple Issue : The VLIW Approach

Very long instruction word (VLIW) refers to processor architectures designed to take advantage of instruction level parallelism (ILP). Whereas conventional processors mostly allow programs only to specify instructions that will be executed in sequence, a VLIW processor allows programs to explicitly specify instructions that will be executed at the same time (that is, in parallel).
Modern superscalar processors are complex, power-hungry devices that present an antiquated view of processor architecture to the programmer in the interests of backwards compatibility — and do a lot of work to achieve high performance while maintaining this illusion.
The alternative to superscalar is aVLIW architecture, but these have traditionally been actively backwards-incompatible, with performancehighly dependent on the (frequently mediocre) abilities ofthe compiler. Neither VLIW nor superscalar are perfect architectures:
each has its own set of trade-offs. This report discusses the relative strengths and weaknesses of the two, focusing on the benefits of VLIW andthe closely-related EPIC architecture as used in Intel’s Itanium processor family.An introduction to the motivation behind VLIW is given, VLIWand EPIC are discussed in detail,and then two case studies are presented: the Analog Devices SHARC family of DSPs, demonstrating the VLIW influences present in a modern DSP; and Intel’s Itanium processor family, which is to date the only implementation of EPIC
VLIWs use multiple, independent functional units. Rather than attempting to issue multiple,
independent instructions to the units, a VLIW packages the multiple operations into one very
long instruction, or requires that the instructions in the issue packet satisfy the same constraints.

Tuesday, September 29, 2015

hardware based speculation



Hardware speculation
·Hardware-based speculation combines three key ideas: dynamic branch prediction to choose which instructions to execute, speculation to allow the execution of instructions before the control dependences are resolved and dynamic scheduling to deal with the scheduling of different combinations of basic blocks.
·Hardware-based speculation follows the predicted flow of data values to choose when to execute instructions. This method of executing programs is essentially a data-flow execution: operations execute as soon as their operands are available.
·The approach is implemented in a number of processors (PowerPC 603/604/G3/G4, MIPS R10000/R12000, Intel Pentium II/III/ 4, Alpha 21264, and AMD K5/K6/Athlon), is to implement speculative execution based on Tomasulo’s algorithm.
·The key idea behind implementing speculation is to allow instructions to execute out of order but to force them to commit in order and to prevent any irrevocable action until an instruction commits.
·In the simple single-issue five-stage pipeline we could ensure that instructions committed in order, and only after any exceptions for that instruction had been detected, simply by moving writes to the end of the pipeline.

Monday, September 28, 2015

dynamic scheduling using tomosulo’s approach



Dynamic Scheduling Using Tomasulo’s Approach :
Tomasulo's algorithm implements register renaming through the use of what are called reservation stations. Reservation stations are buffers which fetch and store instruction operands as soon as they are available
In addition, pending instructions designate the reservation station that will provide their input. Finally, when successive writes to a register overlap in execution, only the last one is actually used to update the register. As instructions are issued, the register specifies for pending operands are renamed to the names of the reservation station, which provides register renaming.


the basic structure of a Tomasulo-based MIPS processor, including both the floating-point unit and the load/store unit.
Instructions are sent from the instruction unit into the instruction queue from which they are issued in FIFO order.
The reservation stations include the operation and the actual operands, as well as information used for detecting and resolving hazards. Load buffers have three functions: hold the components of the effective address until it is computed, track outstanding loads that are waiting on the memory, and hold the results of completed loads that are waiting for the CDB.
Similarly, store buffers have three functions: hold the components of the effective addr ess until it is computed, hold the destination memory addresses of outstanding stores that are waiting for the data value to store, and hold the address and value to store until the memory unit is available.
All results from either the FP units or the load unit are put on the CDB, which goes to the FP register file as well as to the reservation stations and store buffers. The FP adders implement addition and subtraction, and the FP multipliers do multiplication and division.

Sunday, September 27, 2015

concept of dynamic scheduling



Overcoming Data Hazards with Dynamic Scheduling:
The Dynamic Scheduling is used handle some cases when dependences are unknown at a compile time. In which the hardware rearranges the instruction execution to reduce the stalls while maintaining data flow and exception behavior.
It also allows code that was compiled with one pipeline in mind to run efficiently on a different pipeline. Although a dynamically scheduled processor cannot change the data flow, it tries to avoid stalling when dependences, which could generate hazards, are present.

Dynamic Scheduling:
A major limitation of the simple pipelining techniques is that they all use in -order instruction issue and execution: Instructions are issued in program order and if an instruct ion is stalled in the pipeline, no later instructions can proceed. Thus, if there is a dependence between two closely spaced instructions in the pipeline, this will lead to a hazard and a stall. If there are multiple functional units, these units could lie idle. If instruction j depends on a long-running instruction i, currently in execution in the pipeline, then all instructions after j must be stalled until i is finished and j can execute. For example, consider this code:
DIV.D F0,F2,F4 ADD.D F10,F0,F8 SUB.D F12,F8,F14
Out-of-order execution introduces the possibility of WAR and WAW hazards, which do not exist in the five-stage integer pipeline and its logical extension to an in-order floating-point pipeline.
Out-of-order completion also creates major complications in handling exceptions. Dynamic scheduling with out-of-order completion must preserve exception behavior in the sense that exactly those exceptions that would arise if the program were executed in strict program order actually do arise.

Imprecise exceptions can occur because of two possibilities:
1. The pipeline may have already completed instructions that are later in program order than the instruction causing the exception, and

Friday, July 4, 2014

ADVANCED COMPUTER ARCHITECTURE

1.Explain the fundamentals of computer Design?

           Computer technology has made incredible progress in the roughly from last 55 years. This rapid rate of improvement has come both from advances in the technology used to build computers and from innovation in computer design. During the first 25 years of electronic computers, both forces made a major contribution; but beginning in about 1970, computer designers became largely dependent upon integrated circuit technology. During the 1970s, performance continued to improve at about 25% to 30% per year for the mainframes and minicomputers that dominated the industry.

             The late 1970s after invention of microprocessor the growth roughly increased 35% per year in performance. This growth rate, combined with the cost advantages of a mass-produced microprocessor, led to an increasing fraction of the computer business. In addition, two significant changes are observed in computer industry.
  • First, the virtual elimination of assembly language programming reduced the need for object-code compatibility.
  • Second, the creation of standardized, vendor-independent operating systems, such as UNIX and its clone, Linux, lowered the cost and risk of bringing out a new architecture.

These changes made it possible to successfully develop a new set of architectures, called RISC (Reduced Instruction Set Computer) architectures. In the early 1980s. The RISC-based machines focused the attention of designers on two critical performance techniques, the exploitation of instruction-level parallelism and the use of caches. The combination of architectural and organizational enhancements has led to 20 years of sustained growth in performance at an annual rate of over 50%. Figure 1.1 shows the effect of this difference in performance growth rates.

The effect of this dramatic growth rate has been twofold.
• First, it has significantly enhanced the capability available to computer users. For many applications, the highest performance microprocessors of today outperform the supercomputer of less than 10 years ago.
• Second, this dramatic rate of improvement has led to the dominance of micro-processor-based computers across the entire range of the computer design.

Tuesday, January 21, 2014


UNIT I

11.  Explain the hybrid approach for encoding an instruction set?
The hybrid approach reduces the variability in size and work of the variable architecture but provide multiple instruction lengths to reduce code size. 

12.  What are the registers used for MIPS processors.
MIPS has 34, 64-bit general purpose registers (GPRs), named R0, R1…R31. GPRs are sometimes called as integer registers. There are also a set of 32 floating point registers (FPRs), named F0,F1….F31, which can hold 32 single precision values and 32 double precision values. 



13.  Explain the concept behind pipelining.
Pipelining is an implementation technique whereby multiple instructions are overlapped in execution. It takes advantage of parallelism that exists among actions needed to execute an instruction. 

14.  Write about pipe stages and processor cycle.
Different steps in an instruction are completed in different parts of different instruction is parallel. Each of these steps is called a pipe stage or pipe segment.  The time required between moving an instruction one step down the pipeline is called processor cycle. 

15.  Explain pipeline hazard and mention the different hazards in pipeline.
Hazards are situations that prevent the next instruction in the instruction stream from executing during its designated clock cycle. Hazards reduce the overall performance from the ideal speedup gained by pipelining. The three classes of hazards are,
• Structural hazards.
• Data hazards.
• Control hazards. 

16. Explain the concept of forwarding.
Forwarding can be generalized to include passing a result directly to the functional unit that fetches it. The result is forwarded from the pipeline register corresponding to the output of one unit to the input of the same unit.  

17. Mention the different schemes to reduce pipeline branch penalties.
a. Freeze or flush the pipeline
b. Treat every branch as not taken
c. Treat every branch as taken 
d. Delayed branch

18. Consider an unpipelined processor.
Assume that it has a 1ns clock cycle and that it uses 4 cycles for ALU operations and branches and 5 cycles for memory operations. Assume that the relative frequencies of these operations are 40%, 20% and 40% respectively.
Suppose that due to clock skew and setup, pipelining the processor adds 0.2 ns of overhead to the clock.
Ignoring any latency impact, how much speedup in the instruction execution rate will we gain from a pipeline? 
The average instruction execution time on an unpipelined processor is  = clock cycle x Average CPI = 1 ns x ((40% x 4)+(20 x 4)+(40 x 5)) = 4.4 ns.
The average instruction execution time on an pipelined processor is  = 1+0.2ns = 1.2ns  Speedup = Avg. instruction time unpipelined/ Avg. instruction time pipelined = 4.4/1.2 = 3.7 times. 

19. Briefly explain the different conventions for ordering the bytes within a larger object?
Little endian byte order puts the byte whose address is “x….x000” at the least significant position in the double word. Big endian byte order puts the byte whose address is “x….x000” at the most significant position in the double word. 

20. When do data hazards arise? 

Data hazards arise when an instruction depends on the results of a previous instruction in a way that is expressed by the overlapping of instructions in the pipeline.