Microprocessor architecture and instruction cycle
Microprocessor building blocks, buses, the 8085 as reference, instruction/machine cycles and T-states, delay loops, CPI and pipelining with worked numericals.
Drafted with Aria, reviewed by the AiCanCode.org team. Spotted an error? Use Give Feedback at the bottom of the page.
Why it matters
Programmable instruments, PLCs and data loggers are built around a processor that fetches and executes instructions. Understanding the internal blocks (registers, ALU, control unit, buses) and how long each instruction takes lets you estimate execution time, write accurate software delays and judge whether a processor can keep up with a sampling rate.
Key ideas
Building blocks.
- ALU: performs arithmetic (add, subtract, increment) and logic (AND, OR, XOR, compare, shift) on data, usually with the accumulator as one operand, and updates the flags.
- Registers: small, fast storage inside the CPU. General-purpose registers hold data; special registers include the program counter (PC, address of the next instruction), stack pointer (SP, top of the stack in RAM), instruction register (IR, holds the opcode being decoded) and the flag (status) register.
- Flags: Sign (S), Zero (Z), Carry (CY), Parity (P), Auxiliary carry (AC, carry out of bit 3, used for BCD). Conditional jumps test them.
- Control unit: decodes the instruction and generates the timing and control signals (read, write, memory or I/O select) in the right sequence. It can be hard-wired (fast, used in RISC) or microprogrammed (flexible, used in many CISC designs).
- Buses: the address bus (unidirectional, from CPU) selects a memory or I/O location; the data bus (bidirectional) carries the data; the control bus carries RD′, WR′, IO/M′, ALE, interrupt and ready signals. An n-bit address bus reaches 2ⁿ locations.
The 8085 as a reference processor. 8-bit data bus, 16-bit address bus (64 KB), registers A (accumulator), B, C, D, E, H, L (usable as pairs BC, DE, HL), 16-bit SP and PC, five flags. The lower address byte and the data bus share pins AD₀–AD₇; ALE (address latch enable) goes HIGH in the first clock state so an external latch can capture A₀–A₇. The internal clock is half the crystal frequency.
Instruction cycle, machine cycle, T-state.
- Instruction cycle: everything needed to fetch and execute one instruction; it consists of one or more machine cycles.
- Machine cycle: one bus operation, e.g. opcode fetch, memory read, memory write, I/O read, I/O write, interrupt acknowledge.
- T-state: one clock period. In the 8085 an opcode fetch takes 4 T-states (6 for some instructions) and memory or I/O read/write cycles take 3.
- Examples (8085): MOV r, r: 4T (1 machine cycle). MVI r, data: 7T (opcode fetch + memory read). LDA addr: 13T (opcode fetch + 3 reads). DCR r: 4T. JNZ addr: 10T if taken, 7T if not.
Fetch–decode–execute. PC places the address on the bus; memory returns the opcode to the IR; PC increments; the control unit decodes; operands are fetched if needed; the ALU executes; results are written back and flags updated.
Architectural styles.
- Von Neumann: one memory and one bus for program and data (simpler; the bus is a bottleneck).
- Harvard: separate program and data memories and buses, allowing simultaneous instruction fetch and data access (most microcontrollers and DSPs).
- CISC: many complex, variable-length instructions, several cycles each (8085, x86). RISC: few simple fixed-length instructions, load/store architecture, designed for one instruction per clock with pipelining (ARM, AVR).
- Pipelining overlaps the stages of successive instructions. Hazards (data dependences, branches, resource conflicts) cause stalls that raise the effective CPI.
Interrupts. An external request makes the CPU finish the current instruction, push the PC on the stack, jump to an interrupt service routine and return with RET/RETI. Covered in detail under timers and interrupts.
Formulas
Addressable locations = 2ⁿ (n address lines)
T = 1 / f_clk (8085: f_clk = f_crystal / 2)
Instruction time = (number of T-states) × T
CPU time = IC × CPI / f_clk
- IC: instruction count; CPI: average cycles per instruction; f_clk in Hz.
Pipeline cycles = k + N − 1 and Speed-up = N·k / (k + N − 1)
- k: number of pipeline stages; N: instructions; ideal, no stalls.
Worked examples
Example 1 (standard). A program executes 2 × 10⁶ instructions with an average CPI of 1.5 on a 100 MHz processor. Find the execution time.
- CPU time = IC × CPI / f = 2 × 10⁶ × 1.5 / (100 × 10⁶ Hz).
- = 3 × 10⁶ / 10⁸ = 0.03 s.
Answer: 30 ms
Example 2 (GATE level: delay loop). An 8085 runs from a 6 MHz crystal. Find the delay produced by: MVI C, FFH (7T); LOOP: DCR C (4T); JNZ LOOP (10T taken / 7T not taken).
- f_clk = 6 MHz / 2 = 3 MHz → T = 333.3 ns.
- FFH = 255 passes. DCR runs 255 times; JNZ is taken 254 times and not taken once.
- T-states = 7 + 255 × 4 + 254 × 10 + 1 × 7 = 7 + 1020 + 2540 + 7 = 3574.
- Delay = 3574 × 333.3 ns = 1.191 ms.
Answer: 3574 T-states ≈ 1.19 ms
Example 3 (pipeline). A 5-stage pipeline executes 100 instructions with no stalls, each stage taking one clock. Compare with a non-pipelined processor that takes 5 clocks per instruction.
- Pipelined cycles = k + N − 1 = 5 + 100 − 1 = 104.
- Non-pipelined cycles = 5 × 100 = 500.
- Speed-up = 500 / 104 = 4.81 (approaching 5 for long programs).
Answer: 104 cycles; speed-up 4.81
Common mistakes
- Using the crystal frequency as the 8085 clock (it is divided by 2).
- Counting the final JNZ as taken (10T) instead of not taken (7T) in delay loops.
- Confusing machine cycles with T-states, or instruction cycles with machine cycles.
- Thinking the address bus is bidirectional; only the data bus is.
- Assuming a pipeline makes each instruction faster; it raises throughput, not single-instruction latency.
For GATE IN
Typical questions: addressable memory from bus width, T-states and execution time of an instruction sequence, software delay loops (single and nested), register and flag contents after a short 8085 program, CPI and pipeline speed-up, and Harvard versus von Neumann or RISC versus CISC distinctions. Practise tracing short assembly programs while tracking every flag.
Quick check
- How much memory can a 20-bit address bus address?
- How many T-states does MVI A, 32H take on the 8085?
- 4 clock cycles at 2 GHz take how long?
- Which register holds the address of the next instruction?
Answers: 1. 1 MB (2²⁰ bytes) 2. 7 3. 2 ns 4. Program counter
Interview questions
All Digital Electronics and Microcontrollers interview questionsTry answering each one aloud before you open it.
1.What is a microprocessor, and how does it differ from a microcontroller?Concept
A microprocessor is an integrated circuit designed to perform computation tasks and execute instructions. It typically requires external components like memory and input/output interfaces to function. A microcontroller, on the other hand, is a compact integrated circuit that includes a processor, memory, and input/output peripherals on a single chip, making it suitable for embedded applications.
2.Explain the basic architecture of a microprocessor.Concept
The basic architecture of a microprocessor includes the Arithmetic Logic Unit (ALU), Control Unit (CU), and registers. The ALU performs arithmetic and logical operations, the CU directs the operation of the processor, and registers are small storage locations for quick data access. The architecture may also include buses for data, address, and control signals.
3.What is an instruction cycle in a microprocessor?Concept
An instruction cycle is the process by which a microprocessor fetches, decodes, and executes an instruction. It typically consists of three main stages: Fetch, where the instruction is retrieved from memory; Decode, where the instruction is interpreted; and Execute, where the instruction is carried out by the processor.
4.Why is pipelining used in microprocessor architecture?Application
Pipelining is used in microprocessor architecture to increase instruction throughput by overlapping the execution of multiple instructions. It divides the instruction cycle into separate stages, allowing the next instruction to begin before the previous one has completed. This improves the overall efficiency and speed of the processor.
5.What happens if there is a pipeline hazard in a microprocessor?Application
A pipeline hazard occurs when there is a conflict in the pipeline stages, which can lead to incorrect execution of instructions. There are three types of hazards: data hazards, control hazards, and structural hazards. These hazards can cause delays or require the pipeline to be stalled or flushed, reducing the efficiency of the processor.
6.How does a microprocessor handle interrupts?Application
A microprocessor handles interrupts by temporarily halting its current execution to address the interrupt request. It saves the current state of the processor, executes an interrupt service routine (ISR) to handle the interrupt, and then restores the processor's state to resume normal execution. This allows the processor to respond to urgent tasks while maintaining overall program flow.
7.Explain the role of the control unit in a microprocessor.Concept
The control unit (CU) in a microprocessor is responsible for directing the operation of the processor. It interprets the instructions fetched from memory and generates control signals to coordinate the activities of the ALU, registers, and other components. The CU ensures that the correct sequence of operations is followed for each instruction.
8.What is the significance of the program counter in a microprocessor?Concept
The program counter (PC) is a register in a microprocessor that holds the address of the next instruction to be executed. It is automatically incremented after each instruction fetch, ensuring the sequential execution of instructions. The PC is crucial for maintaining the flow of the program and enabling jumps or branches in the instruction sequence.
9.Calculate the time taken to execute an instruction cycle if the clock frequency is 2 GHz and the cycle consists of 4 clock cycles.Numerical
The time taken for one clock cycle is the inverse of the clock frequency. For a 2 GHz clock frequency, the time per cycle is 1 / (2 × 10^9) seconds. Therefore, for 4 clock cycles, the time taken is 4 × (1 / (2 × 10^9)) = 2 × 10^-9 seconds or 2 nanoseconds.
10.If a microprocessor has a 16-bit address bus, how much memory can it address?Numerical
A 16-bit address bus can address 2^16 different memory locations. Since each location typically holds 1 byte, the total addressable memory is 2^16 bytes, which is equivalent to 64 kilobytes (KB).
Finished this topic? Mark it so your progress, study plan and readiness keep up.
Stuck on something here?