ARM Cortex-M and AVR architecture overview
ARM Cortex-M and AVR compared: registers, pipelines, memory maps, NVIC and SysTick, with execution-time, CPI and SysTick reload calculations.
Drafted with Aria, reviewed by the AiCanCode.org team. Spotted an error? Use Give Feedback at the bottom of the page.
Why it matters
Most new embedded designs in mechatronics use either a 32-bit ARM Cortex-M part (STM32, LPC, nRF, SAM, RP2040) or an 8-bit AVR (ATmega, ATtiny, the chip on an Arduino Uno). Choosing between them — and estimating whether a control loop, a communication stack or a filter will meet its deadline — needs a working picture of their registers, pipelines, memory maps and interrupt hardware.
Key ideas
Common ground. Both are RISC, load–store machines: arithmetic works only on registers, and memory is reached with explicit load/store instructions. Both keep program code in on-chip Flash and data in on-chip SRAM, and both integrate timers, serial ports, ADCs and GPIO on the chip.
AVR (Microchip, formerly Atmel) — 8-bit.
- 8-bit data path and ALU; 32 general-purpose 8-bit registers R0–R31. The pairs R26:R27, R28:R29, R30:R31 form the 16-bit pointer registers X, Y and Z.
- Harvard architecture: program memory (16-bit-wide Flash words) and data memory are separate buses. A two-stage pipeline overlaps the fetch of the next instruction with execution of the current one, so most instructions take one clock — roughly 1 MIPS per MHz.
- Data memory map of a typical ATmega (e.g. ATmega328P): registers at 0x00–0x1F, 64 I/O registers at 0x20–0x5F, extended I/O at 0x60–0xFF, SRAM above that. On-chip EEPROM is a separate space for non-volatile parameters.
- Typical part: ATmega328P — 32 KB Flash, 2 KB SRAM, 1 KB EEPROM, up to 20 MHz.
- Interrupts use a fixed vector table; the global I-bit in SREG is cleared on entry, so interrupts do not nest unless the ISR re-enables them. Response is at least four clock cycles before the vector is reached.
- A 16- or 32-bit operation needs several 8-bit instructions (ADD followed by ADC for each higher byte), which is the main performance limit.
ARM Cortex-M — 32-bit. Cortex-M is a family of processor cores licensed by Arm; the chip vendor adds memory and peripherals.
- Registers: R0–R12 general purpose, R13 = stack pointer (MSP/PSP), R14 = link register, R15 = program counter, plus the program status register xPSR.
- Instruction set: Thumb (16-bit) on M0/M0+; Thumb-2 (mixed 16/32-bit) on M3 and above, which gives near-ARM performance with compact code.
- Family: M0/M0+ (ARMv6-M, smallest and lowest power, von Neumann bus); M3 (ARMv7-M, 3-stage pipeline, Harvard buses, hardware divide); M4 (M3 plus DSP/SIMD instructions and an optional single-precision FPU); M7 (6-stage superscalar, caches, optional double-precision FPU); M23/M33 (ARMv8-M with TrustZone security).
- Fixed 4 GB memory map: code region from 0x0000 0000, SRAM from 0x2000 0000, peripherals from 0x4000 0000, private peripheral bus (NVIC, SysTick, debug) from 0xE000 0000. Same addresses on every vendor's chip, which makes code portable.
- NVIC (Nested Vectored Interrupt Controller): programmable priorities, nesting (a higher-priority interrupt pre-empts a lower one), and hardware stacking — on entry the core itself pushes eight registers (R0–R3, R12, LR, PC, xPSR), so an ISR can be a plain C function. Latency is 12 cycles on M3/M4 with zero-wait-state memory; back-to-back interrupts are tail-chained without unstacking.
- SysTick: a 24-bit down-counter in every Cortex-M, intended for the RTOS tick.
- Debug through SWD/JTAG with hardware breakpoints; higher cores add instruction trace.
How to choose. AVR: simple, robust 5 V I/O, very small code, deterministic single-cycle timing, low cost for small jobs. Cortex-M: 32-bit arithmetic, much more memory, DMA, FPU/DSP for control and filtering, RTOS and communication stacks (Ethernet, USB, CAN). Low power is not decided by the core alone — Cortex-M0+ parts and AVR parts both reach very low sleep currents; compare datasheet currents in µA/MHz and in sleep.
Connections. The 8051 (previous topic) is CISC with an accumulator; AVR and Cortex-M are register-rich RISC. Timers, interrupts and the RTOS tick in later topics build directly on NVIC and SysTick.
Formulas
T_clk = 1 / f
- T_clk = clock period (s); f = CPU clock frequency (Hz).
t_exec = N × CPI / f
- t_exec = execution time (s); N = instructions executed; CPI = average clock cycles per instruction (dimensionless). Applies when CPI is known for the actual instruction mix and memory wait states.
CPI = (f × t_exec) / N and MIPS = f / (CPI × 10⁶)
- MIPS = millions of instructions per second.
RELOAD = f × T_tick − 1
- SysTick reload value (counts), T_tick = desired tick period (s). The counter counts RELOAD → 0, so the period is (RELOAD + 1)/f. RELOAD must not exceed 2²⁴ − 1 = 16 777 215.
T_max = 2²⁴ / f
- Longest SysTick period without a prescaler (s).
t_latency = n_cycles / f
- Interrupt latency in seconds for a latency of n_cycles clocks.
Worked examples
Example 1 (standard). An ATmega328P runs at 16 MHz. A delay loop executes 1000 iterations of 5 clock cycles each. Find the delay and the clock period.
T_clk = 1 / f = 1 / (16 × 10⁶ Hz) = 62.5 ns.- Total cycles = 1000 × 5 = 5000 cycles.
t = cycles × T_clk = 5000 × 62.5 ns = 312.5 µs.
Answer: delay = 312.5 µs (clock period 62.5 ns).
Example 2 (GATE level). A Cortex-M4 runs at 72 MHz. (a) Find the SysTick reload value for a 1 ms tick and the longest possible SysTick period. (b) A control routine performs 10 000 additions of 32-bit numbers already held in registers. On the Cortex-M4 each is one 1-cycle ADD. On a 16 MHz AVR each needs ADD + 3 × ADC, all 1-cycle. Compare execution times. (c) Express the M4's 12-cycle interrupt latency in time.
RELOAD = f × T_tick − 1 = 72 × 10⁶ × 1 × 10⁻³ − 1 = 71 999. This is below 16 777 215, so it fits.T_max = 2²⁴ / f = 16 777 216 / 72 × 10⁶ = 0.233 s.- Cortex-M4: 10 000 cycles;
t = 10 000 / 72 × 10⁶ = 138.9 µs. - AVR: 4 cycles per addition → 40 000 cycles;
t = 40 000 / 16 × 10⁶ = 2.5 ms. - Ratio = 2.5 ms / 138.9 µs = 18 — a factor of 4.5 from clock and 4 from data width.
- Latency =
12 / 72 × 10⁶ = 166.7 ns.
Answers: RELOAD = 71 999; T_max ≈ 0.233 s; 138.9 µs vs 2.5 ms (M4 is 18× faster); latency ≈ 167 ns.
Example 3 (CPI). A Cortex-M3 at 40 MHz executes 32 × 10⁶ instructions in 1 s. CPI = 40 × 10⁶ × 1 / 32 × 10⁶ = 1.25, i.e. 32 MIPS. (A CPI far above 1–2 on these cores usually means Flash wait states or a wrong measurement.)
Common mistakes
- Calling Cortex-M a "microcontroller": it is a processor core; the microcontroller is the vendor's chip around it.
- Forgetting that SysTick period is (RELOAD + 1)/f — loading f × T gives one extra count.
- Loading a SysTick value above 2²⁴ − 1; slow ticks need a prescaled timer instead.
- Assuming an 8-bit AVR does 32-bit maths in one instruction. Multi-byte arithmetic multiplies cycle counts.
- Assuming all Cortex-M are Harvard: M0/M0+ use a single (von Neumann) bus.
- Writing assembly-style "save registers" code in Cortex-M ISRs; the hardware already stacks R0–R3, R12, LR, PC and xPSR.
- Assuming ARM always draws more power than AVR — compare actual datasheet figures.
For GATE ME
Microcontroller architecture appears in the mechatronics and instrumentation parts of the syllabus mainly as conceptual MCQs (RISC vs CISC, Harvard vs von Neumann, register widths, role of the interrupt controller) and short numericals on clock period, execution time, CPI/MIPS and timer reload values. Practise converting between cycles, frequency and time without unit slips (ns, µs, ms) and counting cycles for multi-byte operations.
Quick check
- How many general-purpose registers does an AVR have, and which pairs form X, Y and Z?
- What does the NVIC do automatically on interrupt entry?
- A Cortex-M0+ runs at 48 MHz. What SysTick reload value gives a 1 ms tick?
- Which Cortex-M cores offer a hardware FPU?
- How long does a 3-cycle instruction take on a 20 MHz AVR?
Answers: 1. 32 (R0–R31); R26:R27, R28:R29, R30:R31 2. Stacks R0–R3, R12, LR, PC, xPSR, fetches the vector and handles priority/nesting 3. 47 999 4. M4 (optional, single precision), M7 (optional, single or double), and M33 (optional) 5. 150 ns
Interview questions
All Microcontrollers, PLC and Industrial Automation interview questionsTry answering each one aloud before you open it.
1.What is the ARM Cortex-M architecture, and what are its key features?Concept
The ARM Cortex-M architecture is a family of 32-bit RISC microprocessor cores optimized for low-cost and energy-efficient applications. Key features include a simplified instruction set, low power consumption, high performance, and integrated debug support. It is widely used in embedded systems, such as automotive, industrial control, and consumer electronics.
2.Explain the AVR architecture and its typical applications.Concept
The AVR architecture is a family of microcontrollers developed by Atmel, now part of Microchip Technology. It is based on an 8-bit RISC architecture and is known for its simplicity and efficiency. AVR microcontrollers are commonly used in applications like home automation, robotics, and small-scale industrial control due to their ease of use and low power consumption.
3.How does the ARM Cortex-M architecture differ from the AVR architecture?Concept
Both are load–store RISC designs, but AVR is an 8-bit core with 32 eight-bit registers, while Cortex-M is a 32-bit core with sixteen 32-bit registers and the Thumb/Thumb-2 instruction set. Cortex-M has a fixed 4 GB memory map, the NVIC with programmable priorities, nesting and hardware register stacking, and the SysTick timer; AVR has a simpler fixed-priority vector table that does not nest by default. A 32-bit operation is one instruction on Cortex-M but several on AVR, so Cortex-M suits heavier arithmetic, larger memory and communication stacks, while AVR suits small, low-cost control jobs.
4.Why is ARM Cortex-M preferred in industrial automation applications?Application
ARM Cortex-M is preferred in industrial automation due to its high performance, low power consumption, and integrated features like real-time processing and advanced debugging capabilities. These features make it suitable for complex control systems and real-time data processing, which are common in industrial automation.
5.What happens if you use an AVR microcontroller in a high-performance application?Application
Using an AVR microcontroller in a high-performance application may lead to insufficient processing power and memory, resulting in slower execution and potential system bottlenecks. AVR microcontrollers are designed for simpler tasks, so they may not handle complex algorithms or large data efficiently.
6.Explain the role of the NVIC in ARM Cortex-M microcontrollers.Concept
The Nested Vectored Interrupt Controller (NVIC) in ARM Cortex-M microcontrollers manages interrupt handling. It allows for nested interrupts, prioritization, and efficient context switching, enabling real-time processing and responsiveness in embedded systems. NVIC is crucial for applications requiring precise timing and quick response to external events.
7.How does the power consumption of ARM Cortex-M compare to AVR microcontrollers?Application
It depends on the specific chips, not just on the core family. Small cores such as the Cortex-M0+ achieve very low active current per MHz and deep-sleep currents comparable to low-power AVRs, and because a 32-bit core finishes a computation in fewer cycles it can go back to sleep sooner. High-end Cortex-M4/M7 parts running at hundreds of MHz draw far more than an ATmega. The right comparison is the datasheet's µA/MHz and sleep-mode currents for the workload's duty cycle.
8.Calculate the clock cycles required for an ARM Cortex-M processor to execute a loop with 100 iterations, assuming each iteration takes 4 cycles.Numerical
To calculate the total clock cycles required, multiply the number of iterations by the cycles per iteration: 100 iterations × 4 cycles/iteration = 400 cycles.
9.If an AVR microcontroller operates at 16 MHz, how long does it take to execute an instruction that requires 2 clock cycles?Numerical
The time for one clock cycle is the inverse of the frequency: 1 / 16 MHz = 62.5 ns. Therefore, for 2 clock cycles, the time is 2 × 62.5 ns = 125 ns.
10.What are the advantages of using ARM Cortex-M microcontrollers in IoT devices?Application
Cortex-M parts combine a 32-bit core with low-power sleep modes, enough Flash and RAM for network and security stacks, and fast wake-up through the NVIC. Many vendor chips built on them integrate radios (BLE, Wi-Fi, LoRa), crypto accelerators and, on ARMv8-M cores, TrustZone isolation. A common toolchain and RTOS support across vendors also speeds development.
Finished this topic? Mark it so your progress, study plan and readiness keep up.
Stuck on something here?