Real-time operating system basics

RTOS fundamentals: hard vs soft real time, tasks and states, pre-emptive priority scheduling, semaphores, mutexes and queues, priority inversion and deadlock, with utilisation, rate-monotonic and response-time calculations.

Drafted with Aria, reviewed by the AiCanCode.org team. Spotted an error? Use Give Feedback at the bottom of the page.

Why it matters

A modern controller does many things at once: a 1 kHz motor-current loop, a 100 Hz sensor fusion task, a fieldbus stack, a display and a data logger. A simple super-loop cannot guarantee that the fast loop always runs on time while the logger writes to flash. A real-time operating system (RTOS) such as FreeRTOS, Zephyr or VxWorks splits the work into prioritised tasks and gives provable timing — the foundation of drones, robot joints, drives and PLC runtimes.

Key ideas

Real-time means predictable, not fast. A real-time system is correct only if results are both right and on time. Hard real-time: a missed deadline is a failure (airbag, motor commutation, safety interlock). Soft real-time: occasional lateness only degrades quality (video, HMI refresh). Firm: late results are useless but not dangerous.

Super-loop vs RTOS. A super-loop (while(1) calling functions, plus ISRs) is simple and fine for small jobs, but every function's worst-case time adds to every other's response time. An RTOS kernel lets the most urgent ready task run as soon as it becomes ready.

Tasks and states. A task (thread) is an independent function with its own stack and priority. States: running (on the CPU), ready (waiting for the CPU), blocked (waiting for time, a semaphore, a queue message or an event), suspended. Tasks should block when they have nothing to do instead of busy-waiting.

Scheduler. Most RTOSes use fixed-priority pre-emptive scheduling: the highest-priority ready task always runs, and if a higher-priority task becomes ready (because of an interrupt, a timeout or a message) it pre-empts the running one at once. Tasks of equal priority may share the CPU by round-robin time slicing. The kernel keeps time with a periodic tick interrupt (often 1 kHz, from SysTick on Cortex-M).

Context switch. Saving the running task's registers and stack pointer and restoring another's. It costs a few microseconds; frequent switches are overhead.

Inter-task communication and synchronisation.

  • Queue: copies data (messages) from producer to consumer safely; the reader blocks until data arrives.
  • Binary semaphore: signalling, e.g. an ISR "gives" it and a task "takes" it — the standard way to defer interrupt work to a task.
  • Counting semaphore: counts events or free resources.
  • Mutex: mutual exclusion for a shared resource (a UART, an I2C bus, a global struct); it has an owner and usually priority inheritance.
  • Event flags / notifications: wait for one or several conditions.

Classic problems.

  • Race condition: two tasks update shared data without protection.
  • Priority inversion: a high-priority task waits on a mutex held by a low-priority task, which is itself pre-empted by a medium-priority task — so the high-priority task waits for an unrelated one. Priority inheritance (temporarily raising the holder's priority) bounds this; it famously reset the Mars Pathfinder lander until the option was enabled.
  • Deadlock: two tasks each hold a resource the other needs. Avoid by acquiring resources in a fixed order and using timeouts.

Scheduling analysis. Each periodic task i has worst-case execution time C_i and period T_i (deadline usually = period). Rate-monotonic (RM) priority assignment — shorter period, higher priority — is optimal among fixed-priority schemes. Liu and Layland's bound gives a quick sufficient test; response-time analysis gives an exact test for fixed priorities. Earliest deadline first (EDF), a dynamic-priority scheme, can use up to 100 % of the CPU.

ISRs in an RTOS. Keep ISRs very short: read the hardware, clear the flag, give a semaphore or post to a queue using the ISR-safe API, and let a task do the processing.

Memory. Each task needs its own stack, sized for its deepest call chain plus interrupt frames. Safety-critical code usually allocates all objects statically at start-up and avoids heap allocation during operation to prevent fragmentation and non-deterministic timing.

Formulas

U = Σ (C_i / T_i)

  • U = total CPU utilisation (dimensionless); C_i = worst-case execution time of task i (s); T_i = period (s).

U ≤ n (2^(1/n) − 1)

  • Liu–Layland bound for rate-monotonic scheduling of n independent periodic tasks with deadline = period. Sufficient, not necessary. Values: n = 1 → 1.0, 2 → 0.828, 3 → 0.780, large n → ln 2 ≈ 0.693.

U ≤ 1

  • Necessary and sufficient condition for EDF (pre-emptive, deadline = period); also a necessary condition for any scheduler.

R_i = C_i + Σ_(j ∈ hp(i)) ⌈R_i / T_j⌉ × C_j

  • Response time of task i (s) under fixed priorities; hp(i) = tasks with higher priority. Iterate starting from R_i = C_i until the value repeats; the task is schedulable if R_i ≤ deadline.

Overhead = n_switch × t_cs

  • Fraction of CPU used by context switching; n_switch = switches per second; t_cs = time per switch (s).

Worked examples

Example 1 (standard). Three periodic tasks under rate-monotonic scheduling: τ1 (C = 1 ms, T = 4 ms), τ2 (C = 2 ms, T = 8 ms), τ3 (C = 1 ms, T = 10 ms). Is the set schedulable by the utilisation test?

  1. U = 1/4 + 2/8 + 1/10 = 0.25 + 0.25 + 0.10 = 0.60.
  2. Bound = 3 × (2^(1/3) − 1) = 0.780.
  3. 0.60 ≤ 0.780 → the test passes.

Answer: U = 0.60 ≤ 0.780, schedulable under RM.

Example 2 (GATE level). τ1 (C = 1, T = 4), τ2 (C = 2, T = 6), τ3 (C = 3, T = 12), all in ms, RM priorities (τ1 highest), deadline = period. (a) Apply the utilisation test. (b) Find the response time of each task. (c) A 1 kHz tick causes one context switch per tick at 10 µs each; find the overhead.

  1. U = 1/4 + 2/6 + 3/12 = 0.250 + 0.333 + 0.250 = 0.833 > 0.780 → the sufficient test is inconclusive (but U < 1, so the set is not ruled out).
  2. τ1: R1 = C1 = 1 ms ≤ 4 ✓.
  3. τ2: start R = 2 → R = 2 + ⌈2/4⌉ × 1 = 3 → R = 2 + ⌈3/4⌉ × 1 = 3 (converged). R2 = 3 ms ≤ 6 ✓.
  4. τ3: R = 3 → 3 + ⌈3/4⌉1 + ⌈3/6⌉2 = 6 → 3 + ⌈6/4⌉1 + ⌈6/6⌉2 = 7 → 3 + 2 + ⌈7/6⌉2 = 9 → 3 + ⌈9/4⌉1 + ⌈9/6⌉2 = 3 + 3 + 4 = 10 → 3 + ⌈10/4⌉1 + ⌈10/6⌉2 = 10 (converged). R3 = 10 ms ≤ 12 ✓.
  5. Overhead = 1000 × 10 µs = 10 ms per second = 1 %.

Answers: U = 0.833 (bound test inconclusive); R1 = 1 ms, R2 = 3 ms, R3 = 10 ms — all deadlines met; tick overhead 1 %.

Common mistakes

  • Treating the Liu–Layland bound as necessary: failing it does not prove the set unschedulable — do response-time analysis.
  • Using average instead of worst-case execution times.
  • Giving priority by "importance" instead of rate; RM assigns higher priority to shorter periods.
  • Using a binary semaphore where a mutex is needed (no ownership, no priority inheritance).
  • Calling blocking or non-ISR-safe RTOS functions from an interrupt.
  • Busy-wait delays inside tasks, which starve lower-priority tasks.
  • Under-sized task stacks — overflow corrupts neighbouring memory silently.

For GATE ME

RTOS ideas appear mostly as conceptual MCQs (hard vs soft real time, pre-emption, semaphore vs mutex, priority inversion, deadlock, task states) and short numericals on utilisation, the rate-monotonic bound and task frequency/period. Practise computing U with mixed units and remembering the bound values for small n.

Quick check

  1. What is the RM utilisation bound for two tasks?
  2. A task runs every 10 ms for 2 ms. What is its utilisation?
  3. Which kernel object provides priority inheritance?
  4. In which state is a task waiting for a queue message?
  5. Can EDF schedule a set with U = 0.95 (deadline = period)?

Answers: 1. 0.828 2. 0.2 (20 %) 3. Mutex 4. Blocked 5. Yes, since U ≤ 1

Try answering each one aloud before you open it.

  1. 1.What is a real-time operating system (RTOS)?Concept

    A real-time operating system (RTOS) is a specialized operating system designed to manage hardware resources and run applications with precise timing and high reliability. It is used in environments where tasks must be executed within strict time constraints, such as in embedded systems, industrial automation, and robotics.

  2. 2.Explain the difference between a hard real-time system and a soft real-time system.Concept

    In a hard real-time system, tasks must be completed within a strict deadline, and failure to do so can lead to catastrophic consequences. Examples include pacemakers and automotive airbag systems. In contrast, a soft real-time system allows for some flexibility in meeting deadlines, where occasional delays are tolerable, such as in video streaming or online gaming.

  3. 3.Why is an RTOS preferred over a general-purpose operating system in industrial automation?Application

    An RTOS is preferred in industrial automation because it provides deterministic behavior, ensuring that tasks are executed within predictable time frames. This is crucial for maintaining the precision and reliability required in automated processes, where timing is critical for safety and efficiency.

  4. 4.What happens if a task in a hard real-time system misses its deadline?Application

    If a task in a hard real-time system misses its deadline, it can lead to system failure or catastrophic consequences. For example, in an automotive airbag system, a missed deadline could result in the airbag not deploying in time during a collision, potentially causing harm to the occupants.

  5. 5.How does task scheduling work in an RTOS?Concept

    Task scheduling in an RTOS is typically priority-based, where tasks are assigned priorities, and the scheduler ensures that the highest-priority task is executed first. Some RTOSs use preemptive scheduling, allowing a higher-priority task to interrupt a lower-priority one, ensuring timely task execution.

  6. 6.Explain the role of interrupts in an RTOS.Concept

    Interrupts in an RTOS are signals that inform the processor of an event that requires immediate attention. They allow the RTOS to respond quickly to external events by temporarily halting the current task and executing an interrupt service routine (ISR) to handle the event, ensuring timely processing.

  7. 7.Why is memory management important in an RTOS?Application

    Each task has its own stack, and an overflowed stack silently corrupts other tasks' data, so stacks must be sized for the worst-case call depth plus interrupt frames and checked (most RTOSes offer stack watermarks or overflow hooks). General-purpose malloc/free is non-deterministic in time and fragments a small heap, so real-time and safety-critical code allocates tasks, queues and buffers statically or at start-up, or uses fixed-size memory pools. Memory protection units, where available, isolate tasks from one another.

  8. 8.What is the impact of context switching on the performance of an RTOS?Application

    Context switching in an RTOS involves saving the state of a currently running task and loading the state of the next task to be executed. While necessary for multitasking, frequent context switching can lead to performance overhead, as it consumes CPU time and resources, potentially affecting system responsiveness.

  9. 9.Calculate the context switch time if the CPU takes 5 microseconds to save the state and 3 microseconds to load the state.Numerical

    The context switch time is the sum of the time taken to save the state and the time taken to load the state. Therefore, context switch time = 5 microseconds + 3 microseconds = 8 microseconds.

  10. 10.If a task in an RTOS has a period of 10 ms and an execution time of 2 ms, what is its utilization?Numerical

    Utilization is calculated as the ratio of the task's execution time to its period. Utilization = Execution time / Period = 2 ms / 10 ms = 0.2 or 20%.

Finished this topic? Mark it so your progress, study plan and readiness keep up.

Stuck on something here?