Queuing theory: single-server models

Kendall notation, M/M/1 assumptions and steady-state measures, Little's law, effect of service variability (M/D/1, M/G/1) and economic design, with tool-crib, cost and constant-service numericals.

Drafted with Aria, reviewed by the AiCanCode.org team. Spotted an error? Use Give Feedback at the bottom of the page.

Why it matters

Queues are everywhere in a vehicle plant and its service network: machines waiting for a repair crew, mechanics waiting at a tool crib, cars waiting for a single washing bay or an end-of-line test rig, trucks waiting at an unloading dock. Single-server queuing models tell you how long the queue will be and how long people or machines will wait, so you can decide whether a faster server or a second one pays for itself. M/M/1 calculations and Little's law are frequent short GATE questions.

Key ideas

Elements of a queuing system. Calling population (finite or infinite), arrival process, queue (capacity, discipline), service mechanism (number of servers, service-time distribution).

Kendall notation A/B/c/K/N/D. A = arrival distribution, B = service-time distribution (M = Markovian/exponential, D = deterministic, G = general), c = number of servers, K = system capacity, N = population size, D = discipline. M/M/1 usually means M/M/1/∞/∞/FCFS.

M/M/1 assumptions.

  • Arrivals are Poisson with mean rate λ (equivalently, inter-arrival times are exponential with mean 1/λ) — arrivals are independent and random.
  • Service times are exponential with mean 1/μ (service rate μ).
  • One server, first come first served, infinite queue and population.
  • Steady state exists only if ρ = λ/μ < 1. If λ ≥ μ the queue grows without limit; there is no steady state.

Behaviour near ρ = 1. Waiting grows very steeply as utilisation approaches 1: in M/M/1, L = ρ/(1 − ρ), so going from ρ = 0.8 to 0.9 more than doubles the number in the system (4 to 9). A server run at 95–100 % utilisation is therefore not "efficient" — it creates long waits. This is why maintenance crews and test rigs are sized with spare capacity.

Little's law. For any stable queuing system, the average number present equals the arrival rate times the average time spent: L = λW and L_q = λW_q. It does not depend on distributions or discipline.

Effect of variability. For the same λ and μ, a constant service time (M/D/1) halves the average queue length of M/M/1. For a general service-time distribution (M/G/1), the Pollaczek–Khinchine formula shows L_q rises with the service-time variance. Standardising work (for example a fixed-cycle automatic wash) shortens queues even without increasing μ.

Economic design. Total cost per unit time = cost of service capacity + cost of waiting (idle mechanics, idle machines, lost customers). Faster or extra servers cost more but cut waiting cost; the best choice minimises the sum.

Formulas

All for M/M/1, steady state, ρ < 1. λ and μ in the same time unit (e.g. per hour); L in units (customers, machines), W in that time unit.

ρ = λ / μ

  • Server utilisation = probability the server is busy (–).

P₀ = 1 − ρ, P_n = (1 − ρ) · ρⁿ, P(n ≥ k) = ρᵏ

  • Probability of an empty system, of exactly n in the system, and of k or more in the system.

L = λ / (μ − λ) = ρ / (1 − ρ)

  • Average number in the system (in queue + in service).

L_q = λ² / (μ(μ − λ)) = ρ² / (1 − ρ)

  • Average number waiting in the queue.

W = 1 / (μ − λ), W_q = λ / (μ(μ − λ))

  • Average time in the system and in the queue.

L = λ·W, L_q = λ·W_q, W = W_q + 1/μ

  • Little's law and the link between W and W_q (valid generally).

L_n = μ / (μ − λ) = 1 / (1 − ρ)

  • Average length of a non-empty queue (M/M/1).

L_q = (λ²σ² + ρ²) / (2(1 − ρ))

  • M/G/1 (Pollaczek–Khinchine); σ = standard deviation of service time. For M/D/1, σ = 0 so L_q = ρ² / (2(1 − ρ)).

Worked examples

Example 1 (standard) — tool crib. Given: mechanics arrive at a single-attendant tool crib at λ = 12 per hour (Poisson); service is exponential with μ = 15 per hour.

  1. ρ = 12/15 = 0.8; P₀ = 0.2 (attendant idle 20 % of the time)
  2. L = 12 / (15 − 12) = 4 mechanics
  3. L_q = 12² / (15 × 3) = 144 / 45 = 3.2 mechanics
  4. W = 1 / 3 h = 20 min; W_q = 12 / (15 × 3) = 0.267 h = 16 min
  5. Check: L = λW = 12 × (1/3) = 4 ✓; W − W_q = 4 min = 1/μ ✓
  6. P(more than 3 in system) = P(n ≥ 4) = 0.8⁴ = 0.41

Answer: 4 mechanics in the system on average, each losing 20 min per visit.

Example 2 (GATE level) — choosing a faster server. Given: as above; a mechanic's time costs ₹300/h. Option 1: present attendant, ₹200/h, μ = 15/h. Option 2: a trained attendant with a better layout, ₹260/h, μ = 20/h. Which is cheaper per hour?

  1. Option 1: L = 4 → waiting cost = 4 × 300 = ₹1,200/h; total = 1,200 + 200 = ₹1,400/h.
  2. Option 2: ρ = 0.6, L = 12 / (20 − 12) = 1.5 → waiting cost = 1.5 × 300 = ₹450/h; total = 450 + 260 = ₹710/h.

Answer: Option 2 saves ₹690 per hour even though the attendant costs more.

Example 3 (GATE level) — effect of constant service time. Given: λ = 12 per hour as before, but the crib is automated so every service takes exactly 4 min (μ = 15/h, σ = 0).

  1. L_q = ρ² / (2(1 − ρ)) = 0.64 / 0.4 = 1.6
  2. W_q = L_q / λ = 1.6 / 12 h = 8 min

Answer: queue and waiting time are half the M/M/1 values (1.6 versus 3.2; 8 min versus 16 min).

Common mistakes

  • Mixing time units: λ per hour with μ per minute. Convert both to the same unit first.
  • Taking the service time as μ; μ is a rate. A mean service time of 4 min means μ = 15 per hour.
  • Applying M/M/1 formulas when λ ≥ μ; there is no steady state.
  • Confusing L (in system) with L_q (waiting only); they differ by ρ, not by 1.
  • Forgetting that P(n ≥ k) = ρᵏ, not ρᵏ⁺¹; "more than k" means n ≥ k + 1.
  • Using M/M/1 results for a deterministic or low-variability server, which overestimates the queue.

For GATE ME

Expect M/M/1 numericals on ρ, P₀, L, L_q, W, W_q and probabilities of n customers, often with a unit conversion between minutes and hours, and questions where you work back from one measure (for example L_q) to λ or μ. Little's law appears in one-mark conceptual questions. Practise writing all results from ρ alone: L = ρ/(1 − ρ), L_q = ρ²/(1 − ρ).

Quick check

  1. λ = 4/h, μ = 10/h. Find ρ and the probability the server is idle.
  2. λ = 3/h, μ = 7/h. Find L.
  3. λ = 2/h, μ = 5/h. Find W_q in minutes.
  4. A shop holds on average 6 jobs and receives 2 jobs per hour. Average time a job spends in the shop?
  5. If λ = μ in M/M/1, what happens to the queue?

Answers: 1. 0.4; 0.6. 2. 0.75. 3. 2/(5 × 3) = 0.133 h = 8 min. 4. 3 h (Little's law). 5. It grows without limit — no steady state.

Try answering each one aloud before you open it.

  1. 1.What is a single-server queuing model in queuing theory?Concept

    A single-server queuing model is a mathematical representation of a system where there is only one server providing service to incoming customers or jobs. It is used to analyze the behavior of queues, including metrics like average wait time, queue length, and server utilization. The model typically assumes that arrivals follow a Poisson process and service times are exponentially distributed.

  2. 2.Explain the significance of the arrival rate (λ) and service rate (μ) in a single-server queuing model.Concept

    In a single-server queuing model, the arrival rate (λ) represents the average number of customers arriving per time unit, while the service rate (μ) is the average number of customers that can be served per time unit. These rates are crucial for determining the system's performance, such as the average number of customers in the system, the average time a customer spends in the system, and the probability of the server being idle.

  3. 3.Why is the Poisson process often used to model arrivals in queuing systems?Application

    The Poisson process is used to model arrivals in queuing systems because it is a simple and mathematically tractable way to represent random arrival events. It assumes that arrivals are independent and occur at a constant average rate, which is a reasonable approximation for many real-world systems. This process allows for the derivation of key performance metrics and simplifies the analysis of queuing models.

  4. 4.What happens to the queue length and waiting time if the arrival rate exceeds the service rate in a single-server model?Application

    If the arrival rate (λ) exceeds the service rate (μ) in a single-server model, the queue length and waiting time will grow indefinitely over time. This is because the server cannot keep up with the incoming demand, leading to an accumulation of customers in the queue. In practical terms, this situation is unsustainable and indicates that the system is overloaded.

  5. 5.How does the utilization factor (ρ) affect the performance of a single-server queuing system?Application

    The utilization factor (ρ) is the ratio of the arrival rate (λ) to the service rate (μ), expressed as ρ = λ/μ. It indicates the proportion of time the server is busy. A higher utilization factor means the server is busier, which can lead to longer queues and increased waiting times. Ideally, ρ should be less than 1 to ensure the system is stable and can handle the incoming demand without excessive delays.

  6. 6.Explain the concept of Little's Law in the context of queuing theory.Concept

    Little's Law is a fundamental theorem in queuing theory that relates the average number of customers in a system (L) to the average arrival rate (λ) and the average time a customer spends in the system (W). It is expressed as L = λW. This law holds for a wide range of queuing systems, regardless of the arrival process or service distribution, and provides a simple way to calculate one of these metrics if the other two are known.

  7. 7.What is the impact of variability in service times on the performance of a single-server queuing system?Application

    Variability in service times can significantly impact the performance of a single-server queuing system. Higher variability can lead to longer queues and increased waiting times, as the server may experience periods of high demand that exceed its capacity. This variability can also cause fluctuations in server utilization, making it more challenging to predict system performance and optimize resource allocation.

  8. 8.Calculate the average number of customers in the system (L) if the arrival rate (λ) is 5 customers per hour and the service rate (μ) is 8 customers per hour.Numerical

    To calculate the average number of customers in the system (L), we use the formula L = λ / (μ - λ). Substituting the given values, L = 5 / (8 - 5) = 5 / 3 ≈ 1.67 customers.

  9. 9.Determine the average waiting time in the queue (Wq) if the arrival rate (λ) is 3 customers per hour and the service rate (μ) is 6 customers per hour.Numerical

    The average waiting time in the queue (Wq) can be calculated using the formula Wq = λ / (μ(μ - λ)). Substituting the given values, Wq = 3 / (6(6 - 3)) = 3 / 18 = 0.167 hours or 10 minutes.

  10. 10.How can a single-server queuing model be used to improve the efficiency of a production line?Application

    A single-server queuing model can be used to analyze and optimize the flow of products through a production line by identifying bottlenecks and estimating the impact of changes in arrival and service rates. By understanding the relationship between arrival rates, service rates, and queue lengths, managers can make informed decisions about resource allocation, such as adding more servers or adjusting work schedules, to improve overall efficiency and reduce waiting times.

Finished this topic? Mark it so your progress, study plan and readiness keep up.

Stuck on something here?