Robot sensors and vision
Internal and external robot sensors (encoders, resolvers, IMUs, proximity, range, tactile, force/torque) and the machine-vision pipeline with pinhole and stereo geometry.
Drafted with Aria, reviewed by the AiCanCode.org team. Spotted an error? Use Give Feedback at the bottom of the page.
Why it matters
A robot without sensors repeats taught motions blindly; parts must then be presented perfectly every time. Sensors close the servo loops inside every joint, detect parts and people, measure contact forces, and — with machine vision — let the robot find, inspect and pick parts whose position varies. Sensor choice often decides whether an automation project is feasible and affordable.
Key ideas
Internal (proprioceptive) sensors measure the robot's own state.
- Optical incremental encoder — a slotted disc gives N pulses per revolution on two channels A and B in quadrature (90° apart). Counting all edges gives 4N counts/rev and the phase order gives direction. Needs homing at power-up.
- Absolute encoder — each position has a unique code (Gray code or multi-turn electronic); an n-bit encoder resolves 360°/2ⁿ and knows its position at power-up.
- Resolver — rotary transformer; rugged, analog, used in harsh environments.
- Tachogenerator — voltage proportional to speed (now mostly replaced by differentiating encoder counts).
- Potentiometer — cheap absolute position; wear and noise limit it.
- IMU — accelerometers (linear acceleration and gravity direction) plus gyroscopes (angular rate); gyros drift, so they are fused with other sensors.
External (exteroceptive) sensors measure the environment.
- Proximity sensors (non-contact presence): inductive (metals only, a few mm), capacitive (any material, including liquids), photoelectric (through-beam, retro-reflective, diffuse), ultrasonic and infrared ranging.
- Range sensors: ultrasonic time-of-flight (cheap, wide beam, poor angular resolution, affected by soft or angled surfaces), laser triangulation, lidar (time-of-flight laser scanning, accurate, used for mapping and safety scanners).
- Touch/tactile sensors: limit switches, tactile arrays in fingertips, slip sensors.
- Force/torque sensors: six-axis strain-gauge sensors at the wrist measure Fx, Fy, Fz, Mx, My, Mz for force control and collision detection.
- Active vs passive: active sensors emit energy (ultrasonic, lidar, structured light); passive ones only receive it (camera, thermal sensor).
Sensor specifications: range, resolution (smallest detectable change), accuracy, repeatability, linearity, sensitivity, response time/bandwidth, hysteresis.
Machine vision.
- Image acquisition — camera (CCD/CMOS), lens, controlled lighting (backlight for silhouettes, diffuse for shiny parts). Lighting is often the most important design choice.
- Pre-processing — filtering (smoothing, median), contrast enhancement, thresholding to a binary image.
- Segmentation — separate objects from background: thresholding, edge detection (Sobel, Canny), region growing.
- Feature extraction — area, centroid, perimeter, orientation (from moments), holes.
- Recognition/interpretation — template matching, feature-based classification, deep learning.
- Action — convert the pixel location to robot coordinates through camera calibration (intrinsic: focal length, principal point, distortion; extrinsic: camera pose) and hand-eye calibration (camera relative to the gripper or base).
Pinhole camera model: a point at depth Z projects to image coordinate u = f·X/Z. Depth is lost in one image; recover it with stereo vision (two cameras with baseline B; depth from disparity), structured light, or time-of-flight cameras.
Formulas
Incremental encoder: counts per rev = 4·N (quadrature), resolution = 360° / (4·N); through a gearbox of ratio G at the joint: 360° / (4·N·G)
Absolute encoder: resolution = 360° / 2ⁿ
- N — lines (pulses) per revolution; n — number of bits; G — gear ratio.
Ultrasonic / time-of-flight range: d = v·t / 2
- v — wave speed (≈ 343 m/s for sound in air at 20 °C; 3 × 10⁸ m/s for light); t — round-trip time (s).
Pinhole projection: x_image = f·X / Z
- f — focal length; X — object size or offset; Z — depth (same length units). Divide by pixel pitch to get pixels.
Stereo depth: Z = f·B / d; depth resolution: ΔZ ≈ Z² · Δd / (f·B)
- f — focal length in pixels; B — baseline (m); d — disparity (pixels).
Arc error at the tool: Δs = r · Δθ (Δθ in rad)
Worked examples
Example 1 (standard). A joint motor carries a 1000-line incremental encoder read in quadrature and drives the link through a 100:1 gearbox. Find (a) the counts per motor revolution, (b) the angular resolution at the link, (c) the corresponding arc at a tool 0.8 m from the joint axis.
- Counts/rev = 4 × 1000 = 4000.
- Link resolution = 360° / (4000 × 100) = 0.0009°.
- Δθ = 0.0009° × π/180 = 1.571 × 10⁻⁵ rad; Δs = 0.8 × 1.571 × 10⁻⁵ = 1.257 × 10⁻⁵ m.
Answer: (a) 4000 counts, (b) 0.0009°, (c) ≈ 0.0126 mm.
Example 2 (GATE level). A stereo pair has focal length f = 800 pixels and baseline B = 0.12 m. A feature shows a disparity of 32 pixels. Find (a) its depth, (b) the depth uncertainty if disparity can be wrong by ±1 pixel.
- Z = f·B / d = 800 × 0.12 / 32 = 96 / 32 = 3.0 m.
- ΔZ ≈ Z²·Δd / (f·B) = 3.0² × 1 / 96 = 9 / 96 = 0.094 m.
Answer: (a) 3.0 m, (b) about ±0.09 m — stereo accuracy worsens with the square of distance, so use a wider baseline or longer focal length for far objects.
Example 3 (ultrasonic). An echo returns after 5.8 ms. d = 343 × 0.0058 / 2 = 0.995 m.
Common mistakes
- Forgetting to halve the round-trip time in time-of-flight ranging.
- Taking encoder resolution as 360°/N when the signal is decoded in quadrature (×4), or ignoring the gear ratio.
- Confusing resolution with accuracy — a fine encoder does not fix gear backlash or link deflection.
- Using an inductive proximity sensor for plastic parts, or ultrasonic sensors on soft, sound-absorbing surfaces.
- Treating a single camera as giving depth; a single image needs a known plane or object size.
- Mixing pixel and metric units for focal length.
For GATE ME
Expect matching sensor to application (inductive vs capacitive, encoder vs resolver, active vs passive), encoder resolution numericals including quadrature and gearing, time-of-flight range, and simple vision geometry (pinhole size, stereo depth). Conceptual questions may cover the vision pipeline stages. Practise unit conversion between degrees, radians, pixels and millimetres.
Quick check
- A 12-bit absolute encoder resolves what angle?
- Which proximity sensor detects only metals?
- Echo time 10 ms in air at 343 m/s. Distance?
- Stereo: f = 600 px, B = 0.1 m, disparity 20 px. Depth?
- Is a camera an active or passive sensor?
Answers: 1. 360°/4096 ≈ 0.088°; 2. Inductive; 3. 1.715 m; 4. 3.0 m; 5. Passive (unless it projects its own light).
Interview questions
All Robotics interview questionsTry answering each one aloud before you open it.
1.What is a robot sensor and why is it important in robotics?Concept
A robot sensor is a device that detects changes in the environment and sends this information to the robot's control system. Sensors are crucial in robotics because they allow robots to perceive their surroundings, make decisions, and perform tasks accurately. Without sensors, robots would be unable to interact with their environment effectively.
2.Explain the difference between active and passive sensors in robotics.Concept
Active sensors emit energy into the environment and measure the response to gather information. Examples include sonar and lidar. Passive sensors, on the other hand, detect natural energy emitted or reflected by objects in the environment, such as cameras and microphones. The choice between active and passive sensors depends on the application and environmental conditions.
3.How does a vision sensor work in a robotic system?Concept
A vision sensor captures images or video of the environment using cameras. These images are then processed using algorithms to extract useful information, such as object recognition, distance measurement, or motion detection. Vision sensors are essential for tasks that require high precision and adaptability, such as navigation and manipulation.
4.Why are infrared sensors commonly used in obstacle detection for robots?Application
Infrared sensors are commonly used in obstacle detection because they can detect objects in various lighting conditions, including complete darkness. They work by emitting infrared light and measuring the reflection from nearby objects. This makes them reliable for detecting obstacles and avoiding collisions in dynamic environments.
5.What happens if a robot's sensor calibration is incorrect?Application
If a robot's sensor calibration is incorrect, it can lead to inaccurate readings and poor performance. For example, a miscalibrated distance sensor might report incorrect distances, causing the robot to collide with objects or fail to navigate properly. Regular calibration ensures that sensors provide accurate data, which is critical for the robot's functionality.
6.Explain how a robot uses a gyroscope sensor for navigation.Concept
A gyroscope sensor measures the rate of rotation around an axis, providing information about the robot's orientation. In navigation, this data helps maintain stability and control, allowing the robot to move accurately along a desired path. Gyroscopes are often used in conjunction with other sensors, like accelerometers, to improve navigation precision.
7.Why is sensor fusion important in robotic systems?Application
Sensor fusion is important because it combines data from multiple sensors to provide a more accurate and comprehensive understanding of the environment. This approach compensates for the limitations of individual sensors, improving reliability and performance. For example, combining data from a camera and a lidar sensor can enhance object detection and distance measurement.
8.Calculate the distance to an object if a sonar sensor emits a sound wave that returns in 0.1 seconds. Assume the speed of sound is 343 m/s.Numerical
To calculate the distance, use the formula: distance = (speed of sound × time) / 2. The time is divided by 2 because the sound wave travels to the object and back. So, distance = (343 m/s × 0.1 s) / 2 = 17.15 meters.
9.What are the advantages of using lidar sensors over cameras in autonomous vehicles?Application
Lidar sensors provide precise distance measurements and can create detailed 3D maps of the environment, which are crucial for navigation and obstacle avoidance. Unlike cameras, lidar is not affected by lighting conditions, such as darkness or glare. This makes lidar a reliable choice for autonomous vehicles operating in diverse environments.
Finished this topic? Mark it so your progress, study plan and readiness keep up.
Stuck on something here?