3D perception
A VL53L5CX Time-of-Flight sensor turns a depth grid into nine spatial directions.
- • 4 × 4 ranging at 4 Hz
- • Interrupt with 100 ms fallback

Building an autonomous robotic companion around the ESP32-S3
Dum-E is my long-term personal robotics project: a small quadruped designed to perceive its surroundings, react with an expressive face, move organically, listen, speak, and remain maintainable through over-the-air updates. This page is the project hub and will evolve with each development milestone.
Read development phase 1Behavior preview
This hardware footage shows the implemented behavior: the depth grid selects a direction, Brain turns it into a gaze command, and LVGL interpolates the pupils toward the target.
The phase 1 article details the sensor classification, the LVGL animation, and the FreeRTOS message queues.A VL53L5CX Time-of-Flight sensor turns a depth grid into nine spatial directions.
A 240 × 240 GC9A01 round display renders vector expressions and animated gaze.
Four MG90S servos, non-blocking interpolation, and organic movement sequences.
I2S microphone and speaker paths for wake-word detection and audio feedback.
Dual application slots make room for safe OTA firmware updates.
Peripheral ICs transmit events via FreeRTOS queues. The Brain receives these events and forwards the commands.
Sensing
VL53L5CX - detection zone via I2C
Inertial measurement unit
LSM6DSOX - Tilt and movement via I2C
Microphone
INMP441 - Audio acquisition via I2S
Brain
Central thread for control
Display LCD
GC9A01 - With LVGL via SPI (DMA)
Motor kinematics
MG90 - Movements via PWM
Audio Out
I2S - Speaker
Network & OTA
Wi-Fi - Dual firmware slots
The firmware separates sensing, behavior, and rendering. Brain_thread is the sole owner of behavior state, and only the LVGL task manipulates the display.
VL53L5CX driver, interrupt-driven measurements, and spatial classification.
A deterministic event-driven controller selects gaze and facial sequences.
Vector poses are interpolated and flushed to the GC9A01 in six DMA-backed strips.
The same message-queue model is designed to integrate motor commands, audio, networking, and OTA without letting peripheral timing leak into the central behavior logic.
A dedicated thread for servo commands.
One thread for wake-word detection and another for audio output, connected through interrupts.
A final state machine coordinating perception, expression, motion, audio, and wireless maintenance.
Phase 1 — August 2026
Coupling a Time-of-Flight sensor with a vector face rendered on an LCD.
Open milestonePhase 2 — planned
Powering four servos safely and generating smooth, non-blocking quadruped motion.
Phase 3 — planned
Capturing microphone input and managing the speaker, processor load, and PSRAM memory.
Phase 4 — planned
Bringing every subsystem together and updating the robot without freezing its behavior.