ToF distance sensing
Measures nearby obstacles in milliseconds and activates haptic warnings without Wi-Fi or cloud AI.
- Distance classification
- Configurable thresholds
- Always-on local feedback
VisionCane is a safety-first smart cane concept for blind and low-vision people. Local distance sensing warns about nearby obstacles in real time, while AI adds useful context about objects, pathways, signs, and surroundings.
The need
A white cane is an essential mobility tool, but physical contact alone cannot name what is ahead, read a sign, or explain the wider scene. VisionCane is designed to add that missing layer without taking control away from the user.
How it works
Immediate obstacle sensing stays on the device. Visual interpretation runs only when requested. The ESP32-S3 coordinates both paths and prioritizes local warnings.
Measures nearby obstacles in milliseconds and activates haptic warnings without Wi-Fi or cloud AI.
Runs the state machine, coordinates camera and sensors, manages Wi-Fi, and gives safety alerts priority over narration.
Captures a selected frame and returns concise scene context, object identity, OCR, or an answer to a question.
Safety first
VisionCane is designed so that network loss, API timeouts, or uncertain AI responses never disable the local ToF-to-haptic warning loop.
The ToF sensor and vibration motor do not wait for an internet round trip.
A nearby obstacle takes precedence even while an AI response is being spoken.
The system is designed to say when it cannot identify an object confidently.
Core capabilities
The experience is intentionally concise: one press, one useful answer, and no constant stream of narration.
Configurable distance zones translate into distinct vibration patterns, giving immediate physical feedback without relying on the cloud.
A short press captures a frame for AI analysis, prioritizing mobility-relevant objects such as people, vehicles, doors, stairs, curbs, and pathways.
When asked to read, VisionCane captures a frame and returns visible text—or clearly says when the text is not legible.
A long press enters question mode so the user can ask what is ahead or request specific visual context.
Hardware system
The prototype architecture uses widely available, low-power parts. Together they create an independent safety path and an on-demand perception path.
The main brain: coordinates sensors, camera, Wi-Fi, buttons, audio, haptics, and the runtime state machine.
Why it fits: compact, low power, camera-capable, connected, and suitable for reliable embedded control.
Captures a forward-facing JPEG frame when the user requests a scan, reading task, or scene question.
Why it fits: proven ESP32 support, manageable image size, and enough detail for cloud scene analysis.
Measures the actual distance to nearby obstacles and triggers haptic warnings locally—even with no internet.
Why it fits: fast, compact range sensing that does not confuse AI-estimated distance with measured distance.
Turns distance zones into tactile patterns: one pulse for nearby, two pulses for warning, and rapid vibration for critical proximity.
Why it fits: immediate, private, and usable when audio is busy or the environment is noisy.
Short press: scan. Long press: question mode. Double press: repeat the last response when enabled.
Why it fits: direct, discoverable control that does not depend on a touchscreen or precise gestures.
Speaks short scene summaries while leaving the ears open to traffic, voices, bicycles, and other environmental cues.
Why it fits: sealed headphones could mask critical ambient sound; open-ear output preserves awareness.
Supplies regulated portable power to the controller, camera, sensors, haptics, and audio. The prototype target is 4–8 hours, subject to physical testing.
Why it fits: rechargeable energy density for a compact handle, with protection and regulation instead of direct battery-to-GPIO wiring.
Monitors the firmware and resets the controller if a task stalls, helping the device recover from camera, network, or software faults.
Why it fits: embedded safety functions must recover rather than remain frozen after a transient failure.
Sensor fusion
Fusion happens at the ESP32-S3. The ToF measurement always remains authoritative for distance.
VL53L1X provides measured proximity.
Gemini identifies the object and direction.
“Motorcycle nearby.”
Roadmap
The architecture is defined. The next work is disciplined prototyping, testing, and learning with users—not claiming a finished product before it is ready.
Safety invariants, state machine, hardware roles, failure behavior, and acceptance criteria.
DefinedESP32-S3 firmware, ToF-to-haptic loop, camera capture, AI response parsing, and open-ear audio.
NextOffline operation, timeout recovery, malformed AI output, warning priority, power, and physical durability.
PlannedValidate language, haptic patterns, comfort, usefulness, and failure handling with expert supervision.
PlannedEarly collaboration
We are looking to connect with accessibility experts, embedded engineers, mobility specialists, and future pilot partners who want to help test the assumptions behind VisionCane.