INPUT
Stereo
Synchronized cameras converted to depth
INTERACTIVE RETAIL · COMPUTER VISION · TECHNICAL PROTOTYPE
We delivered the computer-vision layer for a gesture-controlled marketing kiosk, using stereo depth to isolate the interaction zone and pose estimation to produce stable 2.5D body tracking.

INPUT
Stereo
Synchronized cameras converted to depth
OUTPUT
2.5D skeleton
Image-space keypoints with usable depth
DELIVERY
Accepted milestones
Background removal, pose and user tracking
Accepted outcome
THE PROBLEM AND HYPOTHESIS
An advertising kiosk can stand in a busy retail or exhibition environment. Several people may be visible at once, and the closest detection in the image is not always the person occupying the intended physical interaction zone.
The client already had stereo-camera hardware and planned downstream gesture recognition and augmented rendering. The missing layer was a stable representation of the active user: body keypoints connected to depth and retained across frames.
PERCEPTION PIPELINE
One annotated pipeline connects the stereo hardware input to the 2.5D output accepted by the downstream interaction system.
01 · Test stage
Synchronised stereo frames
02 · Test stage
Rectification
03 · Test stage
Disparity and depth
04 · Test stage
Range isolation
05 · Test stage
Pose estimation
06 · Test stage
Active-user state
07 · Test stage
2.5D output
ENGINEERING DECISIONS
Background removal reduces the number of irrelevant people presented to the pose model and makes the active-user rule relate to physical distance rather than apparent image size.
Computer vision
The downstream experience needed stable spatial gesture data, not a research-grade reconstruction of the entire body. Combining 2D keypoints with stereo depth reduced technical risk and fitted the hardware already chosen.
Scope judgement
The engagement established the three agreed perception milestones against the supplied hardware and acceptance videos. It did not establish a general-purpose people-tracking platform for arbitrary cameras or environments.
RESPONSIBILITY
Three accepted technical milestones
SELECTED TECHNOLOGY
RELEVANT EXPERIENCE
You need to de-risk the perception layer before building a larger physical product.
You need computer vision to work with existing cameras, GPUs and downstream software.
You need stable spatial tracking in a scene containing several people.
RELATED WORK

AURRENCY
Client-specific anti-theft logic delivered through shared camera, job, scheduling and evidence infrastructure.
View case study
GKM TECH
Customer platform, video-processing backend and specialist annotation workflow.
View case studySTART A CONVERSATION
We can isolate the risky perception problem, test it against the real hardware and provide a clear path into the wider product.