Computer Vision · 2026 · Completed

Customer Behavior Analysis with YOLOv26

An end-to-end retail computer vision system that detects pre-purchase shelf behavior with a custom-trained YOLO model — mAP@50 60.91% across 4 interaction classes on 12,578 frames.

Customer Behaviour Analysis landing page with detection boxes drawn around figures in a collage of paintings
Landing page of the Customer Behaviour Analysis platform — pre-purchase behavior detection for merchandising optimization.

Summary

Problem
Most retail analytics only track completed transactions. What customers do before a sale — touching the shelf, picking up an item, removing it, or walking past without engaging — goes unmeasured.
Solution
A custom-trained Ultralytics YOLO model behind a React + FastAPI stack turns raw shelf video into interaction events, 8 operational KPIs, a planogram-based interaction overlay and product-level engagement metrics.
Result
A working end-to-end prototype built from scratch — custom dataset, annotation, tracking logic and dashboard — that surfaces attention hotspots and high-touch, low-conversion shelf positions.
mAP@50
60.91%
Precision
72.69%
Recall
58.95%
mAP@50-95
25.25%
Training frames
12,578
from 11 real retail video clips
Interaction classes
4
No Interaction, Touching Shelf, Holding Product, Item Removed

Project Summary

Customer Behaviour Analysis is a retail computer vision project that measures pre-purchase shelf behavior — the customer activity that happens before a sale is ever recorded. Most retail analytics platforms only track completed transactions. This system goes earlier in the decision funnel, detecting how customers physically interact with products: touching the shelf, picking up an item, removing it, or walking past without engaging.

Built on a custom-trained Ultralytics YOLO model and a React + FastAPI stack, the platform transforms raw shelf video into operational KPIs, a planogram-based interaction overlay, and product-level engagement metrics — giving retail managers a data layer that point-of-sale systems cannot provide.

Architecture

Data flow diagram in four columns: inputs, processing and AI, storage and APIs, frontend experience

System architecture: retail video input → YOLO inference → interaction tracking → analytics aggregation → dashboard visualization.

Detection outputs pass through an interaction-tracking layer that assigns stable interaction points to shelf zones and links them to product placement data. This enables per-product attention scoring and identification of low-conversion shelf regions — not just raw detection counts.

The FastAPI backend handles inference orchestration, event processing, analytics aggregation, and REST API delivery. The React frontend renders camera playback with shelf overlays, dot-matrix interaction maps, a planogram-based interaction overlay, a behavior-to-purchase conversion funnel, and a treemap for product-level engagement comparison across 8 operational KPIs.

Model Training

The detection model was trained on 12,578 frames extracted from 11 real retail video clips, covering 4 interaction classes: No Interaction, Touching Shelf, Holding Product, and Item Removed. Training used Ultralytics YOLO with a custom annotated dataset built specifically for shelf-facing retail environments.

Class distribution bar chart, training loss curves and two confusion matrices

Training documentation: dataset class distribution, YOLO training loss curves, and confusion matrices across the 4 interaction classes.

Evaluation

MetricValue
Precision72.69%
Recall58.95%
mAP@5060.91%
mAP@50-9525.25%

Four evaluation curves for the four interaction classes

Precision-confidence, recall-confidence, F1-confidence and precision-recall curves, showing detection performance and confidence-threshold trade-offs.

Dashboard

Live view with a store aisle video, interaction counters, legend and camera log

Live monitoring: real-time customer detection, interaction states, shelf-zone overlays and KPI summaries during video inference.

Heatmap view with a dot-matrix grid, interaction counters and legend

Heatmap view: a dot-matrix map of interaction activity across the monitored shelf area, filterable by interaction stage.

Treemap of products such as eye drops, glue and toilet paper with touch, hold and removal counts

Product overlay board: detected interactions mapped onto the physical shelf layout for per-product engagement scoring.

Funnel chart: 1,000 touches, 490 holds, 170 items removed, 45 purchases

Behavior-to-purchase funnel: from shelf touches through holds and removals to purchases, estimating conversion drop-off.

Table of products with interaction counts and summary cards above it

Product-level KPI table: touches, holds and removals per product, with highlight cards for the most-engaged items.

Outcome

The result is a working end-to-end prototype that demonstrates how computer vision can move beyond object detection into applied retail decision support. Specific outputs include identification of attention hotspots by shelf zone, detection of high-touch but low-conversion product positions, and behavioral trend analysis across interaction classes.

The project was built entirely from scratch — custom dataset, custom annotation pipeline, custom tracking logic, and custom dashboard — without relying on pre-labeled retail datasets or off-the-shelf analytics templates. A walkthrough video is on LinkedIn.

Research and Inspiration

The project was inspired by Customer Object Interaction Analytics in Retail Using YOLOv5 Object Detection.

Resources