# Shelf-Lab: what shoppers do at the shelf, from store cameras

Shelf-Lab is a [Cassi.ai](https://www.cassiai.com) product for in-store shelf measurement. It reads the store's existing security cameras and counts, shelf by shelf, who passed, who stopped for 2 seconds or more, who touched the shelf and who took the product, with every head blurred at the source. Each number links to the second of anonymized video that proves it and carries a 95% confidence interval, and the effect of an endcap, a planogram change or an in-store activation is measured against control stores with difference in differences. The vision engine is real and runs on open-source models; the study, comparison and report screens of the demo use simulated data with fictional brands and stores.

## Every shelf action gets negotiated. Most are still proven with sell-out alone.

Retailers sell shelf space and in-store media to manufacturers, and manufacturers fund endcaps, extra display points and activations. What happens in front of the shelf, between the shopper arriving and the sale, rarely makes it into that conversation.

- Topic: In-store execution and shelf decisions. Endcap, planogram change, new shelf position, in-store activation. Trade marketing, category management and retail media teams need to know what each one changed in front of the shelf, because space and co-op funds are priced on it.
- How it is done today: Sell-out, a store audit, a person with a clipboard. Sell-out shows what left the store. Audits and mystery shoppers check whether the endcap was set up. When someone asks what shoppers actually did, the team still sends an observer for a few hours or commissions an eye-tracking study with a small recruited sample.
- The pain: Nobody knows how many stopped and how many put it back. A shopper who stopped, picked up the pack, read the price and put it back leaves no trace in sell-out. Without control stores, an uplift that came from the season or a chain-wide promotion gets credited to the endcap, and the next space is priced on that number.
- Where Cassi.ai comes in: The ceiling cameras become the shelf counter. Shelf-Lab reads the CCTV already installed, blurs every head at the source, and counts passed, stopped, touched and took for each shelf. Each number opens its anonymized clip and carries a 95% confidence interval, and the effect of an action is measured against control stores.

## Shelf questions, answered from the camera and the statistics.

- Passed and stopped ("How many people who walk past this shelf actually stop?"): A pass counts when a person's feet stay in the zone in front of the shelf for 0.5 s. A stop counts when they slow below walking pace for 2 s or more. Stop rate is the second divided by the first, per shelf, per store and per hour.
- Touched and took ("Of those who stopped, how many touched, and how many took it?"): A touch is a wrist, found by pose estimation, inside the shelf area for at least 0.3 s. Took is only claimed with evidence: the shelf changes and stays changed, or the person examines the product in hand right after. In the pilot, the POS receipt closes the count.
- Compare periods ("Did the endcap work, or was it the season?"): Test stores are compared with control stores across the same two periods. The effect is the change in test stores minus the change in control stores, in percentage points, with a 95% interval. When the interval crosses zero, the report says inconclusive.
- Effect per product ("Did the new position take sales from the neighbors?"): The same comparison runs product by product, so a loss on neighboring brands shows up as a negative effect. In the current version a touch is attributed to the shelf zone. Attribution to the exact product, using the store planogram and a reference photo, comes in the pilot.
- Heat maps ("Where do people stand, and where do their hands go?"): Three maps over the camera frame: where people stood, from the feet; where hands touched the shelf, from the wrists; and where heads pointed, an approximate attention that is never called eye tracking. By shelf level, the share of stops shows what eye level is earning.
- Gaps and restock ("Which fronts are empty right now?"): In frames with nobody in front, the shelf is split into cells and scored for texture and edges. A gap that lasts 3 s or more is flagged, with a restock alert when it appears right after a pick. Knowing whether it was the last unit needs the store's inventory system.

## Six steps, from the camera on the ceiling to the report.

1. Camera visit and zones. The team checks which installed cameras frame the shelf, reading RTSP streams or NVR exports. Zones are drawn over a reference frame: the aisle, the engagement area in front of the shelf and the shelf itself. When a camera does not frame the shelf, one extra camera per shelf solves it, and we say so before the pilot.
2. Detection and tracking. An open-source person detector, RF-DETR, runs on every analyzed frame. ByteTrack gives each person a temporary number that expires with the visit. No face recognition. The same person is never followed across days or stores.
3. Pose and anonymization. RTMPose finds wrists, shoulders and head keypoints. The head is covered by an irreversible blur before any clip leaves the store, and a second face detector checks every output frame as a safety net. Audio is removed from every output video.
4. Events frame by frame. Passed, stopped, touched and examined come from the feet, the walking speed and the wrists against the zones. Took is decided by comparing the shelf before and after the touch, in frames with nobody in front. Each event links to the second of anonymized video that proves it.
5. Test and control stores. The pilot design runs 8 weeks: setup, 2 weeks of baseline, 4 weeks with the action in the test stores, and 1 week of reading. Control stores run the same periods without the action. Every rate carries a 95% confidence interval.
6. Report. Funnel, effect against control, effect per product, conclusions and a recommendation, generated from the same numbers as the screens and exported to PDF. Pilot target: the report arrives within 72 hours of the end of measurement.

## What was measured, and the rules the numbers follow.

- 83 to 100%: agreement between the detector and a human count of people, 10 frames in each of 5 real CCTV videos
- 2 s: standing still in front of the shelf before a stop is counted
- 95%: confidence interval on every rate, from the funnel to the effect against control stores
- 4 + 2: test and control stores in the 8 week pilot design

## What it never does

- Never recognizes faces or identifies a person.
- Never estimates age, gender, emotion or ethnicity.
- Never follows the same person across days or across stores.
- Never links an image to an ID number, a loyalty card or a payment method, and never hands raw video to the manufacturer.
- Never calls estimated attention eye tracking, and never reports a product as taken without evidence.

## Who it is for

Trade marketing teams, Category management at retailers, Retail media teams, Shopper and insights teams at manufacturers, Research firms running in-store studies. For shelf decisions that get negotiated with numbers, where the number on the table today is sell-out alone.

## Questions and answers

### What is Shelf-Lab?

Shelf-Lab is a Cassi.ai product that measures shopper behavior in front of the shelf with the store's existing security cameras. It counts who passed, stopped, touched and took each product, blurs every head at the source, and gives every number its anonymized clip and a confidence interval. Cassi.ai is a software engineering company specialized in automation, and its products come from requests made by clients.

### How is it different from Shopper Insight Lab?

Shopper Insight Lab puts recruited participants in a virtual shelf, online, with a cart, a budget and missions, and observes the purchase before anything changes in the store. Shelf-Lab measures real shoppers at the real shelf, through the cameras already installed, and reads the effect of an action that is already in the store.

### Does it work with the cameras the store already has?

The engine reads RTSP streams and exports from the store's NVR. A technical visit confirms whether the cameras frame the shelf; when one does not, an extra camera per shelf solves it. In the pilot, processing can run on a local machine in the store or in a cloud region in Brazil.

### How is privacy handled?

Heads are blurred before any clip leaves the store, each person is a temporary number that expires with the visit, and audio is removed. In the proposal, raw video is kept for 7 days at most and anonymized clips for up to 90 days. The retailer is the data controller, Cassi.ai the processor, and the manufacturer receives aggregated numbers only. The legal basis proposed under Brazil's LGPD is legitimate interest, with an impact assessment done with the retailer's DPO.

### How accurate is it?

On 5 public CCTV videos, a reviewer counted people by hand in 10 frames per video and compared with the detector: agreement ranged from 83.3% to 100%. The engine misses small people behind shelves and people cut at the edge of the frame. In the pilot, accuracy is measured again against human checks and published in the report, and the success criterion is above 90% on passers and stops.

### Does it know whether the shopper took the product?

When there is evidence, yes: the shelf changes and stays changed, or the person examines the product in hand right after touching the shelf. Taking one pack from a stack of identical packs does not change the image, so that case appears as touched without visible change. In the pilot, matching with the POS receipt closes the count of who took it.

### Why not use the heat map the camera already has?

A camera heat map shows where people circulate. Shelf-Lab shows who stopped, who touched and who took, per shelf zone, and measures the effect of an action against control stores, with a confidence interval on each rate. Its attention map is an estimate from head direction and is never presented as eye tracking.

### What in the demo is real and what is simulated?

The vision engine is real: detection, tracking, pose, anonymization and events computed on 5 public CCTV videos, with human checks and stated limitations. The shelf study, period comparison and report screens use a deterministic simulation with fictional brands and stores, flagged on every screen. Privacy and pilot are a proposal. The animation on this page is illustrative.

## Bring the shelf action you need to prove.

Tell us the category, the action and the stores: an endcap, a new position, an in-store activation. We check whether your cameras frame the shelf and show how the study would run, with test and control stores. https://shelflab.cassiai.com/#acesso
