March 21, 2024

Edge AI for mobility: what we learned counting pedestrians and cyclists on the camera

This month we are closing our part of a publicly funded research project on AI cameras that collect data about active mobility: how many people walk and cycle past a given point, and in which direction. Our partner, a start-up building IoT equipment for road infrastructure, led it; we were responsible for the embedded AI part: recording and streaming on the camera systems, choosing and deploying detection models, and testing them in the field.

It was not our first contact with AI cameras: we built our first AI camera demonstrator for road infrastructure back in 2021, and throughout 2023 we evaluated whether a stereo camera with its own vision processor could become the basis of a counting pilot. Below are the lessons we took from it.

Why count on the camera

The goal was never video. A traffic planner needs counts per class, per direction and per quarter of an hour. If the camera computes them itself, no images of people have to leave the device. A sound design keeps images on the device, restricts remote access tightly and sends out only anonymous statistics. That makes the GDPR discussion much simpler, and it matters more now that the European Parliament approved the AI Act this month, with particular attention paid to systems that identify people. Counting that a cyclist passed is a very different thing from recognising who it was.

Roadside devices also have little power and often a weak mobile connection, so sending a few numbers instead of a video stream is not just a privacy choice.

The camera decides which models you can use

On-device inference means the model runs on the camera's vision processor, so it first has to be converted into the format that processor understands. In practice this narrowed the choice far more than any benchmark table would suggest:

  • Lightweight SSD-type detectors ran smoothly on the device and covered both people and vehicles.
  • Compact YOLO variants recognised more classes, but at noticeably lower frame rates.
  • Specialised models, for example vehicle-only ones, can be excellent for one class and of little use once the requirement changes.
  • The newest architectures are not always available in a format the vision processor supports; converting them can take more effort than the rest of the pipeline.

There was a second, less obvious constraint. Because detection runs inside the camera, the device is not designed to take a recorded video back as input, which is how you would normally compare models on a PC. We therefore built recording and streaming functions for the cameras and a way to run each model on the same recorded footage, so that every model could be compared on identical reference video.

Field tests: every result against a manual count

Lab results say little about a street. We therefore tested outdoors on a pavement shared by pedestrians and cyclists, recorded every session and compared the camera's counts with a manual count of the same footage.

  • In light traffic, a lightweight detector combined with depth information and a simple movement filter came close to the manual reference.
  • In busy periods, with groups passing at once, every model struggled: people were merged or missed and bicycles over-counted.
  • Heavier models were not automatically better once they had to run on the device.

The lessons were mostly not about the model:

  • Thresholds are a trade-off, not a fix. Lowering the detection threshold made the model see more, but false detections rose dramatically. Filtering by movement and depth helped more than tuning the threshold.
  • Mounting matters as much as the model. Counts were best where paths did not cross and people did not linger at the edge of the field of view. Installation height also changes results, especially in crowds.
  • Power shapes everything. Low-power requirements narrow the hardware choice and make every on-site test day more expensive, so plan for them from the start.

What we would tell anyone planning camera-based counting

  • Define the classes, directions and acceptable error before choosing hardware. A daily statistic and a safety trigger need very different accuracy.
  • Choose the edge hardware knowing that it limits which models you can run, and check the conversion path for the models you want before you buy.
  • Plan for reference data from day one: recorded video and a manual count for every test, in quiet and busy conditions alike.
  • Test in crowds and at the real mounting position. A quiet pavement makes every model look good.
  • Keep images on the device and send only the numbers the use case needs.

The next question is usually training. Models pre-trained on public datasets get you a long way; closing the gap in crowded scenes needs data from real mounting positions, and synthetic images are one way to reduce the cost of labelling it.

Edge object detection is not a smaller version of cloud AI. It is a different design, where hardware, mounting and power matter as much as the model, and where privacy can be built in rather than added later. If you are working on a device or a smart-city project that needs to see and count, take a look at our IoT services or get in touch.

Quote a project! Get advice.

Let’s talk about your project

Drop us a line or book a short intro call — we’ll get back to you with the right people on our side.

Book an intro call (opens in a new tab)
Call us+48 81 561 85 01
LublinEMBIQ Sp. z o.o.al. Kraśnicka 2720-718 Lublin, Poland
GrazEMBIQ GmbHBrückenkopfgasse 1/68020 Graz, Austria
Let’s inve

Let’s investigate your project concept and its current status together.

Expect an

Expect an initial project scope proposal, time and cost estimation from us.

The consul

The consultancy will be protected by the NDA.