March 21, 2024
This month we are closing our part of a publicly funded research project on AI cameras that collect data about active mobility: how many people walk and cycle past a given point, and in which direction. Our partner, a start-up building IoT equipment for road infrastructure, led it; we were responsible for the embedded AI part: recording and streaming on the camera systems, choosing and deploying detection models, and testing them in the field.
It was not our first contact with AI cameras: we built our first AI camera demonstrator for road infrastructure back in 2021, and throughout 2023 we evaluated whether a stereo camera with its own vision processor could become the basis of a counting pilot. Below are the lessons we took from it.
The goal was never video. A traffic planner needs counts per class, per direction and per quarter of an hour. If the camera computes them itself, no images of people have to leave the device. A sound design keeps images on the device, restricts remote access tightly and sends out only anonymous statistics. That makes the GDPR discussion much simpler, and it matters more now that the European Parliament approved the AI Act this month, with particular attention paid to systems that identify people. Counting that a cyclist passed is a very different thing from recognising who it was.
Roadside devices also have little power and often a weak mobile connection, so sending a few numbers instead of a video stream is not just a privacy choice.
On-device inference means the model runs on the camera's vision processor, so it first has to be converted into the format that processor understands. In practice this narrowed the choice far more than any benchmark table would suggest:
There was a second, less obvious constraint. Because detection runs inside the camera, the device is not designed to take a recorded video back as input, which is how you would normally compare models on a PC. We therefore built recording and streaming functions for the cameras and a way to run each model on the same recorded footage, so that every model could be compared on identical reference video.
Lab results say little about a street. We therefore tested outdoors on a pavement shared by pedestrians and cyclists, recorded every session and compared the camera's counts with a manual count of the same footage.
The lessons were mostly not about the model:
The next question is usually training. Models pre-trained on public datasets get you a long way; closing the gap in crowded scenes needs data from real mounting positions, and synthetic images are one way to reduce the cost of labelling it.
Edge object detection is not a smaller version of cloud AI. It is a different design, where hardware, mounting and power matter as much as the model, and where privacy can be built in rather than added later. If you are working on a device or a smart-city project that needs to see and count, take a look at our IoT services or get in touch.
Quote a project! Get advice.
Drop us a line or book a short intro call — we’ll get back to you with the right people on our side.
E-mail usinfo@embiq.comBook an intro call (opens in a new tab)Let’s investigate your project concept and its current status together.
Expect an initial project scope proposal, time and cost estimation from us.
The consultancy will be protected by the NDA.