Urban Vehicle Detection AI Vision
Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
I need an AI-powered vision solution that can reliably detect vehicles as they pass through busy urban streets. The goal is straightforward: feed a live video stream (or recorded footage) into a model and receive accurate, real-time bounding boxes and class labels for every car, bus, truck, or motorcycle that enters the frame, day or night.
Here is what matters most to me:
• High detection accuracy in crowded, visually noisy city scenes—including partial occlusions, varied lighting, and different camera angles.
• Real-time or near real-time processing (30 fps on an edge device is ideal, but I’m open to server-side inference if latency stays low).
• Easy integration: a clean Python API or REST endpoint that I can wire into an existing dashboard. I already work with OpenCV, so keeping compatibility there is a plus.
• Clear instructions on how to retrain or fine-tune the model when traffic patterns, camera positions, or city regulations change.
Preferred stack: PyTorch or TensorFlow/Keras, YOLOv5/v8 or similar single-stage detector, and standard preprocessing tools such as OpenCV, scikit-image, or Pillow. If you have a different approach that beats these on speed and accuracy, feel free to propose it.
Deliverables
1. Trained model weights and all training scripts.
2. Inference script or microservice (Dockerised) that accepts an RTSP/HTTP stream and outputs JSON or MQTT messages with coordinates, class, and confidence.
3. Short README with setup steps, hardware requirements, and performance benchmarks on sample footage I will supply (urban daytime and nighttime clips).
4. Optional but welcomed: a simple web page that overlays detections on the video for quick validation.
I will provide labelled sample footage to kick-start training and can annotate additional clips if you supply an efficient annotation guideline. Let’s discuss timelines and any hardware constraints you need to meet the real-time target.
Here is what matters most to me:
• High detection accuracy in crowded, visually noisy city scenes—including partial occlusions, varied lighting, and different camera angles.
• Real-time or near real-time processing (30 fps on an edge device is ideal, but I’m open to server-side inference if latency stays low).
• Easy integration: a clean Python API or REST endpoint that I can wire into an existing dashboard. I already work with OpenCV, so keeping compatibility there is a plus.
• Clear instructions on how to retrain or fine-tune the model when traffic patterns, camera positions, or city regulations change.
Preferred stack: PyTorch or TensorFlow/Keras, YOLOv5/v8 or similar single-stage detector, and standard preprocessing tools such as OpenCV, scikit-image, or Pillow. If you have a different approach that beats these on speed and accuracy, feel free to propose it.
Deliverables
1. Trained model weights and all training scripts.
2. Inference script or microservice (Dockerised) that accepts an RTSP/HTTP stream and outputs JSON or MQTT messages with coordinates, class, and confidence.
3. Short README with setup steps, hardware requirements, and performance benchmarks on sample footage I will supply (urban daytime and nighttime clips).
4. Optional but welcomed: a simple web page that overlays detections on the video for quick validation.
I will provide labelled sample footage to kick-start training and can annotate additional clips if you supply an efficient annotation guideline. Let’s discuss timelines and any hardware constraints you need to meet the real-time target.
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.