I am building a real-time vehicle recognition system using Python.
The system should support these input methods:
Local video file upload
Live camera feed using a webcam
camera URL
YouTube video URL
Additionally, I want to implement:
OpenCV point/ROI selection (mouse click points on frame)
Zone detection (detect vehicles only inside selected polygon/zones)
A modular architecture so each input method follows the same processing pipeline.
I am able to load videos individually, but I am struggling to design a clean structure that handles all input types consistently. I also need guidance on how to integrate OpenCV cv2.setMouseCallback() for ROI selection before starting the detection loop.
Questions:
What is the best way to structure a common video input pipeline for:
file path
webcam
RTSP URL
YouTube URL (via pytube or yt-dlp)?
How do I properly implement OpenCV point selection (mouse clicks) on the first frame, and pass those zone coordinates to my detection loop?
Are there standard patterns for combining ROI selection + detection + frame-by-frame processing in real-time systems?
Any example architecture or code pattern that cleanly separates:
input handler
processing/detection (YOLO/OpenCV)
output/logging