ML
VisionSafe: Real-Time Driver Alertness Monitoring
Webcam frame
OpenCVany standard webcam
Find the face
MediaPipe Face Mesh468 landmarks
Locates the eyes, nose, chin and mouth corners
Eyes
Open or closed?
Crop each eye
64 × 64 px
Cut out around the eye-corner landmarks
Classify
MobileNetV2fine-tuned
Trained on the MRL Eye dataset, tested on 16,981 images
Out: each eye labelled open or closed
Head
Looking at the road?
Pick six landmarks
nosechineyesmouth
Matched to a generic 3D face model
Solve the pose
cv2.solvePnP
Works out how the head is rotated from where the points land
Out: pitch, yaw and roll in degrees
Status and risk score
SafeDrowsyDistracted0 to 100
Both eyes closed is drowsy; head tilted past 20° is distracted. The score rises on bad frames and falls on good ones.
Alert
score above 703 s cooldown
A different sound for drowsy and for distracted
The problem
Tired and distracted drivers are a major cause of road accidents, and the systems that watch for it usually rely on cameras or sensors built into the car. This project asks whether an ordinary webcam and a laptop can do the same job: notice when someone's eyes stay closed or their head turns away from the road, and warn them while there is still time to react.
How I approached it
- 01
Data. Trained on the MRL Eye dataset of open and closed eyes, resized to 64 by 64 pixels and split into training, validation and test sets, with random flips and other augmentation on the training images. Used driver photos from the NTHU-DDD2 dataset to test the head-pose logic on real drowsy and alert faces.
- 02
Model. Fine-tuned MobileNetV2 in TensorFlow with CUDA GPU acceleration, starting from ImageNet weights and adding a small head (a 64-unit dense layer, 30% dropout and a sigmoid output). Trained for 20 epochs with Adam at a learning rate of 1e-5 and early stopping. MobileNetV2 was chosen because it is small enough to run on every frame.
- 03
Landmarks. MediaPipe Face Mesh locates 468 facial landmarks in each frame. The eye-corner landmarks define the two crops that go to the classifier, so the model only ever sees eyes.
- 04
Head pose. Six landmarks (nose tip, chin, outer eye corners and mouth corners) are matched to a generic 3D face model with OpenCV's solvePnP, which returns the head's pitch, yaw and roll. I ran this over 3,000 NTHU frames to see how head angle separates drowsy from alert drivers before choosing thresholds.
- 05
Decide. Simple rules turn the readings into a status: both eyes closed is drowsy, a head tilted more than 20 degrees is distracted, anything else is safe. Each frame then moves a risk score between 0 and 100: up 2 for drowsy, up 1 for distracted, down 1 for safe.
- 06
Alert. Once the score passes 70 an alarm plays, with a different sound for drowsy and for distracted and a 3-second cooldown so it does not nag.
- 07
Deploy. Runs locally as an OpenCV app with the readings drawn over the video. For the web version, the browser captures webcam frames and posts them to a FastAPI service, which returns eye state, head angles, status and score.
The hard part
A model that is right 92% of the time is still wrong on hundreds of frames an hour, and every blink is a genuinely closed eye. I had tuned the classifier toward caution, so it catches 99% of closed eyes but calls an open eye closed about 15% of the time. Alerting on single frames would have been unusable. The fix was to never react to one frame: each frame only nudges the risk score, and the alarm waits for the score to pass 70. A blink or a misread moves it a little and fades; eyes that stay shut push it over within moments.
Results
It almost never misses closed eyes
On 16,981 eye images the model had never seen, it was 92% accurate overall. The errors are deliberately one-sided: missing a drowsy driver costs far more than a false alarm.
Eye-state model on unseen test images Closed eyes caught99%Open eyes recognised85%False alarms are absorbed, not sounded
Because the alarm follows a running score, not a single frame, blinks and one-off misreads pass without a sound, while sustained eye closure or looking away triggers it.
It tells drowsy from distracted
Eyes and head pose are judged separately, so the driver hears a different alert for nodding off than for looking away, and knows what to correct.
No special hardware
Everything runs from a standard webcam on an ordinary laptop, with no sensors or dedicated cameras.
Future improvements
The web version's frame rate and response time have room to improve; running the model in the browser, without a round trip to a server for every frame, is the obvious route. Distraction is currently judged on head tilt alone, so adding side glances (yaw) and yawning detection are the next steps.
- Python
- TensorFlow
- CUDA
- MobileNetV2
- OpenCV
- MediaPipe
- FastAPI
- Docker