Home

ML

VisionSafe: Real-Time Driver Alertness Monitoring

GitHub ↗

Webcam frame

OpenCVany standard webcam

OpenCV

Find the face

MediaPipe Face Mesh468 landmarks

Locates the eyes, nose, chin and mouth corners

Eyes

Open or closed?

TensorFlow
  1. Crop each eye

    64 × 64 px

    Cut out around the eye-corner landmarks

  2. Classify

    MobileNetV2fine-tuned

    Trained on the MRL Eye dataset, tested on 16,981 images

Out: each eye labelled open or closed

Head

Looking at the road?

OpenCV
  1. Pick six landmarks

    nosechineyesmouth

    Matched to a generic 3D face model

  2. Solve the pose

    cv2.solvePnP

    Works out how the head is rotated from where the points land

Out: pitch, yaw and roll in degrees

Status and risk score

SafeDrowsyDistracted0 to 100

Both eyes closed is drowsy; head tilted past 20° is distracted. The score rises on bad frames and falls on good ones.

Alert

score above 703 s cooldown

A different sound for drowsy and for distracted

What happens to every frame, from capture to alert.

The problem

Tired and distracted drivers are a major cause of road accidents, and the systems that watch for it usually rely on cameras or sensors built into the car. This project asks whether an ordinary webcam and a laptop can do the same job: notice when someone's eyes stay closed or their head turns away from the road, and warn them while there is still time to react.

How I approached it

  1. 01

    Data. Trained on the MRL Eye dataset of open and closed eyes, resized to 64 by 64 pixels and split into training, validation and test sets, with random flips and other augmentation on the training images. Used driver photos from the NTHU-DDD2 dataset to test the head-pose logic on real drowsy and alert faces.

  2. 02

    Model. Fine-tuned MobileNetV2 in TensorFlow with CUDA GPU acceleration, starting from ImageNet weights and adding a small head (a 64-unit dense layer, 30% dropout and a sigmoid output). Trained for 20 epochs with Adam at a learning rate of 1e-5 and early stopping. MobileNetV2 was chosen because it is small enough to run on every frame.

  3. 03

    Landmarks. MediaPipe Face Mesh locates 468 facial landmarks in each frame. The eye-corner landmarks define the two crops that go to the classifier, so the model only ever sees eyes.

  4. 04

    Head pose. Six landmarks (nose tip, chin, outer eye corners and mouth corners) are matched to a generic 3D face model with OpenCV's solvePnP, which returns the head's pitch, yaw and roll. I ran this over 3,000 NTHU frames to see how head angle separates drowsy from alert drivers before choosing thresholds.

  5. 05

    Decide. Simple rules turn the readings into a status: both eyes closed is drowsy, a head tilted more than 20 degrees is distracted, anything else is safe. Each frame then moves a risk score between 0 and 100: up 2 for drowsy, up 1 for distracted, down 1 for safe.

  6. 06

    Alert. Once the score passes 70 an alarm plays, with a different sound for drowsy and for distracted and a 3-second cooldown so it does not nag.

  7. 07

    Deploy. Runs locally as an OpenCV app with the readings drawn over the video. For the web version, the browser captures webcam frames and posts them to a FastAPI service, which returns eye state, head angles, status and score.

The hard part

A model that is right 92% of the time is still wrong on hundreds of frames an hour, and every blink is a genuinely closed eye. I had tuned the classifier toward caution, so it catches 99% of closed eyes but calls an open eye closed about 15% of the time. Alerting on single frames would have been unusable. The fix was to never react to one frame: each frame only nudges the risk score, and the alarm waits for the score to pass 70. A blink or a misread moves it a little and fades; eyes that stay shut push it over within moments.

Results

  • It almost never misses closed eyes

    On 16,981 eye images the model had never seen, it was 92% accurate overall. The errors are deliberately one-sided: missing a drowsy driver costs far more than a false alarm.

    Eye-state model on unseen test images
    Closed eyes caught99%
    Open eyes recognised85%
  • False alarms are absorbed, not sounded

    Because the alarm follows a running score, not a single frame, blinks and one-off misreads pass without a sound, while sustained eye closure or looking away triggers it.

  • It tells drowsy from distracted

    Eyes and head pose are judged separately, so the driver hears a different alert for nodding off than for looking away, and knows what to correct.

  • No special hardware

    Everything runs from a standard webcam on an ordinary laptop, with no sensors or dedicated cameras.

Future improvements

The web version's frame rate and response time have room to improve; running the model in the browser, without a round trip to a server for every frame, is the obvious route. Distraction is currently judged on head tilt alone, so adding side glances (yaw) and yawning detection are the next steps.