Skip to content
All work

Computer vision platform

Real Vision

Cameras that understand what they see. Real-time detection, multi-camera identity tracking, licence-plate recognition and vehicle damage baselining — narrated in language, running entirely on the client's own hardware.

In production · on-premises and cloud GPU · 2026

  • ~347

    detection classes

  • 4

    model tiers, one GPU

  • 7

    safety event kinds

  • 0

    footage leaves site

01 The problem

  • CCTV records everything and tells you nothing. Finding an event means a person scrubbing footage.
  • Damage disputes come down to one person's word against another's, with no evidence trail.
  • Off-the-shelf detection models recognise a fixed list of objects, which never matches what a specific site actually cares about.

02 The approach

  • Open-vocabulary detection rather than a fixed class list, so the system can be pointed at what a site actually cares about without retraining.
  • A tiered model stack on one GPU: tracking every frame, plate OCR and re-identification only when triggered, a fast vision-language model for live narration, and a heavy one on demand for deep analysis.
  • Everything runs on the client's own machine. No footage leaves the premises.

03 Architecture

Tiered inference — one GPU, four tiers

Tier 0   every frame      YOLOWorld + ByteTrack            ~30–50 ms
Tier 1   triggered        plate OCR · re-ID · Florence-2   on detection
Tier 2   gated            live VLM narration               on scene change
Tier 3   on demand        heavy VLM deep-look · summary    operator request

04 What it does

Open-vocabulary detection

YOLOWorld in hybrid mode unions a general class list with a domain-specific one — roughly 347 classes, deduplicated and order-stable, with domain classes added first so they can never be dropped by the cap.

Multi-camera tracking & re-ID

ByteTrack lanes per camera with cross-view identity fusion, plus deep OSNet appearance embeddings that fall back to histogram matching when the GPU backend is unavailable — so it degrades instead of breaking.

Plate recognition

ONNX-based ANPR with region-aware cleanup that conservatively corrects O/0 and I/1 confusions against known regional plate layouts, and tags each read with its region.

Vehicle condition baselining

Entry damage baseline per vehicle, canonicalised diffing on exit, and evidence photos captured automatically — so a dispute has a record.

Safety events

A seven-kind alarm pipeline with arbitration between the detector and the vision-language model, plus owner notification over Telegram, webhook or email with the evidence photo attached.

Learns without forgetting

A self-correcting label memory that survives restarts, and a continuous training loop behind a quality gate so corrections improve the model instead of drifting it.

Built with

  • Python 3.12
  • Flask
  • YOLOWorld
  • ByteTrack
  • OpenCV
  • Ollama VLM
  • Florence-2
  • OSNet re-ID
  • fast-alpr
  • SQLite WAL
  • CUDA
  • Docker

Deployed for automotive service businesses in the Gulf. Client details are confidential and are not published here.