Our client is a technology company delivering advanced software engineering and quality engineering services for complex, real-world systems. Their work combines backend, embedded, cloud, on-premise, and edge technologies, with a strong focus on video, security, IoT, and mission-critical applications.
The company works on products where reliability, performance, and operational stability matter — including video recording and management platforms, alarm and security systems, video analytics, and connected device ecosystems. This creates an opportunity to join a technically ambitious project that goes beyond standard web development and involves distributed systems, media processing, edge deployments, and production-grade engineering.
About the Role
Our client is looking for an AI / ML Platform Engineer to take ownership of the on-device computer vision and edge inference pipeline for an integrated video monitoring platform.
The product already includes a working analytics pipeline for use cases such as object and person detection, license plate recognition, and scene description using vision-language models. The next stage is to tune, optimize, productionize, and evolve this pipeline so it can reliably run on edge devices using Intel Core Ultra NPU/GPU/CPU through OpenVINO.
This is a hands-on engineering role focused on applied machine learning in production. The selected person will work at the intersection of computer vision, model optimization, edge hardware, video analytics, Linux, cloud fallback architecture, and event integration with downstream alarm and monitoring systems.
Key Responsibilities
- Own and evolve the edge analytics pipeline for video monitoring use cases.
- Build and optimize flows for RTSP ingest, motion gating, object/person detection, crop extraction, scene description, and event generation.
- Tune vision-language model prompts and frame-composition strategies, including multi-cell image composites and structured-output grammars.
- Optimize and benchmark model execution on Intel NPU/GPU/CPU using OpenVINO.
- Work on prompt-length tuning, chunked prefill, generate-hint configuration, device targeting, and runtime performance improvements.
- Migrate and evaluate models as new open-weights releases become available, including next-generation VLMs, YOLO models, segmentation models, and ReID approaches.
- Build and own a model deployment pipeline from Hugging Face models through OpenVINO IR conversion, NPU export, and fleet rollout via OS image builds.
- Integrate AI detection outputs with the wider event pipeline, including event store, event bridge, and downstream alarm or monitoring systems.
- Evaluate third-party detection providers and cloud inference options for selected high-value event types.
- Design edge-to-cloud fallback strategies for scenarios where local inference is not sufficient or where higher-confidence analysis is required.
- Debug model performance, quality degradation, runtime failures, driver issues, and hardware-specific inference problems.
- Collaborate with backend, full-stack, QA, and embedded/edge engineers to make the analytics pipeline reliable, observable, testable, and production-ready.
Requirements
- 4+ years of experience in applied machine learning, computer vision, or ML engineering in production environments.
- Strong Python skills and practical experience with PyTorch and ONNX.
- Hands-on experience deploying computer vision models outside of notebooks and into real production systems.
- Experience with OpenVINO toolkit, including model conversion, runtime API, and device targeting across NPU, GPU, CPU, or heterogeneous execution.
- Practical experience with YOLO-family models or comparable object detection model deployment.
- Understanding of video analytics pipelines, including frame extraction, detection, cropping, batching, latency, throughput, and inference trade-offs.
- Experience with cloud GPU inference or ML-serving infrastructure, such as AWS SageMaker, Bedrock, EC2 GPU instances, Hugging Face Inference Endpoints, or similar platforms.
- Ability to design edge-to-cloud fallback flows for selected high-value events.
- Comfortable working with Linux environments and debugging GPU/NPU driver-level issues when models fail, slow down, or silently degrade.
- Ability to read and understand lower-level inference-related code, including C/C++ code, when needed to diagnose OpenVINO or runtime issues.
- Strong engineering mindset: versioning, benchmarking, deployment pipelines, observability, rollback, reproducibility, and production reliability.
- Ability to communicate trade-offs between model quality, latency, cost, hardware limitations, and product requirements.
Nice to Have
- Hands-on experience with Intel NPU, especially Intel Core Ultra series.
- Familiarity with NPU driver versions, hardware-specific limitations, and known runtime issues.
- Experience running vision-language models in production, such as Qwen-VL, LLaVA, Florence-2, or similar models.
- Experience with real-time video analytics across multiple concurrent camera streams on a single edge device.
- Experience with multi-camera coordination, cross-camera tracking, ReID, or event correlation.
- Experience tuning models under forensic or operational requirements, where false positives and false negatives have different business or safety costs.
- Background in video surveillance, physical security, access control, alarm systems, smart buildings, or IoT.
- Experience with ONVIF, RTSP, video management systems, license plate recognition, or security analytics.
- Experience with model quantization, pruning, batching, caching, memory optimization, or hardware-specific acceleration.
- Familiarity with OS image builds, fleet rollout, edge device management, or over-the-air model deployment.