
AI Engineer (Agentic AI)
- AI
- Language Models
- CUDA
- TensorRT
- Microservices
- Computer Vision
- Gemini
- GPT
- Python
- gRPC
- Docker
- Kubernetes
- AWS
- GCP
- Azure
We are looking for anAI Engineer with strong experience in modern vision and multimodal models, especially on theNvidia ecosystem, who can also build small services around these models to prove and validate product flows.
The ideal candidate is someone who:
Understandsvision-language models (VLMs) andworld foundation models
Knows how to work withcameras and camera control (PTZ, streams, RTSP, frame handling, etc.)
Can wrap these capabilities into asimple, robust service to demonstrate end-to-end AI flows in real use cases.
This role is veryhands-on and experimental, focused on prototyping, validating ideas quickly, and then hardening the most promising ones.
Key Responsibilities
Explore, evaluate, and integratevision-language models (VLMs) andworld foundation models for understanding real-world scenes from cameras.
Design and run experiments forscene understanding, spatial reasoning, and multimodal interaction (image/video + text).
Optimize and deploy models on theNvidia stack (CUDA, TensorRT, DeepStream, or related Nvidia SDKs and tools).
Work withcamera streams (e.g. PTZ, RTSP/IP cameras): frame capture, streaming, basic processing, and latency/performance tuning.
Prototype AI-driven flows that connect camera input β model inference β actionable output or API.
Buildsmall backend services or microservices around AI models to test and demonstrate real product flows.
Expose models through APIs or internal tools so they can be consumed by other teams.
Write clean, maintainable code with basic testing, monitoring, and logging.
Collaborate with other engineers and product stakeholders to move from prototype β proof of concept β production-ready solution.
Must-Have Qualifications
Strong hands-on experience withcomputer vision or multimodal AI, especially:
Vision-language models (e.g. Gemini, GPT-4V, LLaVA, QWIN, or similar)
Scene understanding, detection, or tracking from camera feeds
Practical experience withNvidia software & tools, such as:
CUDA, TensorRT, DeepStream, Nvidia SDKs / frameworks, or similar
Experience working withcamera systems:
RTSP or IP cameras, video streams, frame processing, camera configuration/control
Solid programming skills in at least one of:
Python
Ability todesign and implement a small service around a model:
REST/gRPC APIs, background workers, or simple pipelines
Comfortable working in a fast-paced environment with rapid prototyping and iteration.
Nice-to-Have
Experience withworld models or world foundation models for spatial/scene reasoning, like COSMOS.
Experience deploying models onedge devices or GPU-based systems.
Familiarity with containers and infrastructure:Docker, Kubernetes, and a major cloud provider (AWS/GCP/Azure).
Background in robotics, autonomous systems, or real-time perception.
AI Engineer (Agentic AI) Β· Strattmont