Kairos
Back to jobs

ML Operations Engineer

On-site
DeepIntentBelgrade, RS2 weeks agoWebsite
Aging
Engineering

Compensation

Salary undisclosed
Apply
Share

Description

What You'll Do:

DeepIntent is seeking an MLOps Engineer to join our European Data Infrastructure team. The MLOps Engineer will partner with our Data Science and AI teams to build and scale a high-performance machine learning and AI platform spanning our on-prem data centers and GPU resources - built on Python, Spark, Kubernetes, Docker, and modern orchestration frameworks (Argo Workflows, Airflow).

  • Partner with Data Science and AI Engineering teams to adopt MLOps best practices and migrate training/inference workloads onto the platform
  • Implement and maintain model tracking, versioning and experiment management (MLflow) with observability into model performance and drift
  • Build CI/CD pipelines purpose-built for ML/AI artifacts (model registries, container image pipelines, automated retraining triggers)
  • Continuously improve platform reliability, cost-efficiency and maintainability of the underlying codebase
  • Establish monitoring and observability for ML/AI systems (Prometheus, Grafana) covering GPU utilization, model latency, throughput and custom ML metrics
  • Design and operate ML/AI deployment infrastructure, including GPU cluster architecture, model serving and tool selection across training and inference workloads
  • Build and maintain infrastructure for LLM and generative AI workloads, including model fine-tuning pipelines, vector databases, RAG architectures and inference optimization
  • Collaborate with business units and Product on ML/AI-driven feature development
  • Manage individual project priorities, deadlines and deliverables in a timely manner

Who You Are:

  • Bachelor's degree in Computer Science or similar technical field of study, or equivalent practical experience
  • Strong software engineering skills in complex, distributed, multi-language systems (Python preferred)
  • Hands-on experience with Spark, Docker and Kubernetes in production environments
  • Experience building and operating end-to-end distributed systems
  • Experience developing and maintaining ML systems built with open-source MLOps tools (e.g., MLflow, Argo, Metaflow, Airflow, Kubeflow)
  • Strong understanding of software testing, benchmarking, and CI/CD practices
  • Solid understanding of Linux systems administration
  • Familiarity with LLM/generative AI tooling and concepts (model serving frameworks, embeddings, vector stores, RAG pipelines) is a strong plus
  • Familiarity with GPU infrastructure - CUDA fundamentals, GPU scheduling/orchestration in Kubernetes, and driver/toolkit management (NVIDIA drivers, CUDA toolkit, cuDNN)
  • Ability to work closely with data scientists and understand their tooling and workflows (Jupyter, notebooks, experiment tracking)
  • An enthusiastic learner with a genuine thirst for keeping up with the fast-moving ML/AI infrastructure landscape

Stack

LLMsGenerative AIPythonEmbeddingsData ScienceGPUCI/CDSparkAirflowMLOpsVector DatabasesDistributed SystemsMachine LearningFine-tuningKubernetesDockerCUDARAGMLflow
Posted
Aug 31, 2026
Last seen
Sep 2, 2026
First seen
Sep 2, 2026

Similar roles

Browse more AI jobs