Matcha
Joblogic logo

MLOps Engineer

Joblogic

Apply

Cloud field service management software streamlining scheduling, jobs, and efficiency.

$209.4M raised· Backed by Vista Equity Partners, Axiom Equity, LGPS Central
FULL TIMERemote · Asia-Pacific474 employeesPosted Sep 17
linuxpythonazureawsci/cdansibledockerdatabricksopentelemetrymlopsllmopsinfrastructure as codeterraformkubernetessystem administration

Remote startup roles in your inbox

Matcha reads all job descriptions to surface the handful that actually matter in a zero noise email. Simple by design.

Describe your next role, cut the noise

The Joblogic Story

Established in 1998, Joblogic is the UK’s #1 Field Service Management (FSM) software platform. We are a global business with offices in the UK, Pakistan, and Vietnam. Since our management buy-out in 2013, we have grown from ~£500K ARR to ~£35M+ ARR and expanded our team from 11 to 500+ people.

Recently, we secured a strategic growth investment from Vista Equity Partners — a global technology investor specialising in enterprise software. This investment includes over £100 million in new primary capital and will fuel our next phase of growth by accelerating our AI-first roadmap, expanding our platform into CAFM (Computer-Aided Facilities Management) capabilities, and supporting our expansion across Europe and beyond.

With Vista’s backing, we’re transforming from a successful UK business into a global scaling SaaS rocket ship, and we’d love for you to join us on our journey to £100M ARR across international markets.

Joblogic provides software to service contractors who install and maintain the built environment. Our platform helps businesses streamline operations, improve profitability, ensure compliance, and achieve rapid growth. With over 100,000 users across industries including HVAC, plumbing, electrical maintenance, facilities management, and building fabric maintenance, we are entering a new era of intelligent automation, predictive maintenance, and data-driven decision-making for service firms.

About the Role 

We are building Joblogic’s AI Agent Platform — a multi-tenant system for designing, versioning, evaluating, and running AI agents that handle customer conversations over email, voice, SMS, and WhatsApp on behalf of our customers. It runs on a Python/FastAPI service estate (API, Celery workers, knowledge-base worker, voice gateway, service-bus listener) and React portals, hosted on Azure Container Apps, a sizeable Linux VM fleet across Azure and AWS, Databricks, and a set of vendor platforms. 

We are looking for a DevOps Engineer to own how all of it is deployed, kept running, secured, and paid for. You will be the go-to engineer for anything that runs, deploys, or costs money across Azure, AWS, and Databricks, and you will build the release and observability machinery that lets the team ship agent versions and machine learning models safely and roll them back in one step when they misbehave. 

To be clear about the bar: deep strength in Linux and cloud infrastructure is what we are hiring for and is not negotiable. The MLOps and LLMOps side matters just as much to the role, but if your experience there is lighter than your infrastructure experience, we would still like to hear from you — we will invest in building it. 

You will work closely with the AI, data, and product engineering teams, and your work will directly determine how reliably tens of thousands of field-service businesses experience intelligent automation. 

What You’ll Do 

  • Run the Linux fleet — provision, harden, patch, configure, back up, and capacity-plan the Linux VM estate across Azure and AWS, and respond to incidents on it. 
  • Own CI/CD — build and maintain the pipelines for a multi-service Python and React estate: build, test, database-migration orchestration, environment promotion, and rollback. 
  • Gate and promote agent versions — run the draft → approved → deployed pipeline, wire online evaluators in as release gates, and keep one-click rollback through deployment revision tracking. 
  • Operate the multi-provider LLM path — rate limiting, token streaming, per-tenant token budgets, caching, circuit breakers and fallbacks across model providers, with cost and latency metrics flowing into monitoring. 
  • Build the training-to-serving path — stand up GPU training jobs, a model registry, approval-gated deployment jobs, and model serving for the team’s machine learning and computer-vision models. 
  • Administer Databricks — workspaces, cluster and usage policies, serverless compute, Unity Catalog, service principals, deployment bundles, and cost attribution. 
  • Build observability — instrument services with OpenTelemetry into Application Insights, run agents or collectors on VMs, and define alerting, SLOs, and error budgets for the agent runtime and voice gateway. 
  • Unify the control plane — bring Azure and AWS hosts under common management for patching, inventory, and compliance baselines enforced through configuration management. 
  • Own identity and secrets — managed identities, workload identity federation for pipelines, secret rotation, least-privilege access, and audit trails. 
  • Own continuity — design backup and restore paths for the data and services the platform depends on, test them on a schedule, and keep documented recovery objectives rather than assumptions. 
  • Own reliability — lead incident response and blameless reviews that end in concrete, tracked follow-ups. 

Essential Experience and Skills 

  • 3+ years running production Linux fleets (Ubuntu, RHEL, or equivalent): systemd, storage, networking and firewalls, SSH and key management, TLS and reverse proxies, backups and restores. This is non-negotiable. 
  • Configuration management with Ansible (or equivalent) and patching at fleet scale, plus hardening against recognised benchmarks such as CIS. 
  • Fluent Bash and working Python; comfortable debugging from logs, metrics, and system tools under incident pressure. 
  • Strong Azure experience: container hosting, container registry, Key Vault, service bus, managed PostgreSQL and Redis, virtual networks and private endpoints, VMs and scale sets, managed identities and RBAC. 
  • Infrastructure as code (Bicep, Terraform, or equivalent) and pipeline-as-code with federated, secretless service connections. 
  • Working AWS experience: EC2 fleet management, IAM, VPC, S3, CloudWatch, and managed model access with quotas and guardrails. 
  • Containers and release engineering: multi-service Docker builds, zero-downtime deploys, database-migration orchestration, environment parity, and a tested rollback path. You have built and run a CI/CD pipeline for a team, including the branching model and merge gating. 
  • Observability in practice: OpenTelemetry instrumentation, a metrics and logs backend such as Application Insights or Grafana, and alert design that pages on symptoms rather than causes. 
  • Defining SLIs, SLOs, and error budgets, and running incident response and blameless reviews. 
  • MLOps fundamentals: experiment tracking and a model registry (MLflow or equivalent), orchestrating GPU training, and serving models in production. 
  • LLMOps fundamentals: tracing and online evaluation in production, and LLM serving and routing concerns such as rate limiting, streaming, load balancing, per-tenant budgets, caching, fallbacks, and provider key management. 
  • Security practice: secretless authentication, least privilege, blast-radius thinking, auditability, and dependency and supply-chain hygiene. 
  • Committed to continuous learning, proactive problem-solving, and timely issue identification, with a keen interest in staying current with a fast-moving field. 
  • Strong communicator, experienced in collaborating with cross-functional teams using tools such as Jira and Slack. 
  • Creative and innovative thinker, consistently contributing fresh ideas and solutions in alignment with current technological trends. 

Nice to Have 

  • Databricks administration: workspaces, usage policies, Unity Catalog, service principals, deployment bundles, and the Databricks Terraform provider. 
  • Kubernetes and Helm; GitHub Actions. 
  • Self-hosted LLM serving (vLLM, SGLang, TensorRT-LLM) and inference servers such as ONNX Runtime, Triton, KServe, or Ray Serve. 
  • Prometheus and Grafana; load testing of token-streaming endpoints with k6 or Locust. 
  • FinOps: cost allocation and budgets across Azure, AWS, Databricks, and model providers. 
  • Other LLM observability stacks (LangSmith, Langfuse, Braintrust, Arize Phoenix). 
  • Operating voice and telephony infrastructure (Twilio, ElevenLabs, Azure Speech). 
  • Awareness of the OWASP Top 10 for LLM Applications and prompt-injection-aware infrastructure such as guardrail services on retrieval inputs. 
  • Agent sandboxing (Firecracker, gVisor); SOC 2 or ISO 27001 control experience. 
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent hands-on experience. 

What We Offer 

  • Professional Working environment 
  • Market Competitive Salary 
  • Life Insurance & Medical Insurance (Including Family) 
  • OPD 
  • Provident Fund 
  • Gym Facility 
  • Maximum 45 Weekly Hours (Monday–Friday) 
  • Remote Working (During Pandemic Situation) 
  • Company trip 
  • 29 Annual Leaves 
  • 8 Sick & uncapped Compassionate Leaves (As per Company Policy) 
  • Have a chance to work onsite with the UK team 

Similar remote jobs