Matcha
Buildkite logo

Senior Engineer - Platform

Buildkite

See similar roles
Apply

CI platform for teams; trusted by top global organizations.

$20.3M Series A, $21.0M Series B raised· Backed by OpenView, OneVentures, Airtree Ventures
FULL TIMERemote · United StatesRemote · Asia-Pacific157 employeesPosted Sep 30
kubernetesawsterraforminfrastructure as codeci/cddistributed systemsobservabilityincident responseautomationnetworkingcapacity managementsystem designreliability engineeringplatform engineeringmentoring

Remote startup roles in your inbox

Matcha reads all job descriptions to surface the handful that actually matter in a zero noise email. Simple by design.

Describe your next role, cut the noise

About Buildkite

At Buildkite, our mission is to unblock every developer on the planet. We’ve rethought how software delivery should work and built a platform that is fast, reliable, secure, and able to scale to the needs of demanding high-growth technology companies globally, including Airbnb, Shopify, Canva, PagerDuty, Lyft, and Pinterest.

Job Overview

We’re hiring a Senior Platform Engineer to join our Platform engineering organisation. The Platform team builds and operates the shared foundations that help Buildkite teams ship and scale products securely and reliably.

This is a hands-on role for a senior platform engineer with strong cloud and infrastructure foundations. You’ll design and operate reliable Kubernetes platforms across multiple clusters, solve distributed-systems challenges, automate infrastructure at scale, and take ownership of the systems you build. 

We’re also hiring a Platform Engineer at intermediate level. If you’re earlier in your career than this role describes but the work below appeals, apply anyway and we’ll consider you.

🔧 About the Team

Our mission is to build the foundations that let Buildkite teams ship, operate, and scale products securely and reliably. We make the reliable path the easiest path through paved roads for deployment, observability and security.

Buildkite’s customers run some of the world’s most demanding CI/CD and emerging AI workloads. As software development accelerates, CI/CD is becoming a major scaling bottleneck. A shared-platform change can affect every product team, while one unusual workload can expose the platform’s next limit. The infrastructure underneath must behave accordingly.

Our current challenges include:

  • Self-service, repeatable foundations. Build secure defaults, reusable architecture, automation, tooling, and documentation so teams can provision and operate consistent environments across regions and isolated use cases without tickets or operational hand-offs.
  • Kubernetes at production scale. Improve networking, upgrades, patching, capacity management, and recovery without disrupting teams or customers.
  • Global and regional resilience. Develop repeatable regional environments, customer routing, disaster recovery, and tested failover. Design for traffic growth, partial failures, hot shards, noisy neighbours, queue pressure, and datastore limits.
  • Safe, observable operations. Make infrastructure and application changes easier to deploy, understand, and reverse. Improve observability, capacity planning, incident response, documentation, and runbooks, and demonstrate readiness through tests, dashboards, and recovery exercises.

We measure our impact through cluster availability, customer-routing latency, deployment speed and success, self-service adoption, and recovery time.

🚀 What You’ll Do

  • Own substantial parts of Buildkite’s AWS and multi-cluster Kubernetes platform from design through production operation.
  • Design reusable, self-service automation for provisioning and operating secure, consistent, and isolated platform environments.
  • Deliver regional architecture, request routing, disaster recovery, and tested failover capabilities.
  • Improve Kubernetes networking, upgrades, patching, scaling, observability, capacity management, and recovery.
  • Make infrastructure and deployments safer through secure defaults, progressive delivery, automated guardrails, clear health signals, and fast rollback.
  • Partner with product teams to diagnose distributed-system problems across service, datastore, network, and platform boundaries.
  • Turn incidents, capacity limits, and operational signals into durable engineering improvements.
  • Write well-tested, observable, documented, and operable systems; lead design discussions, review code, share context, and mentor other engineers.

🎨 Skills & Experience We Value

You do not need every item below. We’re looking for demonstrated depth in several areas, sound judgement, and the ability to learn the rest.

Core

  • Clear communication and effective collaboration in a distributed team.
  • Technical influence: navigating ambiguity, explaining trade-offs, building consensus, and driving cross-team work
  • Developer experience: Experience building CI/CD, internal-platform, or self-service capabilities that make the safe path easier for other engineers.
  • Experience owning production systems through growth, failures, migrations, and incidents.
  • Production AWS, Kubernetes, and infrastructure-as-code experience.
  • Infrastructure as Code: Experience managing production infrastructure with Terraform or an equivalent tool, including DNS, IAM, monitoring, and secrets.
  • Strong reliability practices, including observability, incident response, recovery testing, and operational improvement.
  • Experience building platforms or self-service capabilities that make secure, reliable delivery easier for engineers.

Useful additional experience

  • Multi-region or multi-cluster systems, traffic routing, and disaster recovery.
  • Progressive delivery, deployment guardrails, and rollback.
  • Platform-level authentication, identity, or compliance controls implemented through engineering rather than paperwork.
  • Datastores: Experience operating relational databases, caches, or queues, including migrations, replication, failover, and workload isolation.

🗓 A Typical Day Might Include

  • Designing a reusable pattern for deploying or upgrading a Kubernetes environment.
  • Working through how a service, its data, and its traffic operate safely across regional clusters.
  • Reviewing Terraform, an architecture proposal, or a failover plan.
  • Investigating a capacity hotspot, queue backlog, network constraint, or datastore limit with a product team.
  • Improving a deployment guardrail, health signal, runbook, or recovery exercise.
  • Turning an incident action into automation or a safer operational workflow.
  • Pairing with or mentoring another engineer on a difficult reliability problem.

✨ Why Join Buildkite

Buildkite combines meaningful technical problems with visible impact. We’re remote-first and async by default: we document our reasoning, use meetings intentionally, and protect time for focused work. You’ll be trusted to exercise judgement, encouraged to experiment, and supported by colleagues who care about both the work and one another.

Benefits include:

  • Work from anywhere in the world for up to 6 weeks a year, on top of a fully remote setup — great work doesn't have to happen from one desk.
  • 20 days paid annual leave plus 10 days sick & carer's leave — actively encouraged, we want you to be at your best and time off is a part of that.
  • 16 weeks paid parental leave for primary carers (6 for secondary) — becoming a parent is life changing and we're committed to supporting you and your family through it.
  • A remote-work allowance for home-office kit, co-working, or working elsewhere — we want you to have everything you need to be set up for success in your role.
  • Confidential EAP & mental health support — sometimes life gets hard. You should be able to access the resources and support you need.
  • Support for learning and growth — with time and budget for conferences, courses, and coaching so you have the opportunity to grow, improve your craft, and build mastery.
  • Retirement and health coverage (based on your region) — build towards your future, your way.

Exact benefits may vary slightly by country. We'll walk through the specifics for your location.

Like all of Buildkite, this role is 100 % remote, however it does need to be located in either the ANZ or PST timezone.

🌈 Equal Opportunity Employer

At Buildkite, we value diversity and celebrate all types of skills, backgrounds, and experiences. We’re dedicated to fostering an inclusive environment and providing reasonable accommodations throughout our recruitment process.

If you need any accommodations or support during the application or interview process, please reach out to us at accommodations@buildkite.com.

More remote jobs at Buildkite

Similar remote jobs