Matcha
Fireworks AI logo

Fireworks AI careers: remote jobs, culture and where they hire

The inference platform built by the PyTorch team, serving 40 trillion tokens a day on open models.

Team of 283$1.8B raised, $17.5B valuation (Series D, July 2026)Backed by Index Ventures, Sequoia, Benchmark, Lightspeed and NvidiaHybrid in San Mateo, New York, London and Singapore, a few US-remote roles

Who backs Fireworks AI

Fireworks AI has raised about $1.8 billion. The $25M Series A closed on March 27, 2024, led by Benchmark, whose general partner Eric Vishria joined the board, with Sequoia Capital, Databricks Ventures and angels including Frank Slootman and Alexandr Wang. Sequoia led the $52M Series B on July 11, 2024 at a $552M valuation, joined by NVIDIA, AMD and MongoDB Ventures. The $250M Series C on October 28, 2025 (about $230M primary plus a $20M secondary) was co-led by Lightspeed, Index Ventures and Evantic at a $4B post-money valuation, with Sequoia, NVIDIA, AMD, MongoDB and Databricks participating, when annualized revenue had passed $280M. On July 15, 2026 the company closed a $1.505B Series D at a $17.5B valuation led by Atreides Management, Index Ventures and TCV, with Nvidia, Lightspeed, Evantic, Bessemer, Menlo Ventures, Insight Partners, Ontario Teachers' Pension Plan, Lone Pine and 20VC. The money goes to compute (capacity already comes from more than 20 suppliers, plus a Microsoft partnership announced in March 2026) and to engineering headcount, which the CEO said she intends to triple by the end of 2026.

What is Fireworks AI?

Fireworks AI runs open models in production for other companies. The platform offers serverless and dedicated inference, fine-tuning, reinforcement fine-tuning and reserved GPU capacity across hundreds of open models for text, image, embedding, audio and multimodal work, on top of FireAttention, a custom CUDA kernel the company says delivered up to 12x faster inference than vLLM on long-context workloads. The founders started it in late 2022, before ChatGPT launched, and chose inference over training on the argument that training scales with a small pool of researchers while inference scales with the whole world's users.

By July 2026 the platform served more than 40 trillion tokens a day for over 10,000 companies and crossed $1 billion in annualized revenue, five times the year before. Customers include Cursor, Notion, Uber, DoorDash, Shopify, Upwork, Samsung, Harvey, Doximity and GitLab. Cursor serves a model at about 1,000 tokens per second on Fireworks, 13 times faster than a standard Llama 70B deployment, Notion cut a feature's latency from about two seconds to 350 milliseconds, and roughly 95% of tokens served now come from models specialized on customer data rather than stock open models.

Who founded Fireworks AI

Lin Qiao, CEO and co-founder. Spent seven years at Meta, where she led PyTorch and rebuilt the company's AI infrastructure stack, growing her team from five engineers to more than 300 and running a system doing over 5 trillion inferences a day by the time she left. Started at IBM in 2003 and was later a technical lead at LinkedIn, with database publications in PVLDB and SIGMOD. B.S. and M.S. in Computer Science from Fudan University and a Ph.D. from UC Santa Barbara; she founded Fireworks at 48 after deliberately delaying a startup idea she had in 2015 until she had learned to manage people.

Dmytro Dzhulgakov, CTO and co-founder. Spent over ten years at Meta after interning at Google and Facebook: core developer of Caffe2, co-creator of the ONNX interchange format with Microsoft and Amazon, co-founder of Facebook's first company-wide production AI platform, and a PyTorch core maintainer who drove its production features through the 1.0 release. From Kharkiv, Ukraine; B.Sc. and M.Sc. in Applied Mathematics from the National Technical University Kharkiv Polytechnic Institute.

Dmytro Ivchenko, Co-founder. Led PyTorch development for ranking systems at Meta, after earlier engineering work at LinkedIn. Graduate of Kyiv Polytechnic.

James Reed, Co-founder. Runs multimedia inference (image, video and audio) at Fireworks, including what he describes as the lowest-latency Stable Diffusion XL serving platform on the market and the official inference backend for Stability AI. Before that, Staff Software Engineer on the PyTorch team at Meta (2017 to 2022), leading torch.fx, PiPPy, TorchScript and ONNX work. B.S. in Computer Engineering from Virginia Tech, summa cum laude, December 2016.

Chenyu Zhao, Co-founder. Previously the Google Vertex AI lead. B.S. from UC Berkeley (2010 to 2013).

Benny Chen, Co-founder. Led advertising infrastructure at Meta before co-founding Fireworks.

Pawel Garbacki, Co-founder. Co-founder and researcher, working on model architecture, fine-tuning, alignment, multimodality, long context and inference optimization. Nine years at Facebook, latterly as a Principal Engineer and core ML lead for News Feed, after time at Google and IBM's T.J. Watson lab. B.Sc. in Mathematics and M.Sc. in Computer Science from the University of Warsaw, M.Sc. from VU Amsterdam, and a Ph.D. in Computer Science from Delft University of Technology and VU Amsterdam.

Where does Fireworks AI hire remotely?

Fireworks AI is headquartered at 900 Concar Drive in San Mateo, California and hires in the United States, the United Kingdom and Singapore, with offices in San Mateo, San Francisco, New York (295 Madison Avenue), London and Singapore. Most roles are hybrid at one of those offices; the few remote openings are United States only and sit in sales enablement and enterprise sales. The team is primarily US-based, and London staff are expected to work across time zones and travel to San Mateo for planning.

Open remote jobs at Fireworks AI

Describe your next role, cut the noise

Fireworks AI has no live openings on Matcha right now. Set your preferences above and we will email you when new remote roles at companies like Fireworks AI open up.

Why Fireworks AI is exciting to work for

The scale is real and recent: 40 trillion tokens a day, revenue up 5x in a year, and a $17.5B valuation reached less than four years after founding. The engineering problems sit at the layer where that shows up, serving large models with low tail latency, sharing GPUs across thousands of customers, and keeping a fleet up when a coding assistant's users depend on it. A PhD new grad posting puts it plainly: "here you get a production fleet as your testbed and your work ships."

The people you would learn from built PyTorch. Four of the seven co-founders worked on PyTorch at Meta, and the CEO ran the team that took Meta's AI infrastructure to over 5 trillion inferences a day. Aishwarya Srinivasan, who works at Fireworks, wrote in 2025 that it is "one of the most densely talented teams I've ever had the privilege to work with. Every single day feels like a learning experience, and I don't say that lightly."

The company hires at every level from new grad to Head of. New grads get a senior engineer as a mentor, a structured ramp on scoped real problems, and one application that routes to teams across inference, training, cloud infrastructure, developer platform and product engineering.

Fireworks AI vibe check

Work week
Three days on site: postings cite a San Mateo in-office policy of Monday, Wednesday and Friday, and product marketing roles state "in the office three days per week."
Pace
Reviewers describe exceptional peers and a demanding pace with no formal career ladder (Glassdoor 3.6/5 overall and 2.9/5 on work-life balance from about 8 reviews, as summarized by JobsByCulture, May 2026).
Equity
Nearly every posting is tagged "Offers Equity"; the company is currently hiring a Compensation Partner to build out its options and RSU framework.
Notable
On-call is being formalized: the Reliability Engineering hire owns SLOs, error budgets and on-call expectations across engineering, and every team is expected to run what it builds. GTM roles travel 15 to 25% for offsites, kickoffs and conferences.

What to expect from the culture

Two principles come straight from the CEO: "customer first" and "high velocity of innovation", with the caveat that infrastructure "must also be stable and reliable." The Reliability Engineering posting turns that into a rule, "Every team owns the reliability of what they build," and the boilerplate on every posting promises "no bureaucracy, just results."

Extreme ownership is the phrase Lin Qiao uses most. She joined Facebook in 2015 partly to learn how a company culture works, and what struck her was employees fixing bugs in code they did not write and reporting a broken dashboard in the lobby: "It's magical. It brings the best out of every person." At Fireworks she says people who show that ownership "rise up without you asking them to", and GTM postings ask for people who thrive in "ambiguity, extreme ownership, and the pace of a high-growth startup." She hires on aptitude over experience ("I look for hunger, motivation, and a fast-learning mindset"), and the EMEA engineering posting follows suit, written as a bucket for "exceptional builders at different levels of experience rather than a fixed job spec" with anywhere from 3 to 10 years fitting. The structure has stayed relatively flat while headcount went from about 50 to 200 in roughly a year, and the CTO once embedded at a customer site for months.

AI tooling is assumed, not optional. Even the BS/MS new grad posting requires "fluency with AI coding harnesses and a habit of using them well", and engineering roles ask for "agentic development" as a core part of how you build, "with judgment about when to trust the output." Qiao has said that "coding interview in the past is going to [be replaced] by how good you are at using coding agents", since coding agents now perform at roughly a junior engineer's level.

How to get hired at Fireworks AI

Fireworks AI screens for first-principles thinking before anything else. The EMEA engineering posting says it outright: reasoning about unfamiliar problems from the ground up "is the primary interview signal and can outweigh other qualifications", and you are evaluated on it directly in the loop. Candidate reports collected on Blind and Glassdoor (summarized by Design Gurus) describe a roughly 30-minute recruiter call, a 45 to 60 minute technical screen mixing coding with ML-systems questions on inference serving and GPU utilization, a take-home for some roles (one candidate built a simple chat playground on the Fireworks API, reviewed in detail at the onsite), and an onsite of four to seven rounds spanning coding, one or two system design sessions on model serving and GPU scheduling, and cross-team conversations. Software engineering loops take about five weeks end to end.

Level is set during the loop rather than by the posting; the Developer Advocate role, for example, is open at Senior or Staff with the level "determined through the interview process based on your experience." New grad roles accept a Bachelor's or Master's completed within the last six months or by summer 2027, with December 2026 graduates prioritized, and PhD start dates flex around thesis defense. Go-to-market hires go through Firecademy, a monthly in-person onboarding week in San Mateo.

Fireworks AI hiring FAQ

How much does Fireworks AI pay?

Fireworks AI publishes a base band on every US posting, as of September 2026: Member of Technical Staff for LLM, cloud or full-stack infrastructure $175,000 to $220,000; Software Engineer and Reliability Engineering $200,000 to $290,000; Research, AI Training Infrastructure and Evals $210,000 to $320,000; Applied Machine Learning Engineer $170,000 to $240,000; AI Field Engineer $200,000 to $260,000; Product Manager $170,000 to $300,000; Enterprise Account Executive $300,000 to $330,000; BDR $90,000 to $100,000. New grads get $160,000 to $180,000 (BS/MS) or $200,000 to $300,000 (PhD). London bands run £200,000 to £220,000 for a Partnerships Lead and £70,000 to £100,000 for a BDR; Singapore engineering pays SGD 150,000 to 250,000. Equity comes on top.

Is Fireworks AI remote?

Mostly no. Fireworks AI is a hybrid company: 64 of its 79 open roles require three days a week in San Mateo, San Francisco, New York, London or Singapore, and 2 are fully on site. The exceptions are a US-remote Sales Enablement Manager ($175,000 to $200,000), Sales Enablement Lead ($210,000 to $250,000) and Enterprise Account Executive ($300,000 to $330,000). No engineering role is listed as remote.

Does Fireworks AI hire new grads or interns?

Fireworks AI hires new grads but has no internship program. Three new grad tracks are open for 2026: Member of Technical Staff (BS/MS), and PhD tracks in Research and in Systems Infrastructure covering inference performance, distributed training and cloud infrastructure. There are no internship listings on the careers board, and an August 2026 guide for students confirms no formal program exists yet.

What is Fireworks AI's tech stack?

Python plus a systems language (Go, C++, Rust) for services, CUDA for kernel work on FireAttention, PyTorch throughout, and Kubernetes on multi-cloud GPU capacity from more than 20 suppliers. Infrastructure postings want familiarity with open inference engines like vLLM, SGLang and TensorRT-LLM, plus LLM serving concepts such as disaggregated prefill and KV-cache memory estimation. Security runs on CrowdStrike, Incident.io and Console AI.

How does Fireworks AI make money?

Fireworks AI charges for usage: pay-per-token serverless inference, dedicated deployments, reserved capacity and fine-tuning. Sacra estimates blended revenue of about $28,000 per customer company a year and a gross margin around 50%, with a 60% target, and the CEO said at the Series D announcement that serving a model of equivalent quality to a closed model costs customers five to ten times less.

Want remote roles like these in your inbox? Tell us what you are looking for.

Describe your next role, cut the noise