Infrastructure Engineer
Mercor · San Francisco, United States +1 · 20d ago
About Mercor
Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
About the Role
As an Infrastructure Engineer at Mercor, you’ll build and scale the systems that power our rapid growth. You’ll ensure our infrastructure is highly available, cost-effective, and able to handle explosive traffic and compute demands. You’ll work closely with engineers across product, research, and operations to design scalable architectures, streamline deployments, and improve observability.
We're hiring across Infrastructure: Platform, Developer Productivity, Production Engineering, Storage and Databases. We team-match after the first screen, so apply even if your background leans toward one area.
What You'll Work On
Scale our multi-tenant services that fronts every model provider we use, so dozens of internal teams and products get isolated quotas, routing, observability, and capacity on demand without filing a ticket.
Run Temporal, Postgres, MongoDB, and our Kubernetes fleet at a scale where "it worked last month" isn't a guarantee, and design for the next 10x.
Own the reliability program: define SLOs that matter, kill noisy alerts, make CI/CD deploys boring, and turn every incident into a durable fix rather than a runbook entry.
Build the developer platform for an engineering org that ships dozens of times a day, and where coding agents are now first-class users of our CI, sandboxes, and deploy pipelines.
Design network and identity boundaries across production, preprod, and research compute so teams move fast without cross-environment risk.
Make the cost picture legible: attribute spend across data centers, and inference providers, and find the architectural changes that bend the curve.
Build the tooling that lets the rest of engineering self-serve: our on-call is already AI-triaged and auto-assigned; you'll decide what gets automated next.
What We're Looking For
We care far more about how you reason about systems than which tools you've used. Strong candidates typically have:
A deep grasp of reliability and scalability fundamentals: failure modes, backpressure, idempotency, capacity planning, and how distributed systems actually break under load.
Experience operating production systems that real users depend on, and the scars to prove it.
Comfort writing code in a language like Python or Go, and reading code in whatever the incident requires.
Familiarity with cloud infrastructure (we're on AWS) and infrastructure-as-code (we use Terraform). You don't need to be an expert; you need to be curious and fast to pick it up.
Working knowledge of containers and how they're deployed. Kubernetes experience is a plus, not a gate.
High ownership: you see a gap, you write the proposal, you ship it, you carry the pager for it.
You Might Be a Great Fit If
You've been the person who understood why the system fell over when nobody else did.
You've built internal platforms or tooling that other engineers loved using.
You've operated databases, message queues, or workflow engines at meaningful scale.
You've come from backend or DevOps and want to own infrastructure end to end.
How We Work
Small, senior team with direct access to the head of infra and to engineering leadership.
In-person five days a week in SF, or NYC. Infra is a team sport for us.
We use AI coding agents heavily and expect you to as well; the interesting work is deciding what they should do, not typing.
Weekly on-call rotation with structured handoffs; on-call load is a metric we actively drive down.
Why Mercor
Impact: Your work powers how the world’s leading AI labs train and test their models.
Learning: Get early insights into frontier model capabilities months before the market.
Growth: Work on both infrastructure and research-adjacent projects with fast paths to ownership.
Benefits
Bi-annual performance bonus structure
Generous equity grant vested over 4 years
Up to $15k Relocation bonus
$10K housing bonus (if you live within 0.5 miles of our office)
$1.5K monthly stipend for meals
Free Equinox membership
$200 monthly laundry reimbursement
$200 monthly personal wellness reimbursement
Health, Dental, Vision insurance