Unknown Company

Distributed Systems Engineer, GPU Infrastructure

northern, ky • Posted 2 days ago
Remote Full Time IT & Technology

Engineering San Francisco or remote - Full-time

Own the clusters, GPU fleet, and production systems that turn inference research into a reliable service

About Coral Bricks

Our mission is to make frontier intelligence affordable and accessible to everyone. Frontier models are finally here - but almost nobody can afford to use them freely. People have token anxiety: they meter every call, ration every context window, and settle for weaker models because the best ones are priced out of everyday use.

We're building the inference platform that ends that, starting with the workloads that feel the squeeze hardest: research and coding agents that swarm across multiple models, plan, call tools for hours, and reason over big context. Classic LLM serving was never built for them - rate limits that throttle real workloads, queues that stretch a 20-minute job into a 4-hour one, costs that grow with every agent turn. Same models, same prompts - many times the tokens per second at a fraction of the cost.

The team is small, technical, and shipping. We also build in the open: a lot of the day-to-day happens in our Discord, where the developers building on Coral Bricks tell us what broke, compare numbers with us, and push on what we work on next.

The role

You'll own the operational systems behind our inference platform: the clusters, GPU fleet, deployment machinery, and control plane that keep models available and traffic moving. When research produces a faster serving technique or a new model drops, you'll turn it into a repeatable, observable, production launch.

This role is distinct from our inference research role. You won't be measured on inventing a new attention kernel. You'll be measured on whether we can provision capacity, place workloads, ship changes, recover from failures, and operate a growing fleet without heroics.

This is a founding-team role with broad ownership. You'll work across cloud infrastructure, distributed systems, networking, storage, deployment, and the serving layer where they meet.

What you’ll work on

  • Own our GPU clusters and fleet across cloud providers: capacity, provisioning, machine images, drivers, networking, storage, health, and cost.
  • Build the control-plane systems that place workloads, manage capacity, drain and replace unhealthy nodes, and recover cleanly from failures.
  • Turn model launches into a reliable process: bring up new weights, validate serving configurations, roll out safely, watch production behavior, and roll back when needed. You'll hear how a launch landed from the developers in our Discord, not only from the dashboards.
  • Build deployment and release systems for inference servers and the services around them, with fast feedback and clear failure modes.
  • Create the observability we need to operate the fleet: metrics, logs, traces, dashboards, alerts, and tools that make incidents diagnosable instead of mysterious.
  • Improve reliability at every layer - autoscaling, load balancing, failover, backpressure, graceful degradation, and capacity planning.
  • Automate recurring operational work so the fleet can grow faster than the team operating it.

You probably have

  • Strong backend or distributed-systems fundamentals and experience owning production services end to end.
  • Experience with Linux, containers, networking, and at least one major cloud platform. You can debug across application, host, and infrastructure boundaries.
  • Good instincts around reliability: staged rollouts, observability, failure isolation, incident response, and simple systems that are easy to operate.
  • Comfort working from symptoms to root cause. A failed launch, an unhealthy node, or a latency spike is a systems problem to investigate, not a ticket to hand off.
  • A high work ethic and excitement about early-stage startups. The pace is fast, the problems are open-ended, and everyone does a bit of everything.
  • A bias toward shipping and automation. You fix the immediate problem, then build the mechanism that keeps it from becoming routine work.

Bonus

  • Experience operating GPU or accelerator fleets, including NVIDIA or AMD drivers, topology, health checks, and failure modes.
  • Experience with Kubernetes, Nomad, Slurm, ECS, or another cluster scheduler - especially if you've had to work below its happy path.
  • Familiarity with vLLM, SGLang, TensorRT-LLM, PyTorch distributed, NCCL, or other model-serving and collective-communication systems.
  • Experience with multi-cloud capacity, bare-metal provisioning, or scarce-resource scheduling.
  • You've built an internal platform, scheduler, deployment system, or piece of infrastructure that other engineers trusted in production.

Compensation

$120,000-$200,000 base salary, plus 0.25%-$2.0% equity. Where you land depends on experience, and cash and equity move together - take less of one and we'll weight the other.

Equity vests over four years with a one-year cliff. Health, dental, and vision coverage, and flexible time off.

Founding engineers shape the platform, the technical direction, and the team we build around it.

#J-18808-Ljbffr

Distributed Systems Engineer, GPU Infrastructure in northern at Unknown Company

This position is listed as full time and able to be worked remotely.

Back to Job Search