Unknown Company

AI Systems Engineer

san francisco, ca • Posted 5 days ago
Onsite Full Time General

AI Systems EngineerTransluce is a fast-moving research lab building the public tech stack for understanding and debugging AI systems. We build world-class, AI-backed analysis tools and use these to set industry standards for evaluation. We are a non-profit with a mission to steer the development of AI for the public good.We are looking for an exceptional AI systems engineer to lead the design and development of our core ML stack, building systems that can scale to thousands of GPUs and performantly query trillion-token databases.

As an early member of a highly collaborative team, you will be free to innovate and move fast, building high-impact systems from the ground up. As part of a mission-focused non-profit, your work will have high direct impact (e.g. used by governments to inform AI policy) and cross-organisational reach (open-source tools the entire community can build on).Core ResponsibilitiesSet overall code culture and tooling for a fast-growing orgHelp to solve our core technical challenges across verticals.

Examples include:Docent: High-concurrency container-based evals with quick ability to iterate on interventions to agentic trajectoriesDeterministic sandbox execution of code that can efficiently restore state from checkpointsInterpretability: Inference stacks that are as performant as vLLM but flexible enough to allow complex model introspection and intervention, steering, configurable sampling, etc., and that can scale to 400B+ parameter modelsBehavior elicitation: Distributed RL training and roll-outs allowing thousands of concurrent rollouts across machinesBuild great internal tools to speed up the teamHelp tone-set in the organization around best practices for building and path-set on what infra we should buildHelp other team members think through infra challengesQualities of a Strong CandidateExceptional programmer fluent in PythonBare metal optimization: know GPUs, other accelerators in and out (low-level performance + optimization + parallel programming)Experience engineering at scale (distributed systems, reliability, architecture design)Leader on global code quality and health (designing good primitives, managing complexity and scale)Bonus: can set up LLM pipelines, e.g. multiple specialized LLMs interacting with each other in a performant and reliable wayBonus: experience with open-source community management

Back to Job Search