Thinking Machines Lab Inc. in San Francisco is seeking an engineer to help build and scale the core data infrastructure for distributed training pipelines and multimodal data catalogs.
You will work with researchers to accelerate experiments, develop datasets, and improve reliability across LLM research platforms, using Spark, Kafka, Beam, Ray, and Delta Lake.
This role emphasizes end-to-end ownership, collaboration across teams, and building from the ground up in a fast-moving environment.
#J-18808-LjbffrData Infrastructure Engineer for Large-Scale AI Pipelines in san francisco at Unknown Company
This position is listed as full time and onsite.