Unknown Company

Principal AI Systems Architect, MoE Runtime & Memory Hierarchy

san jose, ca • Posted 6 days ago
Remote Full Time IT Management & IT Project Management

A leader in memory solutions we deliver best-in-class products. For consumers and businesses alike, we have committed ourselves to quality, performance, and customer service. It is our mission to design a broad range of memory solutions focused on high performance, exceptional reliability, and outstanding customer service. We provide our partners with top-notch services, from instant technical support and professional services to streamlined and efficient delivery processes that leverage our many years of experience. Our cutting-edge products drive higher performance, increase efficiency, and deliver unwavering reliability. Our continued commitment to innovation ensures we are positioned to deliver superior solutions for our customers, partners, and communities.

About the Role

Own the end-to-end technical architecture for sparse MoE inference and model-aware memory and storage management. This person turns the strategic direction into concrete object models, APIs, data paths, cache policies, runtime integrations, performance models, and prototype specifications.

The role should bridge model architecture, runtime software, operating-system memory management, accelerator execution, storage firmware, and performance engineering. It is the central technical owner for the AI-HLC architecture.

Primary responsibilities

  • Architect the model-aware residency system. Define experts, expert bundles, shared experts, KV blocks, recurrent state, activations, routing metadata, and agent state as first-class objects with explicit lifecycle, identity, placement, quality, deadline, and security attributes.
  • Design HLC policies and mechanisms. Own hot, warm, cold, compressed, resident, prefetched, and evicted states; admission and eviction policies; routing-frequency telemetry; reuse prediction; prefetching; compression decisions; and recovery behavior.
  • Develop quantitative workload models. Model bytes per token, expert reuse distance, cache hit rate, cold-miss bursts, decode throughput, UFS or NVMe bandwidth, memory-write bandwidth, interconnect latency, thermal limits, and cost per token.
  • Own both deployment tracks. For mobile and edge, design compressed expert storage, SoC-side decode, pinned residency, heterogeneous CPU/GPU/NPU execution, and UFS-aware scheduling. For enterprise and near-data systems, evaluate host-only execution, CXL memory, computational SSDs, in-drive acceleration, multi-drive sharding, and activation-oriented host-device protocols.
  • Integrate with inference runtimes. Work with low-level and production runtimes such as llama.cpp, vLLM, SGLang, ExecuTorch, LiteRT, vendor accelerator stacks, and custom runtime components. Determine what can be implemented above the operating system and what requires kernel, firmware, compiler, or hardware changes.
  • Architect agentic-state storage. Define how agent sessions, reusable prefixes, KV state, checkpoints, tool results, files, embeddings, intermediate artifacts, and shared multi-agent state move across accelerator memory, DRAM, CXL, local SSD, and remote storage.
  • Design the runtime-to-system interface. Specify APIs for expert requests, prefetch deadlines, routing hints, cache state, quality constraints, memory pressure, compression eligibility, placement, cancellation, and telemetry.
  • Lead performance and correctness reviews. Ensure that optimizations improve p95 and p99 latency, energy, and total system cost without silently changing routing behavior, model quality, security isolation, or data correctness.
  • Guide implementation. Review C++, Python, kernel, firmware, and simulator designs; mentor the senior engineer; and translate prototype findings into production architecture.
  • 15+ years in systems software, storage software, high-performance computing, operating systems, ML infrastructure, embedded systems, or related engineering.
  • BS/MS degree in Computer Science or equivalent experience
  • At least 5 years in AI/ML infrastructure, transformer inference, model serving, or accelerator software.
  • Demonstrated experience architecting complex systems across multiple software or hardware layers.
  • Strong understanding of transformer execution, MoE routing, expert FFNs, KV caches, quantization, model formats, batching, scheduling, and inference performance.
  • Strong C++ and Python skills, with hands-on experience debugging concurrency, memory, I/O, and performance problems.
  • Deep Linux systems knowledge, including virtual memory, mmap, page cache, faults, NUMA, DMA, asynchronous I/O, process telemetry, and memory-pressure behavior.
  • Experience with at least one major storage interface or stack such as NVMe, UFS, PCIe, block I/O, SSD firmware, object storage, or distributed caching.
  • Ability to translate architecture into testable requirements, interface specifications, and implementation milestones.

Preferred qualifications

  • Android or embedded Linux experience, including Perfetto, cgroups, DMA buffers, vendor delegates, or mobile thermal and power analysis.
  • Experience with NPU, DSP, GPU, FPGA, or custom accelerator programming.
  • Familiarity with entropy coding, model compression, INT4 or mixed-precision inference, tensor packing, and accelerator-native weight layouts.
  • Experience with CXL, computational storage, SmartSSD-class platforms, remote memory, RDMA, GPU-direct I/O, or disaggregated inference.
  • Familiarity with agent frameworks, context engineering, MCP/A2A integration, workflow persistence, and multi-agent execution.
  • Experience designing secure multi-tenant protocols or executing AI workloads in trusted hardware and firmware environments.

About Lexar Enterprise

Building upon the foundation and the credibility of the well-established Lexar brand, Lexar Enterprise designs and manufactures memory and storage solutions that include embedded storage, mobile memory, solid-state drives, and memory modules in commercial, industrial, and automotive grades. Our cutting-edge products are deployed globally to drive higher performance, increase efficiencies, and deliver unwavering reliability. Our continued commitment to innovation ensures we are positioned to deliver superior solutions for our customers, partners, and communities

Lexar Enterprise. 1737 N First Street, Suite 680, San Jose, CA USA

Unsolicited resumes sent to Lexar from third-party recruiters and placement agencies do not:

(a) constitute any type of relationship with Lexar or

(b) obligate Lexar to pay fees should we elect to hire from those resumes. Please do not contact or present candidates directly to any Lexar personnel.act or present candidates directly to any Lexar personnel.

#J-18808-Ljbffr

Principal AI Systems Architect, MoE Runtime & Memory Hierarchy in san jose at Unknown Company

This position is listed as full time and able to be worked remotely.

Back to Job Search