As an Agentic AI Software Engineer, you will own the design and development of intelligent systems that reason over real-time network telemetry, automate operational decisions, and surface actionable insights to cluster operators. You will integrate frontier LLM APIs alongside purpose-trained smaller models fine-tuned for networking-specific tasks to enable autonomous network management. You will solve hard problems at the intersection of LLM orchestration, retrieval-augmented generation, and low-latency infrastructure. This role spans architecture through production deployment across Aria’s cluster software stack.
Responsibilities
- Design and build agentic workflows like tool use, multi-step reasoning, and RAG pipelines that operate over fine-grained GPU-cluster network telemetry
- Integrate LLM APIs (OpenAI, Anthropic, Google or equivalent) into production systems with reliable orchestration, structured outputs, and robust failure handling
- Fine-tune and distill smaller task-specific models (LoRA, PEFT, or equivalent) for networking-domain tasks including anomaly detection, fault classification, and predictive analytics
- Architect the AI layer of Aria Cluster Software, defining how agentic components interact with the telemetry pipeline, control plane, and operator-facing interfaces
- Define evaluation frameworks to measure AI system accuracy, latency, and decision quality against real network operational outcomes
- Partner with telemetry and backend engineers to shape data schemas and streaming infrastructure that serve AI workloads efficiently
- Ship AI capabilities end to end with appropriate observability, rollback mechanisms, and confidence calibration
Qualifications
- Bachelor’s degree (or foreign equivalent) in Computer Science, Software Engineering, Computer Engineering, or a closely related technical field
- Minimum 5 years of software engineering experience, with at least 2 years in hands-on AI/ML product development
- Experience designing and implementing agentic workflows including multi-step reasoning, retrieval-augmented generation, and dynamic tool invocation
- Experience fine-tuning or training smaller models using LoRA, PEFT, distillation, or equivalent techniques
- Working knowledge of ML frameworks such as PyTorch, JAX, Tensorflow, etc.
- Experience building AI-powered control or automation systems that act on live system state is preferred
- Experience applying AI/ML to networking, infrastructure, or observability domains is a significant plus