Unknown Company

Infrastructure Software Engineer - Topology and Observability

mountain view, ca • Posted Today
Remote Full Time Networks & Systems

Infrastructure Software Engineer - Topology and Observability

Mountain View, CA or Remote

Joyent powers the global cloud infrastructure and developer platform providing back-end services for Samsung's billions of devices. Joyent's data center footprint is within 100ms latency to 70% of the world's population, while our multi-cloud, Kubernetes-based developer platform extends our reach to additional resource regions. We're operating at hyperscale to power workloads that bring capability and delight to Samsung's employees and customers.

Job Summary

Our infrastructure spans multiple layers — physical network, compute hosts, virtualization, managed services, and the tenant workloads running on top of them. Each layer has its own configuration management, inventory, and monitoring systems. What we lack is a tool that cuts across those layers and answers, in one view, "which hosts, network paths, and services does this workload depend on?"
In this role you will design and build a system that unifies scattered configuration, inventory, and observability data into a single infrastructure knowledge graph, and delivers multi-layer topology maps and service dependency maps on top of it. Users can overlay physical, logical, and service layers selectively, drill down hierarchically from a region to an individual node, and jump from any component to its related alerts, tickets, runbooks, and dashboards. The system serves as a shared operational view for the network, compute, SRE, and security teams, and becomes the foundation for fast root cause analysis, failure-domain analysis, and pre-change impact assessment.

Job Responsibilities

Design and maintain an infrastructure topology data model spanning physical, logical, and service layers — schemas, node and edge metadata, and cross-layer relationships

Build read-only, low-impact collectors and parsers that gather data from configuration management tools, network devices, hypervisors, and metrics/log systems

Build pipelines that reconcile declared configuration against observed state to detect and report drift

Develop the interactive web frontend and backend services providing multi-layer, hierarchical drill-down navigation, directional traffic and dependency flow visualization, and deep links into related systems

Design dependency-based blast radius calculation, failure-domain analysis, and change scenario simulation

Provide topology and dependency data as the foundation for alert correlation and automated root cause analysis (RCA), and integrate with incident response, alerting, and ticketing systems and AIOps pipelines — including human-in-the-loop review and approval workflows

Define data accuracy and freshness metrics; own the reliability and performance of the system itself

Gather requirements from consuming teams and expand adoption in stages (static artifacts → hosted service)

Understanding of hypervisor-based virtualization and compute host operations

Understanding of the data models and limitations of observability stacks (Prometheus, OpenTelemetry, logging and tracing systems)

Sound judgment in designing data collection that minimizes impact on production systems

Ownership -Take ownership of projects, ensuring excellence in execution and accountability for results. Foster a sense of responsibility and pride in delivering high-quality work

Innovation - Drive innovation by proposing and implementing creative solutions to challenges. Stay abreast of industry trends and technologies, bringing fresh ideas to the table

Customer focus - Understand and prioritize customer needs, striving to exceed expectations in every interaction. Collaborate with cross-functional teams to ensure the delivery of customer-centric solutions

Teamwork - Embrace a collaborative and inclusive approach, working seamlessly with colleagues to achieve common goals

Education & Experience

5+ years in infrastructure, platform, SRE, or network automation, with hands-on experience across both cloud and on-premises environments

Working knowledge of L2/L3 networking: routing protocols (e.g. BGP), overlay networks (e.g. VXLAN), and the Linux networking stack

Experience building Python data pipelines, including schema definition and validation tooling

Experience modeling and processing data in analytical stores (columnar databases, graph databases, or similar)

Experience visualizing graph/topology data on the web (D3, Cytoscape, or similar) and handling the performance and readability challenges of large graphs

Preferred qualifications

Experience operating large-scale multi-region cloud or IaaS environments

Experience building or integrating CMDBs, inventories, or network sources of truth (NetBox, Nautobot, or similar)

Experience with service dependency mapping, failure-domain analysis, or what-if / digital-twin style simulation

AIOps experience: alert correlation and noise reduction, anomaly detection, automated root cause analysis, and building incident response automation pipelines

Experience applying LLM-based agents to operational workflows — particularly using LLMs to parse and summarize unstructured sources (documents, free-form configuration, alert text) wrapped in deterministic validation and human review

Experience with automated diagram generation (draw.io, Graphviz, or similar)

Experience collaborating with security teams on exposure surface or attack path visualization

What success looks like (first 12 months)

3 months: Data model and collectors for the core layers are operational, and a validated physical and logical topology exists for at least one environment

6 months: Two or more teams use the system in real incident analysis, and configuration-vs-observed drift is reported on a regular cadence

12 months: The system runs as a hosted service, and impact assessment and change what-if analysis are part of standard procedures

What this role is not

A monitoring operations role focused on dashboard maintenance

A pure frontend or pure network engineering role — we are looking for someone who combines infrastructure understanding with data and visualization skills

Compensation and Benefits

Compensation for this position will vary among specific regions due to geographical differentials in the labor market, and actual pay will be determined considering factors such as relevant skills, experience, and comparison to other employees in the role. Therefore, the annual base compensation range for this role (depending on the geographical location) is expected to be between $ and $ .

Regular full-time employees (salaried or hourly) have access to benefits including Medical, Dental, Vision, Life Insurance, 401(k), Employee Purchase Program, Vacation and Sick leave, electronic reimbursement and many more. In addition, regular full-time employees (salaried or hourly) are eligible for bonus compensation based on individual, department, and company performance.

About Joyent

Joyent, a wholly-owned subsidiary of Samsung, is the open cloud company. Joyent builds technology, at thepinnacle of scale, performance, stability, and security to accelerate the transformation toward the mobile andcloud-centric world. Joyent designs, builds and manages market competitive cloud computing solutions andservices for Samsung Electronics and its partners at global scale.

Joyent is committed to employing a diverse workforce and providing Equal Employment Opportunities for allindividuals regardless of race, color, religion, gender, age, national origin, marital status, sexual orientation,gender identity, status as a protected veteran, genetic information, status as a qualified individual with adisability, or any other characteristic protected by law.

Balance Work/Life with time off to truly relax and reboot.

Work Remotely

We work seamlessly together as one from our worldwide offices and offer telecommuting.

Referral Bonus

Refer someone from your network who gets hired and we'll show our appreciation through our referral bonusprogram.

Let us help you plan for your future retirement with Matched 401K Contributions

Discounts

Who doesn't like a deal? Get discounts on Samsung and affiliate company products.

Health

We care about your and your family’s wellbeing. Stay healthy with our medical, dental and vision plans.

Training and Education

Grow your career with training resources and certifications

Next Generation Tech

We work, build and collaborate with next generation technologies in data, AI and compute

We use, sponsor, and collaborate extensively with open source projects

#J-18808-Ljbffr

Infrastructure Software Engineer - Topology and Observability in mountain view at Unknown Company

This position is listed as full time and able to be worked remotely.

Back to Job Search