Principal Platform Engineer
The Principal Platform Engineer is an enterprise-level individual contributor within the Technology Infrastructure Hosting team. The role shapes hosting-platform vision, strategy, roadmaps, governance, priorities, investment recommendations, and delivery outcomes for complex, business-critical initiatives across multiple teams and systems. It develops reusable modules, golden paths, self-service capabilities, engineering standards, evaluation methods, and operational controls across Windows Server, Active Directory and identity integrations; AIX/POWER and Linux; VMware, Nutanix, Hyper-V and related compute; storage, backup, SAN and NAS; middleware; networking and firewalls; Datadog, Nagios and related observability; ServiceNow CMDB, ITSM and workflow integrations; and adjacent infrastructure services. The engineer remains hands-on with code and production systems and converts approved designs, domain standards, runbooks, lifecycle obligations, and recurring work into secure, scalable, reliable, cost-effective, and version-controlled capabilities using IaC, configuration management, CI/CD, deterministic orchestration, and governed AI where it adds value.
This role operates with broad decision-impacting opportunities over platform engineering priorities and outcomes without direct people-management responsibility. It aligns senior leaders and cross-functional stakeholders, establishes governance for consistent delivery, coaches senior technical leaders and engineers, resolves systemic cross-domain problems, and challenges unsupported completion claims with evidence. Engineering-whether manual, scripted, IaC-based, or AI-assisted-must improve scalability, developer and operator experience, reliability, security, lifecycle currency, recoverability, observability, auditability, service and financial outcomes.
Our flexible/hybrid work schedule includes 3 in-person days in our Salisbury, NC office and 2 remote days. Applicants must be currently authorized to work in the United States on a full-time basis.
Core Responsibilities
- Platform strategy and roadmap.
- Governance and portfolio leadership.
- Cross-domain platform engineering.
- Platform product and experience management.
- Infrastructure as Code and CI/CD.
- Runbook-to-automation engineering.
- Generative and Agentic AI for operations.
- Prompt and context engineering.
- Agent harnesses and orchestration.
- Continuous agentic improvement.
- Evaluation and quality engineering.
- Financial and capacity stewardship.
- Security, risk, and compliance by design.
- Reliability and observability.
- Platform lifecycle and resilience.
- Basis Engineering and Definition of Done.
- Technical leadership and organizational capability.
Required Infrastructure Domain Breadth
Windows and identity
Windows Server lifecycle, Active Directory/GPO, DNS, privileged access, patching, configuration baselines, hybrid identity dependencies, PowerShell/DSC and recovery.
AIX and Linux
AIX/POWER, HMC/PowerVM/NIM, RHEL or equivalent Linux, package and kernel lifecycle, SSH/PAM/sudo, Ansible, filesystem and multipath dependencies, patching and recovery.
Compute and virtualization
VMware vSphere/vCenter/ESXi, Nutanix AOS/Prism, Hyper-V or comparable platforms; capacity, firmware/compatibility, cluster resilience, migration, lifecycle and automation.
Storage, backup and SAN
Block/file/storage services, Fibre Channel zoning and multipath, SAN/NAS lifecycle, backup policy, immutable protection, capacity, replication and tested restore or failover.
Middleware and runtime
Application servers, web tiers, messaging, integration runtimes or comparable middleware; configuration, certificates, dependencies, patching, HA, logging and performance.
Networking and firewalls
IP/DNS/load-balancing dependencies, routing, segmentation, firewall policy, ports and protocols, network change validation, telemetry and rollback coordination.
Observability
Datadog, Nagios or comparable platforms; coverage, tagging, service checks, metrics/logs/traces, SLI/SLO, actionable alerts, dependency suppression, synthetic health and evidence retention.
Service management
ServiceNow CMDB/Discovery, CI identification and reconciliation, completeness/correctness/compliance, service mapping, catalog/workflow, change, incident, problem and knowledge integration.
Required Qualifications
- 12+ years of progressive infrastructure, platform, site reliability, or infrastructure software engineering experience, including principal-level ownership of complex cross-platform outcomes in large enterprise production environments.
- Demonstrated production experience across at least four hosting-domain groups-Windows/AD; AIX/Linux; VMware/Nutanix/Hyper-V; storage/backup/SAN; middleware; networking/firewalls; observability; and ServiceNow-with deep hands-on expertise in at least two.
- Hands-on ability to design, code, test, deploy and operate production automation using Python and at least two of PowerShell, Ansible, Terraform/OpenTofu, ARM/Bicep, shell scripting or comparable technologies.
- Demonstrated experience applying Generative AI to engineering or operational workflows, including prompt engineering, context management, structured outputs, retrieval grounding, tool calling, and evaluation.
- Direct experience designing or operating agentic workflows or autonomous/semi-autonomous systems with tool integration, state/memory, human approvals, observability, bounded execution, and failure recovery.
- Experience creating evaluation harnesses, regression suites, test datasets, red-team cases, and measurable release gates for AI-enabled workflows.
- Expert ability to define platform strategy and roadmaps, establish governance, lead complex multi-team initiatives, and align senior business and technology stakeholders around priorities and measurable outcomes.
- Working mastery of Lean-Agile or SAFe principles, platform product management, business analysis, experience design, user stories, estimation, planning, dependency management, and incremental delivery.
- Demonstrated financial analysis and FinOps capability, including demand and capacity forecasting, unit-cost and consumption analysis, cost allocation, licensing and vendor trade-offs, optimization, and benefits realization.
- Strong knowledge of platform resource provisioning and configuration, CMDB and service mapping, patch and update management, incident/change/problem management, monitoring and logging, lifecycle and capacity management, documentation and reporting, security administration, backup, disaster recovery, high availability, runbook engineering, and production support.
- Ability to translate ambiguous operational problems and approved direction into reusable engineering products, measurable outcomes, implementation patterns, and clear technical standards.
- Proven ability to influence senior engineers, suppliers, technical stakeholders, and leaders and to challenge design assumptions through implementation evidence.
- Bachelor's degree in computer science, engineering, information systems, or a related field, or equivalent demonstrable experience.
- Able to work in the office a minimum of 3 days per week including every Tuesday, and Wednesday (excluding holidays & paid time off).
Preferred Qualifications
- Experience engineering and automating hybrid hosting estates spanning enterprise data centers, retail or edge compute, cloud, managed-service delivery, and multiple supplier boundaries.
- Experience with Microsoft Agent Framework, Azure AI/Foundry agent services, Semantic Kernel, LangGraph, or comparable orchestration frameworks.
- Experience integrating ServiceNow CMDB, Discovery, service mapping, catalog/workflow, change, incident and knowledge processes with source control, IaC, observability and operational automation.
- Experience with container platforms, Kubernetes, GitOps, golden paths, platform engineering, software supply-chain security, and artifact provenance.
- Familiarity with NIST AI RMF and the Generative AI Profile, CIS Benchmarks, NIST Cybersecurity Framework, zero trust, PCI-related controls, and regulated enterprise environments.
- Experience building self-service infrastructure products used by multiple engineering teams.
- Published technical writing, implementation guidance, open-source contributions, conference participation, or a sustained technical community presence.
- Relevant advanced certifications or directly equivalent demonstrated depth in Microsoft/Windows, Red Hat or IBM AIX, VMware, Nutanix, storage/SAN, networking/security, ServiceNow, observability, cloud, DevOps, Kubernetes or Terraform.
Salary Range:
$163,280 - $244
Principal Platform Engineer -Infrastructure Automation & Agentic Engineering in salisbury at Unknown Company
This position is listed as full time and able to be worked remotely.