Infrastructure Software Engineer - Topology and Observability
Mountain View, CA or Remote
Joyent powers the global cloud infrastructure and developer platform providing back-end services for Samsung's billions of devices. Joyent's data center footprint is within 100ms latency to 70% of the world's population, while our multi-cloud, Kubernetes-based developer platform extends our reach to additional resource regions. We're operating at hyperscale to power workloads that bring capability and delight to Samsung's employees and customers.
Job Summary
Our infrastructure spans multiple layers — physical network, compute hosts, virtualization, managed services, and the tenant workloads running on top of them. Each layer has its own configuration management, inventory, and monitoring systems. What we lack is a tool that cuts across those layers and answers, in one view, "which hosts, network paths, and services does this workload depend on?"
In this role you will design and build a system that unifies scattered configuration, inventory, and observability data into a single infrastructure knowledge graph, and delivers multi-layer topology maps and service dependency maps on top of it. Users can overlay physical, logical, and service layers selectively, drill down hierarchically from a region to an individual node, and jump from any component to its related alerts, tickets, runbooks, and dashboards. The system serves as a shared operational view for the network, compute, SRE, and security teams, and becomes the foundation for fast root cause analysis, failure-domain analysis, and pre-change impact assessment.
Job Responsibilities
Design and maintain an infrastructure topology data model spanning physical, logical, and service layers — schemas, node and edge metadata, and cross-layer relationships
Build read-only, low-impact collectors and parsers that gather data from configuration management tools, network devices, hypervisors, and metrics/log systems
Build pipelines that reconcile declared configuration against observed state to detect and report drift
Develop the interactive web frontend and backend services providing multi-layer, hierarchical drill-down navigation, directional traffic and dependency flow visualization, and deep links into related systems
Design dependency-based blast radius calculation, failure-domain analysis, and change scenario simulation
Provide topology and dependency data as the foundation for alert correlation and automated root cause analysis (RCA), and integrate with incident response, alerting, and ticketing systems and AIOps pipelines — including human-in-the-loop review and approval workflows
Define data accuracy and freshness metrics; own the reliability and performance of the system itself
Gather requirements from consuming teams and expand adoption in stages (static artifacts → hosted service)
Skills & Competencies
Understanding of hypervisor-based virtualization and compute host operations
Understanding of the data models and limitations of observability stacks (Prometheus, OpenTelemetry, logging and tracing systems)
Sound judgment in designing data collection that minimizes impact on production systems
Ownership -Take ownership of projects, ensuring excellence in execution and accountability for results. Foster a sense of responsibility and pride in delivering high-quality work
Innovation - Drive innovation by proposing and implementing creative solutions to challenges. Stay abreast of industry trends and technologies, bringing fresh ideas to the table
Customer focus - Understand and prioritize customer needs, striving to exceed expectations in every interaction. Collaborate with cross-functional teams to ensure the delivery of customer-centric solutions
Teamwork - Embrace a collaborative and inclusive approach, working seamlessly with colleagues to achieve common goals
Education & Experience
5+ years in infrastructure, platform, SRE, or network automation, with hands-on experience across both cloud and on-premises environments
Working knowledge of L2/L3 networking: routing protocols (e.g. BGP), overlay networks (e.g. VXLAN), and the Linux networking stack
Experience building Python data pipelines, including schema definition and validation tooling
Experience modeling and processing data in analytical stores (columnar databases, graph databases, or similar)
Experience visualizing graph/topology data on the web (D3, Cytoscape, or similar) and handling the performance and readability challenges of large graphs
Preferred qualifications
Experience operating large-scale multi-region cloud or IaaS environments
Experience building or integrating CMDBs, inventories, or network sources of truth (NetBox, Nautobot, or similar)
Experience with service dependency mapping, failure-domain analysis, or what-if / digital-twin style simulation
AIOps experience: alert correlation and noise reduction, anomaly detection, automated root cause analysis, and building incident response automation pipelines
Experience applying LLM-based agents to operational workflows — particularly using LLMs to parse and summarize unstructured sources (documents, free-form configuration, alert text) wrapped in deterministic validation and human review
Experience with automated diagram generation (draw.io, Graphviz, or similar)
Experience collaborating with security teams on exposure surface or attack path visualization
What success looks like (first 12 months)
3 months: Data model and collectors for the core layers are operational, and a validated physical and logical topology exists for at least one environment
6 months: Two or more teams use the system in real incident analysis, and configuration-vs-observed drift is reported on a regular cadence
12 months: The system runs as a hosted service, and impact assessment and change what-if analysis are part of standard procedures
What this role is not
A monitoring operations role focused on dashboard maintenance
A pure frontend or pure network engineering role — we are looking for someone who combines infrastructure understanding with data and visualization skills
Compensation and Benefits
Compensation for this position will vary among specific regions due to geographical differentials in the labor market, and actual pay will be determined considering factors such as relevant skills, experience, and comparison to other employees in the role. Therefore, the annual base compensation range for this role (depending on the geographical location) is expected to be between $135000 and $190000.
Regular full-time employees (salaried or hourly) have access to benefits including Medical, Dental, Vision, Life Insurance, 401(k), Employee Purchase Program, Vacation and Sick leave, electronic reimbursement and many more. In addition, regular full-time employees (salaried or hourly) are eligible for bonus compensation based on individual, department, and company performance.
About Joyent
Joyent, a wholly-owned subsidiary of Samsung, is the open cloud company. Joyent builds technology, at the pinnacle of scale, performance, stability, and security to accelerate the transformation toward the mobile and cloud-centric world. Joyent designs, builds and manages market competitive cloud computing solutions and services for Samsung Electronics and its partners at global scale.
How To Apply
To apply, please submit a brief introduction, a copy of your resume, and a link to your Github or LinkedIn profile to jobs@joyent.com with Infrastructure Software Engineer - Topology and Observability in the subject. We are an equal-opportunity employer, building a diverse and inclusive team. Qualified applicants with criminal histories will be considered for the position in a manner consistent with the Fair Chance Ordinance.
Joyent is committed to employing a diverse workforce and providing Equal Employment Opportunities for all individuals regardless of race, color, religion, gender, age, national origin, marital status, sexual orientation, gender identity, status as a protected veteran, genetic information, status as a qualified individual with a disability, or any other characteristic protected by law.
Disclaimer: This job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee. Duties, responsibilities and activities may change or new ones may be assigned at any time with or without notice.
Vacation
Balance Work/Life with time off to truly relax and reboot.
Work Remotely
We work seamlessly together as one from our worldwide offices and offer telecommuting.
Referral Bonus
Refer someone from your network who gets hired and we'll show our appreciation through our referral bonus program.
Retirement Benefits
Let us help you plan for your future retirement with Matched 401K Contributions
Discounts
Who doesn't like a deal? Get discounts on Samsung and affiliate company products.
Health
We care about your and your family’s wellbeing. Stay healthy with our medical, dental and vision plans.
Training and Education
Grow your career with training resources and certifications
Next Generation Tech
We work, build and collaborate with next generation technologies in data, AI and compute
Open Source Tech
We use, sponsor, and collaborate extensively with open source projects