xelys jobs xelys jobs

Senior Platform Security Engineer (Zero Trust and Platform Security)

Submer

full-remoteseniorpermanentsecuritybackend Full remote 65 days ago via WTTJ

See how well this job matches your profile

Sign up to get an AI match score and generate a tailored application in seconds.

Get your match score

Tags

CloudStackKubernetesSlurmArgo WorkflowsGoPythonJavaGPU InfrastructureNVIDIA MIGZero Trust

About the role

Role Overview

Senior Platform Security Engineer focused on Zero Trust and Platform Security within a GPU-native cloud platform for AI training/inference and HPC workloads. You will bridge the existing production platform with a next-generation orchestration architecture, owning critical orchestration components with deep hands-on engineering.

Missions / Responsibilities

  • Design, build, and operate the compute orchestration layer for a cloud-native GPU platform.
  • Maintain and evolve CloudStack-based deployments while contributing to the design of the next-generation compute platform.
  • Collaborate with networking, storage, and platform engineers to integrate core platform primitives.

Requirements

  • Experience running GPU-heavy infrastructure for AI training, inference, or HPC workloads.
  • Proven experience with large-scale distributed compute environments (neo-cloud, hyperscaler, or HPC provider).
  • Strong experience with CloudStack internals (extending and maintaining platform functionality).
  • Production experience operating cloud orchestration platforms.
  • Experience designing/operating control-plane services for infrastructure platforms.
  • Experience maintaining or extending large Java codebases (preferably in infrastructure/platform contexts).
  • Strong programming skills in Go and Python for cloud-native platform components.
  • Familiarity with workflow orchestration systems, e.g. Argo Workflows.
  • Deep practical knowledge of Kubernetes internals and Slurm scheduling systems.
  • Experience building/operating compute orchestration layers for large-scale clusters.
  • GPU infrastructure experience including passthrough, NVIDIA MIG, scheduling, and GPU lifecycle management.
  • Understanding of GPU virtualization/passthrough (e.g., QEMU PCI passthrough, NVIDIA MIG).
  • Familiarity with virtual/distributed networking: OVS, OVN, VPC networking, RDMA/RoCE, ECMP, EVPN/VXLAN, leaf-spine fabrics.
  • Ability to solve complex implementation/operational problems across CloudStack, Kubernetes, Slurm, and workflow orchestration; improve orchestration quality via code/automation and practical design decisions.
  • Comfortable mentoring peers and improving documentation, operational workflows, and platform reliability.
  • Ability to independently own major compute-orchestration initiatives from design through rollout and operational stabilization.

Nice to Have

  • Strong platform security/Zero Trust alignment (implied by role title) applied to orchestration/control-plane components.

About Submer

Submer is a fast-growing scale-up focused on building a GPU-native cloud platform for AI and high-performance workloads. The company operates and evolves production platform components and next-generation orchestration architecture for large-scale compute environments.

Scraped 7/20/2026