Alle Stellen

Work with NVIDIA's DGX Cloud team to maintain high-performance DGX Cloud clusters for AI researchers and enterprise clients worldwide. DGX Cloud delivers a fully managed AI platform on major cloud providers, optimizing AI workloads using high-performance NVIDIA infrastructure.

Tasks

  • Build, implement, and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real-time monitoring, logging, and alerting.
  • Define SLOs/SLIs, monitor error budgets, and streamline reporting.
  • Support services before they launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews.
  • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.
  • Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds.
  • Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.
  • Lead triage and root-cause analysis of high-severity incidents.
  • Practice balanced incident response and blameless postmortems.
  • Participate in on-call rotation to support production services.

Requirements

  • BS in Computer Science or related technical field, or equivalent experience.
  • 10+ years of experience operating production services.
  • Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture.
  • Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet).
  • Proficiency in at least one high-level programming language (e.g., Python, Go).
  • In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards.
  • Proficient knowledge of SRE principles, encompassing SLOs, SLIs, error budgets, and incident handling.
  • Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.
  • Operating GPU-accelerated clusters with KubeVirt in production.
  • Applying generative-AI techniques to reduce operational toil.
  • Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions.
  • Experience operating and troubleshooting production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis.
Bist du Teil dieses Unternehmens?

Dieses Unternehmensprofil wurde automatisch erstellt. Wenn du für NVIDIA Switzerland AG arbeitest, kannst du das Profil jetzt beanspruchen und verifizieren – kostenlos und in wenigen Minuten.

Verifizierte Profile erhalten ein Siegel und können ihre Seite, Stellen und Bewerbungen direkt verwalten.
Über uns
NVIDIA Switzerland AG ist die Schweizer Gesellschaft des internationalen Technologieunternehmens NVIDIA. Das Unternehmen ist in der Entwicklung, Vermarktung und Bereitstellung von Technologien, Produkten und Dienstleistungen in Bereichen wie künstliche Intelligenz, beschleunigtes Computing, Grafikprozessoren, Robotik, virtuelle Realität, Informatik und professionelle Visualisierung tätig. NVIDIA entwickelt weltweit Chips, Systeme, Software und Plattformen für Rechenzentren, Forschung, Gaming, Automobilindustrie, Industrieanwendungen, Gesundheitswesen und weitere technologieintensive Branchen. Die Schweizer Gesellschaft unterstützt die Aktivitäten von NVIDIA in der Schweiz und ist im Bereich IT-Dienstleistungen registriert.
Das Team

As an NVIDIAN, you will be immersed in a diverse, supportive environment where everyone is inspired to do their best work.

Ähnliche Stellen