HPC Engineer

Verda
Berlin

At Verda, we're building a full-stack AI cloud, covering everything from data centers and hardware to our own cloud platform that the world's leading AI teams use to do serious AI work.

We strive to make a positive mark on the world through the infrastructure we build and give leading teams a service they can truly depend on. Headquartered in Helsinki, we operate globally with offices in London and San Francisco.

Join Verda while it’s still being built - not once it’s finished.

Why Verda

  • Cash and equity compensation along with various fringe benefits (healthcare, lunch, wellbeing, and more).

  • Profitable operations with rapid, sustained growth.

  • 30+ nationalities, with 6 different ones on the management team.

  • A real chance to make an impact and work alongside world class engineers, researchers, and partners across the global AI ecosystem.

Practicalities

  • Work mode: Remote (EU)

  • Level: Mid / Senior

  • Employment type: Full time and permanent

About the role

GPUs only deliver value once they're wired into a cluster that researchers can actually train on. As our HPC Engineer, you'll own the baremetal and virtualized clusters behind our AI cloud - from the InfiniBand fabric and shared filesystems up through the workload orchestration layer. You'll be the person who keeps large-scale GPU clusters healthy, performant, and ready for the next workload.

Your responsibilities

  • Administer baremetal GPU/HPC clusters end to end, from provisioning through day-two operations

  • Administer virtualized clusters, including the hypervisor and GPU virtualization stack underneath them

  • Design, deploy, and tune InfiniBand fabrics, including topology planning, subnet management, and performance validation

  • Deploy and operate shared/parallel filesystems supporting training and inference workloads, balancing performance, capacity, and reliability

  • Troubleshoot and resolve issues across the full fabric stack: fibers, transceivers, NICs, switches, drivers, and firmware

  • Partner with remote-hands and data center teams to diagnose hardware faults and execute physical-layer fixes and cluster expansions

  • Operate and tune Slurm (or equivalent) workload scheduling deployments used by customers and internal teams

  • Keep issue tracking, IPAM, and DCIM records accurate as clusters are built, changed, and decommissioned

  • Participate in on-call rotations and incident response for cluster-level issues

  • Collaborate with platform, network, and storage teams to integrate new clusters into the broader AI cloud

Your key competencies

  • Solid Linux skills, with specialization in memory management, PCIE topologies and virtualization being a bonus

  • Deep Infiniband knowledge, including fabric design, subnet management, and performance tuning

  • Solid experience with IB clustering, troubleshooting/debugging, understanding the ecosystem of fibers+transceivers+NICs+switches and all the things that could possibly fail in them

  • Experience about shared filesystems (e.g. Lustre, GPFS/Spectrum Scale, WekaFS, or similar)

  • Ability to work with remote hands teams to diagnose and resolve hardware issues remotely

  • Knowledge about NCCL, CUDA, DOCA and the Nvidia stack

  • Knowledge about Slurm and/or other workload scheduling solutions, bonus points for Slinky/slurm-bridge

  • Understanding the importance of keeping issue tracking/IPAM/DCIM up to date

  • Comfort operating production clusters where uptime and performance directly affect customer workloads

  • Scripting/automation ability (e.g. Python, Bash, Ansible) for repeatable cluster operations

Nice to have

  • RoCEv2 knowledge (and/or Spectrum-X)

  • Understanding of agentic guardrails, especially when applied to administration of complex systems

  • Ability to think beyond what is needed right now vs. some given trajectory or roadmap

  • Experience with GPU health-checking and diagnostics tooling (e.g. DCGM, field diagnostics)

  • Experience with baremetal provisioning/orchestration tooling (e.g. MAAS, Foreman, custom PXE/iPXE pipelines)

  • Familiarity with GPU-aware virtualization or containerization (Kubernetes device plugins, KVM/QEMU with GPU passthrough, SR-IOV)

  • Exposure to observability stacks (Prometheus, Grafana, Loki) for cluster-level monitoring

What's next

We're building fast and this role needs the right person behind it. There's no artificial deadline, but when we find who we're looking for, we move. If this sounds like your next move, apply now.

Please submit your application through our Careers page. We don't accept applications sent by email.

Veröffentlicht am 2026-08-18

Empfohlene Jobs

Sozialpädagoge /in - Behindertenarbeit *

PerZukunft Arbeitsvermittlung GmbH&Co.KG
Berlin

Für unseren Auftraggeber suchen wir eine/n Sozialpädagogen/in oder Heilpädagogen/in. Die Tätigkeit wird in Berlin ausgeführt. Die wöchentliche Arbeitszeit beträgt 40 Stunden. Die Arbeitszeiten sind v…

Details Anzeigen
Veröffentlicht am 2026-08-03

Executive Assistant (m/f/d)

Alstom
Berlin

Req ID:527781    At Alstom, we build what millions depend on daily - rail transportation systems that can't afford to fail. We're more than 85,000 people worldwide creating the industry's most div…

Details Anzeigen
Veröffentlicht am 2026-08-18

Service Engineer IT Experten (m/w/d)

CANCOM
Berlin

Bei CANCOM erwartet dich ein innovatives, agiles und nachhaltiges Umfeld: Mehr als 5.300 Mitarbeiter arbeiten tagtäglich daran, mit Hilfe moderner IT-Lösungen die Zusammenarbeit und den Austausch in …

Details Anzeigen
Veröffentlicht am 2026-07-18

IT Administrator (m/w/d) im Help Desk

Medicover GmbH
Berlin

Medicover GmbH Über Medicover: Die Medicover Gruppe wurde 1995 von schwedischen Unternehmern gegründet und hat sich seitdem als internationaler Anbieter hochwertiger, integrierter Gesundheitsdien…

Details Anzeigen
Veröffentlicht am 2026-08-18

Member of Commercial Staff

nomos
Berlin

About Nomos We strive to make energy free for European households. We achieve this by developing new energy products with Europe’s leading technology companies and embed energy as a cost-saving fe…

Details Anzeigen
Veröffentlicht am 2026-08-18

(Senior) Manager Cyber Security (w/m/d)

KPMG Karriere Deutschland
Berlin

Cyber Security ist Deine Expertise? Dann sei Teil von KPMG, der Clear Choice für Cyber Security, und bring Dich hier ein: Du prüfst und berätst Organisationen im Bereich IT-Sicherheit, Security Ma…

Details Anzeigen
Veröffentlicht am 2026-08-15

Bauleiter (m/w/d)

FERCHAU GmbH
Berlin

Das ist zukünftig dein Job *Eigenverantwortliche Leitung und Überwachung von Bauprojekten *Koordination und Steuerung aller beteiligten Gewerke auf der Baustelle *Sicherstellung der termin-, kosten- u…

Details Anzeigen
Veröffentlicht am 2026-07-03

Altenpfleger (m/w/d)

Trenkwalder Deutschland
Berlin

Altenpfleger (m/w/d) Du suchst eine Aufgabe mit Sinn, bei der der Mensch im Mittelpunkt steht? Zur Verstärkung unseres Teams suchen wir zum nächstmöglichen Zeitpunkt Altenpfleger sowie Gesundheits-…

Details Anzeigen
Veröffentlicht am 2026-08-12

Helfer (m/w/d) - Verpacken - ab 19,55 EUR - Schichtarbeit

PerZukunft Arbeitsvermittlung GmbH&Co.KG
Berlin

Du [bist] bereit für eine Karriere mit Zukunft? Dann haben wir den richtigen Job für Dich! Denn wir suchen genau Dich als Verpacker (m/w/d) 19,55 € PLUS Schichtzuschläge, für ein Unternehmen aus der…

Details Anzeigen
Veröffentlicht am 2026-08-06

Proxy Platform Engineer / L3 Support (m/w/d)

Perso Plankontor
Berlin

JOBBEZEICHNUNG Wir suchen einen Senior Proxy Engineer (L3/SME) mit Gesamtverantwortung für Architektur, Sicherheit, Performance und Stabilität unserer Enterprise-Proxy-Plattform auf Basis von Squ…

Details Anzeigen
Veröffentlicht am 2026-04-11