HPC Engineer
At Verda, we're building a full-stack AI cloud, covering everything from data centers and hardware to our own cloud platform that the world's leading AI teams use to do serious AI work.
We strive to make a positive mark on the world through the infrastructure we build and give leading teams a service they can truly depend on. Headquartered in Helsinki, we operate globally with offices in London and San Francisco. Join Verda while it’s still being built - not once it’s finished.Why Verda
Cash and equity compensation along with various fringe benefits (healthcare, lunch, wellbeing, and more).
Profitable operations with rapid, sustained growth.
30+ nationalities, with 6 different ones on the management team.
A real chance to make an impact and work alongside world class engineers, researchers, and partners across the global AI ecosystem.
Practicalities
Work mode: Remote (EU)
Level: Mid / Senior
Employment type: Full time and permanent
About the role
GPUs only deliver value once they're wired into a cluster that researchers can actually train on. As our HPC Engineer, you'll own the baremetal and virtualized clusters behind our AI cloud - from the InfiniBand fabric and shared filesystems up through the workload orchestration layer. You'll be the person who keeps large-scale GPU clusters healthy, performant, and ready for the next workload.
Your responsibilities
Administer baremetal GPU/HPC clusters end to end, from provisioning through day-two operations
Administer virtualized clusters, including the hypervisor and GPU virtualization stack underneath them
Design, deploy, and tune InfiniBand fabrics, including topology planning, subnet management, and performance validation
Deploy and operate shared/parallel filesystems supporting training and inference workloads, balancing performance, capacity, and reliability
Troubleshoot and resolve issues across the full fabric stack: fibers, transceivers, NICs, switches, drivers, and firmware
Partner with remote-hands and data center teams to diagnose hardware faults and execute physical-layer fixes and cluster expansions
Operate and tune Slurm (or equivalent) workload scheduling deployments used by customers and internal teams
Keep issue tracking, IPAM, and DCIM records accurate as clusters are built, changed, and decommissioned
Participate in on-call rotations and incident response for cluster-level issues
Collaborate with platform, network, and storage teams to integrate new clusters into the broader AI cloud
Your key competencies
Solid Linux skills, with specialization in memory management, PCIE topologies and virtualization being a bonus
Deep Infiniband knowledge, including fabric design, subnet management, and performance tuning
Solid experience with IB clustering, troubleshooting/debugging, understanding the ecosystem of fibers+transceivers+NICs+switches and all the things that could possibly fail in them
Experience about shared filesystems (e.g. Lustre, GPFS/Spectrum Scale, WekaFS, or similar)
Ability to work with remote hands teams to diagnose and resolve hardware issues remotely
Knowledge about NCCL, CUDA, DOCA and the Nvidia stack
Knowledge about Slurm and/or other workload scheduling solutions, bonus points for Slinky/slurm-bridge
Understanding the importance of keeping issue tracking/IPAM/DCIM up to date
Comfort operating production clusters where uptime and performance directly affect customer workloads
Scripting/automation ability (e.g. Python, Bash, Ansible) for repeatable cluster operations
Nice to have
RoCEv2 knowledge (and/or Spectrum-X)
Understanding of agentic guardrails, especially when applied to administration of complex systems
Ability to think beyond what is needed right now vs. some given trajectory or roadmap
Experience with GPU health-checking and diagnostics tooling (e.g. DCGM, field diagnostics)
Experience with baremetal provisioning/orchestration tooling (e.g. MAAS, Foreman, custom PXE/iPXE pipelines)
Familiarity with GPU-aware virtualization or containerization (Kubernetes device plugins, KVM/QEMU with GPU passthrough, SR-IOV)
Exposure to observability stacks (Prometheus, Grafana, Loki) for cluster-level monitoring
What's next
We're building fast and this role needs the right person behind it. There's no artificial deadline, but when we find who we're looking for, we move. If this sounds like your next move, apply now.
Please submit your application through our Careers page. We don't accept applications sent by email.
Empfohlene Jobs
Assistent für Projektleiter
Assistent für Projektleiter im Bereich Elektrotechnik gesucht Aufgaben Unterstützung in der Projektüberwachung Kundengespräche Word Excel Simaris Qualifikation Gute Laune und …
Malerhelfer / Malerhelferin - Neubau - 12-15,06 €/Stunde
Es wird ein Malerhelfer (m/w/d) für den Innenbereich gesucht! Packen Sie mit bei der Errichtung neuer Wohnungen/Wohnkomplexen an. Alles was Sie benötigen ist handwerkliches Geschick und Motivation. …
Kommissionierer (m/w/d) mit Staplerschein
Stellenbeschreibung Seit 2008 sind wir ein zuverlässiger Partner für zahlreiche Unternehmen und unterstützen unsere Kunden bei der Besetzung offener Positionen. Zurzeit suchen wir Kommissionierer (m/…
Regionalleitung (m/w/d)
ATHERA – wir suchen Hände mit Herz ATHERA verbindet die persönliche Atmosphäre starker Teams vor Ort mit der Sicherheit, Struktur und den Entwicklungsmöglichkeiten einer innovativen und wachsenden…
Pflegefachkraft - Neurologie
Strukturierte Einarbeitung Attraktive Vergütung nach TVöD Betriebliche Altersversorgung (VBL) Möglichkeit der bezuschussten Altersvorsorge CNE – zeitgemäßes, digitales Fortbildungsangebot …
Verkäuferin Backwaren (m/w/d) EDEKA Riebe 10249 Berlin-Friedrichshain
Sie arbeiten in Teilzeit bis Vollzeit in unserem Mark in 10249 Berlin, Platz der Vereinten Nationen 14. Backen Sie mit an! Mit Freude bedienen und beraten Sie unsere Kundschaft und verkaufe…
Kabelziehhelfer (m/w/d)
Vollzeit Berlin Ihre Aufgaben Unterstützung bei der Verlegung von Kabeln gemäß technischen Vorgaben und Plänen für Bahninfrastruktur Vorbereitung und Kalibrierung von Kabelschutzrohren fü…
Mitarbeiterin/Mitarbeiter Kriminaldatenverwaltung (m/w/d)
Mitarbeiterin/Mitarbeiter Kriminaldatenverwaltung (m/w/d) Wir stehen für Berlin Die Arbeit bei der Polizei Berlin sorgt nicht nur für mehr Sicherheit in der Hauptstadt - sie sichert ebenso indiv…
Metallbauer / Schlosser Tür- & Tortechnik (m/w/d)
ACT NOW! Bei ENGIE arbeiten Sie beim europaweiten Marktführer für effizienten Energieeinsatz. Mehr als 5.500 Mitarbeiter:innen an 50 Standorten in Deutschland packen beim Thema Klimaneutralität richt…
Metallbauer/Schlosser (m/w/d)
Im Rahmen der privaten Arbeitsvermittlung suchen wir noch einen Metallbauer/Schlosser/Mechaniker (m/w/d). Sie sollten eine abgeschlossene Berufsausbildung zum Metallbauer/Schlosser/Mechaniker (m/w/d)…