Kraken logo
Kraken·

Senior AI Compute Infrastructure Engineer - Kraken

Fully remoteFull-timeSeniorUTC-8–+3EuropeLATAMAfrica+1 more#kubernetesCommission

Building the Future of Open Finance

Payward—the parent company behind Kraken—has spent 15 years building a globally accessible financial infrastructure platform. We are now building a dedicated AI Compute and Infrastructure team to power the next generation of model training, inference, and experimentation.

The Role

This team owns the infrastructure layer that enables Kraken to run AI workloads with control, speed, reliability, and cost discipline. You will work directly with AI/ML researchers, platform engineers, and security teams to build production-grade compute infrastructure.

Key Responsibilities

  • Cluster Operations: Own and operate GPU/accelerator clusters, including drivers, runtimes, kernels, and workload isolation.
  • Infrastructure Design: Build systems that enable local model execution to reduce dependency on external providers.
  • Orchestration: Improve scheduling, quota management, and utilization across heterogeneous accelerator environments.
  • Inference Optimization: Optimize pipelines for latency, throughput, and cost using stacks like vLLM, Triton, or TensorRT.
  • Observability: Build comprehensive monitoring for GPU utilization, memory pressure, token throughput, and spend.
  • Reliability: Drive incident response, alerting, and post-incident improvements for always-on infrastructure.
  • Tooling: Develop internal tooling to make GPU usage visible and accessible to non-infrastructure teams.

What You Bring

  • 5+ years of infrastructure engineering experience, specifically with GPU compute, ML infrastructure, or distributed systems.
  • Hands-on experience operating GPU clusters in production environments.
  • Strong systems engineering fundamentals: Linux, networking, storage, containers, and Kubernetes.
  • Proficiency in Python for automation and operational workflows.
  • Experience with ML serving frameworks (e.g., vLLM, Triton, TensorRT, Ray Serve).
  • Practical understanding of performance tradeoffs (batching, concurrency, memory usage, GPU utilization).
  • Ability to translate complex infrastructure tradeoffs for diverse stakeholders.

Nice to Haves

  • Experience at a frontier AI lab, hyperscaler, or high-frequency trading firm.
  • Familiarity with custom silicon (TPUs, AWS Trainium, Gaudi).
  • Experience with distributed training frameworks (DeepSpeed, Megatron-LM, FSDP).
  • Experience debugging low-level issues (CUDA, NCCL, kernel/driver).
  • Proficiency in systems languages like Rust, C++, or Go.

Timezone overlap

UTC-8–+3

Benefits

Commission

Open to

Europe · LATAM · Africa · NA

Sign in to track applications and earn points.

More roles at Kraken

Similar remote roles