Inference Optimization Engineer

PolarGrid
PolarGrid

Ottawa, ON, Canada · Remote

Posted on Jul 29, 2026

About the Role

PolarGrid is building the infrastructure layer for real-time AI inference. We're looking for an Inference Optimization Engineer to squeeze every bit of performance out of our stack. You'll work directly on the systems that serve inference to our customers, making them faster, cheaper, and more efficient.

This is a deep technical, systems-focused role. You'll own latency, throughput, and cost per token as real metrics you're responsible for improving. You'll build the repeatable benchmarking and optimization process that takes new models and hardware from an initial baseline to a validated production configuration.


What You'll Do

  • Profile and optimize inference pipelines end to end using representative customer workloads, from request handling and scheduling through distributed GPU execution
  • Tune serving frameworks such as vLLM, TensorRT-LLM, and SGLang for specific latency, throughput, and cost targets
  • Build automated benchmarking and performance regression tooling across models, frameworks, precisions, hardware, and workload profiles
  • Characterize customer workloads and translate TTFT, ITL, concurrency, and context-length requirements into deployment configurations
  • Implement and evaluate quantization strategies across model families, measuring both performance gains and model-quality regressions
  • Work with hardware teams to match model configurations and parallelism strategies to GPU topology, NVLink, and interconnect bandwidth
  • Benchmark new hardware such as RTX Pro 6000s and B300s, identifying the best engine, precision, parallelism, and deployment configuration for each workload
  • Bring new model architectures into production, including checkpoint conversion, framework support, distributed configuration, and correctness validation
  • Contribute to continuous batching, speculative decoding, KV-cache optimization, prefill/decode disaggregation, and request-scheduling work
  • Read, debug, and modify inference framework internals when configuration-level tuning is not enough
  • Work with the platform team to canary performance improvements, measure them under production traffic, and turn successful configurations into repeatable deployment recipes

What We're Looking For

  • Strong GPU systems fundamentals, with the ability to work across Python, C++, CUDA, or Triton when optimization requires going below framework configuration
  • Hands-on experience with at least one major inference serving framework such as vLLM, TGI, TensorRT-LLM, or SGLang
  • ⁠Deep understanding of transformer architecture and where inference bottlenecks actually live
  • Ability to read, debug, and modify inference framework internals rather than treating them as black boxes
  • Experience building benchmarking, load-generation, or performance-regression infrastructure
  • Comfortable profiling with Nsight Systems, Nsight Compute, PyTorch Profiler, or similar tools
  • Experience with quantization and precision tradeoffs in production, including validating numerical correctness and model quality
  • Experience optimizing multi-GPU or multi-node inference across high-speed interconnects
  • Understanding of distributed inference, NCCL, GPU topology, and communication bottlenecks
  • ⁠You care about numbers: TTFT, ITL, P95/P99 latency, throughput, GPU utilization, and tokens/sec/dollar

Bonus Points

  • Experience writing custom CUDA or Triton kernels
  • Familiarity with speculative decoding, MoE routing optimizations, or prefill/decode disaggregation
  • Experience with inference request routing, scheduling, or admission control
  • ⁠Experience upstreaming performance improvements to vLLM, SGLang, TensorRT-LLM, or related projects
  • ⁠Open-source contributions to inference or ML systems projects

Why PolarGrid

You'll work on real hardware at scale, not toy benchmarks. The performance improvements you ship go directly to customers and directly affect our unit economics. Small team, real ownership.