Log in
Log in
Top pay

Machine Learning Engineer — Inference Optimization

Featherless AI · company site35w
Posted 8 months ago — may be filled

1 · Can you apply from ?

Open to
Remote (world)Exact words from the ad · found 27 Sep 2026

2 · What reaches you

Pay not listed
The ad gives no pay, so we can’t work out what reaches you. Ask the company.

3 · How you get paid

Unknown — ask the company

How often
Unknown
Ask the company
First money
Unknown
Ask the company

4 · Your working hours

Not stated

The ad doesn’t say which hours. Ask the company.

5 · Trust

Company site · found 27 Sep 2026
No one should ask you to pay to work.
Report this post

They ask for

Experience not statedNo degreePyTorchLLMs and generative AICUDAROCmTritonQuantization (fp16/bf16/int8/fp8)KV-cache optimizationSpeculative decodingModel pruningTensorRTONNX RuntimeGPU/CPU profiling

Full description

Shown as posted, in English

ABOUT THE ROLE We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale. You’ll work at the intersection of research and production—turning cutting-edge models into fast, reliable, and cost-efficient systems that serve real users. This role is ideal for someone who enjoys deep technical work, profiling systems down to the kernel/GPU level, and translating research ideas into production-grade performance gains. WHAT YOU’LL DO - Optimize inference latency, throughput, and cost for large-scale ML models in production - Profile and bottleneck GPU/CPU inference pipelines (memory, kernels, batching, IO) - Implement and tune techniques such as: - Quantization (fp16, bf16, int8, fp8) - KV-cache optimization & reuse - Speculative decoding, batching, and streaming - Model pruning or architectural simplifications for inference - Collaborate with research engineers to productionize new model architectures - Build and maintain inference-serving systems (e.g. Triton, custom runtimes, or bespoke stacks) - Benchmark performance across hardware (NVIDIA / AMD GPUs, CPUs) and cloud setups - Improve system reliability, observability, and cost efficiency under real workloads WHAT WE’RE LOOKING FOR - Strong experience in ML inference optimization or high-performance ML systems - Solid understanding of deep learning internals (attention, memory layout, compute graphs) - Hands-on experience with PyTorch (or similar) and model deployment - Familiarity with GPU performance tuning (CUDA, ROCm, Triton, or kernel-level optimizations) - Experience scaling inference for real users (not just research benchmarks) - Comfortable working in fast-moving startup environments with ownership and ambiguity NICE TO HAVE - Experience with LLM or long-context model inference - Knowledge of inference frameworks (TensorRT, ONNX Runtime, vLLM, Triton) - Experience optimizing across different hardware vendors - Open-source contributions in ML systems or inference tooling - Background in distributed systems or low-latency services WHY JOIN US - Real ownership over performance-critical systems - Direct impact on product reliability and unit economics - Close collaboration with research, infra, and product - Competitive compensation + meaningful equity at Series A - A team that cares about engineering quality, not hype

Apply on the company site
Top pay
JobFeatherless AI · company sitePosted 35w ago

Machine Learning Engineer — Inference Optimization

Posted 8 months ago — may be filled

1 · Can you apply from ?

Open to
Remote (world)Exact words from the ad · found 27 Sep 2026

4 · Your working hours

Not stated

The ad doesn’t say which hours. Ask the company.

5 · Trust

Company siteFound 27 Sep 2026
No one should ask you to pay to work.
Something wrong?Report this post

They ask for

Experience not statedNo degreePyTorchLLMs and generative AICUDAROCmTritonQuantization (fp16/bf16/int8/fp8)KV-cache optimizationSpeculative decodingModel pruningTensorRTONNX RuntimeGPU/CPU profiling

Full description

Shown as posted, in English

ABOUT THE ROLE We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale. You’ll work at the intersection of research and production—turning cutting-edge models into fast, reliable, and cost-efficient systems that serve real users. This role is ideal for someone who enjoys deep technical work, profiling systems down to the kernel/GPU level, and translating research ideas into production-grade performance gains. WHAT YOU’LL DO - Optimize inference latency, throughput, and cost for large-scale ML models in production - Profile and bottleneck GPU/CPU inference pipelines (memory, kernels, batching, IO) - Implement and tune techniques such as: - Quantization (fp16, bf16, int8, fp8) - KV-cache optimization & reuse - Speculative decoding, batching, and streaming - Model pruning or architectural simplifications for inference - Collaborate with research engineers to productionize new model architectures - Build and maintain inference-serving systems (e.g. Triton, custom runtimes, or bespoke stacks) - Benchmark performance across hardware (NVIDIA / AMD GPUs, CPUs) and cloud setups - Improve system reliability, observability, and cost efficiency under real workloads WHAT WE’RE LOOKING FOR - Strong experience in ML inference optimization or high-performance ML systems - Solid understanding of deep learning internals (attention, memory layout, compute graphs) - Hands-on experience with PyTorch (or similar) and model deployment - Familiarity with GPU performance tuning (CUDA, ROCm, Triton, or kernel-level optimizations) - Experience scaling inference for real users (not just research benchmarks) - Comfortable working in fast-moving startup environments with ownership and ambiguity NICE TO HAVE - Experience with LLM or long-context model inference - Knowledge of inference frameworks (TensorRT, ONNX Runtime, vLLM, Triton) - Experience optimizing across different hardware vendors - Open-source contributions in ML systems or inference tooling - Background in distributed systems or low-latency services WHY JOIN US - Real ownership over performance-critical systems - Direct impact on product reliability and unit economics - Close collaboration with research, infra, and product - Competitive compensation + meaningful equity at Series A - A team that cares about engineering quality, not hype