Voltar para vagas

Machine Learning Engineer — Inference Optimization

Featherless AI

Pleno100% remotoRemoto global

About the Role

We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale. You’ll work at the intersection of research and production—turning cutting-edge models into fast, reliable, and cost-efficient systems that serve real users.

This role is ideal for someone who enjoys deep technical work, profiling systems down to the kernel/GPU level, and translating research ideas into production-grade performance gains.

What You’ll Do

-

Implement and tune techniques such as:

-

Quantization (fp16, bf16, int8, fp8)

-

KV-cache optimization & reuse

What We’re Looking For

Nice to Have

Why Join Us

Originally posted on Himalayas

Compartilhar:

Crie uma conta grátis para se candidatar

O cadastro leva menos de um minuto e libera a candidatura a esta e a todas as outras vagas do site.