MLOps Engineer job description template
An MLOps engineer builds and runs the platform that vision models are trained, deployed and monitored on. The role is closer to infrastructure and platform engineering than to modelling, and becomes essential once a team is running more models than it can manage by hand.
Before you post
Quantify the scale — inference volume, data size, number of models, GPU fleet. This audience evaluates roles on problem size more than on domain, and vague scale reads as a small problem. If your workload has genuinely interesting properties, such as high-throughput video, lead with that.
The template
About the role
We are looking for an MLOps engineer to build and operate the platform our vision models run on. You will own training and serving infrastructure, and make it fast, reliable and affordable at [scale].
What you will do
- Build and operate training and inference infrastructure
- Improve GPU utilisation, latency and serving cost
- Extend CI/CD to cover models, data and evaluation
- Build monitoring for model performance and infrastructure health
- Support engineering teams in deploying and operating their models
What we are looking for
- Strong Kubernetes and cloud infrastructure experience
- Experience operating machine learning workloads in production
- Python and infrastructure-as-code
- Monitoring and observability practice
Keep this list to three to five items. Long mandatory lists disproportionately deter the candidates you most want.
Nice to have
- GPU scheduling and optimisation
- Triton, TorchServe or equivalent serving experience
- Video and large media pipeline experience
- Cost optimisation for GPU workloads
Plain text version
Select all and paste into your ATS, then replace everything in brackets.
MLOPS ENGINEER ABOUT THE ROLE We are looking for an MLOps engineer to build and operate the platform our vision models run on. You will own training and serving infrastructure, and make it fast, reliable and affordable at [scale]. WHAT YOU WILL DO - Build and operate training and inference infrastructure - Improve GPU utilisation, latency and serving cost - Extend CI/CD to cover models, data and evaluation - Build monitoring for model performance and infrastructure health - Support engineering teams in deploying and operating their models WHAT WE ARE LOOKING FOR - Strong Kubernetes and cloud infrastructure experience - Experience operating machine learning workloads in production - Python and infrastructure-as-code - Monitoring and observability practice NICE TO HAVE - GPU scheduling and optimisation - Triton, TorchServe or equivalent serving experience - Video and large media pipeline experience - Cost optimisation for GPU workloads ABOUT US [Two or three sentences on the company, the product, and why the problem matters.] DETAILS - Location: [city / hybrid / remote] - Salary: [range] - Apply: [link]
Calibrating the level
Adjust the requirements to match the level you are actually hiring for. Asking for senior capability at a mid-level budget is the most common cause of a stalled search.
Junior
Maintains pipelines and deployment tooling within an existing platform.
Mid
Owns serving infrastructure and its reliability and cost.
Senior
Designs the ML platform and sets the standards teams build against.
Staff / Principal
Owns platform architecture and the economics of running models at scale.
Post it where the right people are
JobsInVision is a computer vision job board — your listing reaches engineers who work in this field specifically, rather than a general software audience.