About Us
Aivar Innovations is an AI-native services company and AWS Preferred Partner building governed, production-grade agentic AI systems. We partner with enterprises to deploy intelligent agents that automate complex business processes—from intelligent customer interactions to enterprise knowledge systems.
Experience: 4-8 years | Machine Learning, LLM Training, Inference & MLOps
The Role:
Kubogent is an AI/ML platform for teams that need to train, fine-tune, deploy, and operate machine learning models across Kubernetes environments.
We are looking for a Senior Machine Learning Engineer who can help turn real-world ML workflows into reliable product capabilities inside Kubogent.
This is a hands-on engineering role. You will work across the model lifecycle - from preparing data and training models to building reusable training jobs, inference workflows, experiment tracking, evaluation, and model lifecycle pipelines. You should be comfortable working directly with models and data, while also thinking about how those workflows should be exposed as repeatable, production-grade platform capabilities for other teams to
use.
You will work closely with the engineers building Kubogent's Kubernetes, accelerator, scheduling, storage, observability, and platform layers. Your job is not only to train models, but to help define how training and inference should work as a product.
What you'll do
● Build Kubogent's model training capabilities. Design and implement reusable training and fine-tuning workflows for different classes of machine learning models, including large language models.
● Build training-job abstractions. Help define how users configure datasets, models, hyperparameters, compute requirements, checkpoints, outputs, and distributed training behavior, and translate those workflows into robust Kubogent training jobs.
● Build inference workflows. Work on model loading, serving, batching, runtime configuration, evaluation, scaling behavior, and performance characteristics for different model types and inference runtimes.
● Train and fine-tune models. Perform hands-on model development, training, fine-tuning, evaluation, debugging, and iteration using modern machine learning frameworks and tooling.
● Own data preparation workflows. Work with raw datasets, understand data quality problems, clean and transform data, build preprocessing pipelines, create training/evaluation splits, and ensure datasets are reproducible and usable by downstream training jobs.
● Build MLOps pipelines. Implement repeatable workflows covering experimentation, training, evaluation, model registration, versioning, promotion, deployment, and rollback.
● Integrate experiment and model lifecycle tooling. Build integrations with tools such as MLflow for experiment tracking, metrics, artifacts, model metadata, model registries, and reproducibility.
● Evaluate models systematically. Define and implement evaluation pipelines, quality metrics, benchmark datasets, regression checks, and acceptance criteria appropriate to different model types.
● Improve training and inference efficiency. Investigate memory use, accelerator utilization, throughput, latency, batch sizing, precision, checkpointing, parallelism, and other factors that affect model performance and cost.
● Make ML workflows production-grade. Handle failure recovery, checkpointing, reproducibility, observability, artifact management, configuration, and operational concerns that are often missing from notebook-only workflows.
● Help shape the product. Work with the broader Kubogent team to decide which concepts should become APIs, workflows, defaults, templates, or reusable platform primitives rather than remaining one-off scripts.
What we're looking for
● 4-8 years of relevant engineering experience, with substantial hands-on work building and training machine learning systems.
● You have trained machine learning models yourself. You understand the practical workflow around datasets, preprocessing, experimentation, training, validation, evaluation, debugging, and iteration.
● Hands-on experience training or fine-tuning large language models. You understand tokenization, dataset preparation, training/fine-tuning workflows, checkpoints, evaluation, and the practical constraints involved in working with LLMs.
● Strong data handling skills. You are comfortable taking imperfect real-world data and performing cleaning, normalization, filtering, transformation, sampling, labeling or relabeling, dataset validation, and train/validation/test preparation.
● Strong Python skills and hands-on experience with mainstream ML frameworks such as PyTorch or TensorFlow.
● Experience with the modern LLM ecosystem, including libraries or frameworks such as Hugging Face Transformers and related model, tokenizer, dataset, and training tooling.
● Working experience with MLflow or comparable MLOps tooling. You understand experiment tracking, artifacts, parameters, metrics, model versioning, registries, and reproducible runs.
● You understand model inference beyond calling an API. You have worked with model serving or inference runtimes and understand concerns such as batching, latency, throughput, memory usage, model loading, precision, and accelerator constraints.
● You can build software around ML workflows. You are comfortable turning experimental code into maintainable libraries, services, jobs, pipelines, or platform components that other engineers can depend on.
● You understand reproducibility. You think about dataset versions, code versions, environments, configuration, seeds, checkpoints, artifacts, and the metadata required to reproduce a training run.
● You are comfortable debugging across layers. When a training or inference workload fails or performs poorly, you can reason about the model, data, framework, runtime, hardware, and surrounding system rather than treating them as isolated pieces.
● You communicate tradeoffs clearly. You can explain why a particular training method, evaluation strategy, serving runtime, precision, dataset choice, or optimization is appropriate instead of selecting tools by convention.
Strong pluses:
● Experience with distributed training technologies such as PyTorch Distributed, FSDP, DeepSpeed, Ray Train, or similar systems.
● Experience with parameter-efficient fine-tuning techniques such as LoRA or related approaches.
● Experience serving LLMs or other large models using runtimes such as vLLM, NVIDIA Triton, TensorRT-LLM, TGI, or similar systems.
● Experience working directly with GPUs or other ML accelerators and understanding memory, utilization, mixed precision, multi-device training, and performance bottlenecks.
● Experience running ML workloads on Kubernetes.
● Experience with workflow or pipeline systems used for ML workloads.
● Experience with large datasets and data-processing technologies such as Pandas, Polars, PyArrow, Spark, or similar tools.
● Experience designing model evaluation, benchmark, safety, or regression pipelines.
● Experience working on an ML platform, developer platform, or internal tooling used by other ML engineers or data scientists.
This role is a strong fit for someone who enjoys both sides of machine learning engineering: working directly
with models and data, and building the systems that make those workflows repeatable for everyone else