Ilya Fedotov

Ilya Fedotov — AI infrastructure and MLOps

Engineering lead in AI infrastructure and MLOps. Eight years in infrastructure, four of them in MLOps. I work on cloud AI platforms, GPU infrastructure and production model inference, and I lead the teams that run them. The work sits between architecture and delivery: choose the design, ship it, operate it, and account for what the compute costs.

Remote. Open to travel. Russian native, English advanced.

// experience

2025.09 → now

Head of Engineering

GPU cloud and AI inference platform · Singularity Compute

MLOps lead until mid-2026, head of engineering since.

  • Own the engineering roadmap for the platform and the priorities that follow from it.
  • Lead inference and its optimisation: model placement across accelerators, splitting models across devices, and diagnosis down through the accelerator stack.
  • Load-test and compare inference configurations on time to first token, generation speed and throughput; evaluate the techniques that trade memory or accuracy for latency.
  • Worked on a model gateway — a compatible inference API, streaming responses and per-token accounting — and its integration with backend, authorisation and billing.
  • Hiring, planning and delivery across the engineering group.
  • Work with stakeholders: turn customer and partner needs into technical plans someone can build.
2024.10 → now

Head of MLOps

SingularityNET · in parallel

Product owner for Compute, the platform that became a company of its own.

  • Own the roadmap of a multi-tenant ML platform spanning public clouds and owned hardware.
  • Design and operate GPU infrastructure for training and high-throughput inference, with performance, reliability and cost as the standing constraints.
  • Lead cross-functional teams of backend, MLOps, DevOps and ML engineers: priorities, decomposition, technical review, hiring, delivery.
2022.12 → 2024.10

AI R&D, MLOps

SingularityNET

  • Platform engineering for serverless and microservice architectures, across conventional IT, ML, DL and blockchain projects.
  • Training and inference pipelines, feature store, continuous delivery, data warehouses and data lakes.
  • Incident response, risk analysis, project management.
2019.09 → now

Deputy head of the laboratory

Internal developer platform and ML platform on owned hardware · NaInt · in parallel

System engineer from 2019, team lead of system engineering from 2021, deputy head since 2023.

  • Lead cross-functional teams — MLOps, DevOps and ML software engineering — building the platforms that internal products and custom AI work for clients run on.
  • Define technical strategy, oversee software registration, organise knowledge transfer between teams.
  • Hiring and mentoring; rebuilding R&D and product teams; setting up outstaff technical teams.
  • Earlier: data centre architecture for machine learning workloads, on-premise training and inference, ML and DL services in production, continuous integration and delivery.
2022.01 → 2022.03

DevOps CI/CD internship, L2

EPAM Systems

  • Mid-level training programme covering the delivery lifecycle.
2018.07 → 2019.08

IT infrastructure engineer

Online Communications · St Petersburg · part-time

  • Automation for internal services.

// stack

platformsKubernetes, Docker, Helm, Linux, AWS, GCP, EKS, bare metal automationTerraform, Ansible, ArgoCD, GitOps, GitLab CI/CD, GitHub Actions inferencevLLM, Triton Inference Server, ONNX Runtime, Ray, tensor parallelism, quantisation, speculative decoding, KV-cache reuse acceleratorsNVIDIA GPUs, CUDA, NCCL, Mellanox InfiniBand agentsTool and function calling, multi-agent systems, agent harnesses, agentic RAG, LLM gateways and tracing ml toolingMLflow, DVC, feature stores observabilityPrometheus, Grafana, Loki, ELK, InfluxDB data / codePython, Bash, SQL, PostgreSQL, Redis, S3, MinIO, Ceph leadershipPlatform roadmap, technical strategy, hiring and mentoring, delivery management, SLA/SLO ownership, cost optimisation

// open source

agents Harness-engineering · deepagents_multiagent · BobaClaw An agent-first blueprint for repositories where agents plan, edit, test and evaluate — the repo as an executable work environment, not just source. Multi-agent experiments, and another Claw built to do what is needed rather than what big tech expects. inference vllm-omni-sparse-attention · central-llm-gateway Sparse attention on top of vLLM, and a fully self-hosted LLM platform with authentication, quotas, metrics and tracing. networks BibaVPN A DPI-resistant SOCKS5 and HTTP tunnel over TLS and WebSocket, for self-hosted single-VPS setups. My most-starred repository by a wide margin.   github.com/Eljaja Tool-calling datasets for Home Assistant, a Telegram-only Android VPN, and the rest.

// publication

2026.03 Good-Enough LLM Obfuscation (GELO) Anatoly Belikov, Ilya Fedotov. arXiv:2603.05035, cs.CR and cs.LG. Privacy-preserving LLM inference on untrusted accelerators.

// education

2023 MSc, Internet of Things and applied AI St Petersburg State University of Telecommunications (Bonch-Bruevich) 2021 BSc, Automated information processing and control systems St Petersburg State University of Telecommunications (Bonch-Bruevich)