~ $ whoami

Imran Nazir

AI Applied Engineer II  ·  QuillBot (Learneo)

I build the infrastructure that LLMs run on at scale.

Civil Engineering grad → self-taught → ML systems engineer. Serving 50M+ users.

the stack I run ↓

LLM Inference Stack

production topology1,240 req/s
CLIENTS
web · mobile
API consumers · agents
EDGE & SECURITY
CDN · API gateway · WAF
rate limiting · auth
LLM GATEWAY — LiteLLM
OpenAI-compatible API
provider routing · load balancing
retries · fallbacks
rate limits · budgets · cost tracking
SEMANTIC ROUTER
prompt / task classification
model selection
quality · cost · latency policy
reasoning vs fast path
frontier
FRONTIER MODELS
OpenAI · Anthropic
Gemini · Bedrock · Azure
called through LiteLLM
self-hosted
private network (vpc)
kubernetes cluster
llm-d CONTROL PLANE
request scheduling
KV-cache-aware routing
GPU-aware placement
FAST MODELS
Qwen · Llama
REASONING
DeepSeek · GPT-OSS
INFERENCE ENGINES
vLLM · SGLang · Triton
prefix cache · PagedAttention
continuous batching
speculative decoding
MODEL OPTIMIZATION
LoRA · QLoRA
AWQ · GPTQ
FP8 · INT4
GPU NODE POOLS
H100 · H200 · B200 · A100
tensor · pipeline · expert parallelism
Observabilityinstruments every tier above
metrics
Prometheus
dashboards
Grafana
logs
CloudWatch
traces
OpenTelemetry

Imran·designs & operates this stack

01.

about

I studied Civil Engineering at BIT Sindri, Dhanbad — graduating with an 8.99 CGPA in 2023. While most of my classmates went into construction, I was in my dorm room learning to code. No CS degree, no bootcamp — just curiosity and an obsession with building things that work.

I taught myself software engineering, data engineering, and then ML infrastructure from the ground up. That journey took me from a Full-Stack intern to a Data Engineer at Episource (acquired by Optum), and eventually to building production ML systems at Optum (UnitedHealth Group) for 3+ years — where I also managed networking for both dev and prod Kubernetes clusters, including firewall security, VPC configuration, and compliance policies.

Today I'm an AI Applied Engineer II at QuillBot, where I own the LLM serving infrastructure for 50M+ users. I work across the full inference stack — from model optimization (LoRA, QLoRA, quantization) to serving layers (vLLM, SGLang, Triton) to orchestration on Kubernetes. I shipped Rengoku, a production Kubernetes sidecar that handles async LLM inference at 5,000+ pages per batch.

I also write. My articles cover GPU inference internals, connection-level overload protection on Kubernetes, and capacity-aware LLM serving. I believe in sharing what I learn in the open.

50M+

users served

5K+

pages / batch

5 yrs

in ML systems

Imran Nazir
locationBangalore, India
roleAI Applied Eng. II
companyQuillBot · Learneo
focusLLM Infra · GPUs
eduCivil Eng → ML Infra
02.

work experience

QuillBot (Learneo)

current

AI Applied Engineer II

Sep 2026 – Present

Remote · Bangalore

  • ›Building and maintaining LLM serving infrastructure for 50M+ users
  • ›Owning model deployment, versioning, monitoring, and rollback pipelines
  • ›Stack: vLLM, Kubernetes, Python, MLflow

Optum (UnitedHealth Group)

MLOps and Data Platform Engineer

Jul 2023 – Sep 2026

Mumbai, India

  • ›Built Rengoku — a production Kubernetes sidecar for async LLM inference (SQS → Qwen 2.5B → S3), achieving 5,000+ pages/batch via SOLID principles and dependency injection
  • ›Led migration of model serving to AWS SageMaker Async via Terraform IaC; designed CI/CD pipelines with GitHub Actions
  • ›Architected data pipelines with Kafka, DynamoDB, and Snowflake; led MLOps automation testing to improve end-to-end model stability

Episource → Optum

Data Engineer Intern

Jan 2022 – Jun 2023

Remote

  • ›Optimised memory load in Airflow DAGs for complex multi-step pipeline orchestration
  • ›Built Docker microservice architecture with AppWrite; integrated external APIs into No-Code apps
  • ›Implemented access control using Casbin and OSO middleware

B.Tech Civil Engineering · BIT Sindri, Dhanbad

CGPA 8.99  ·  2019–2023

foundation
03.

projects

Rengoku

Work · Optum
KubernetesAWS SQSPythonDockerS3SNSSOLID

Production-grade Kubernetes sidecar acting as async middleware between AWS SQS and a self-hosted LLM inference container (Qwen 2.5B). Handles the complete async pipeline — SQS polling, inference invocation, S3 result upload, SNS/SQS notifications, and auto pod termination when the queue drains.

  • ✓5,000+ pages per batch in production
  • ✓SOLID principles + dependency injection
  • ✓Cloud-agnostic via subclassing

LLM-Powered RAG System

Personal
LangChainLangGraphFastAPIVector DBPython

Production-grade RAG system using LangChain for orchestration and a vector database for semantic retrieval. Implements LangGraph-based stateful multi-step reasoning with tool invocation. Hybrid retrieval (semantic + keyword) deployed as a FastAPI service with request batching and response caching.

  • ✓+40% response relevance over baseline
  • ✓Stateful multi-step reasoning via LangGraph
  • ✓FastAPI service with batching + caching

Medallion Lakehouse Pipeline

Personal
DatabricksPySparkDelta LakeUnity CatalogDLT

Production-style Medallion Architecture on Databricks — Bronze → Silver → Gold layers with Auto Loader for incremental ingestion, Delta Live Tables for pipeline orchestration, and Unity Catalog for fine-grained governance. Spark execution tuned with AQE, partition pruning, and predicate pushdown.

  • ✓AQE + partition pruning optimisations
  • ✓Delta Live Tables for declarative pipelines
  • ✓Unity Catalog governance layer
07.

contact

Let's talk.

Open to conversations about LLM infrastructure, MLOps architecture, or GPU systems. Happy to connect on collaborations, technical writing, or speaking opportunities.

imrannaz326@gmail.com ↗