/ About

Hi, I'm Rohit.

I started as a full-stack developer and grew into building large systems — the kind with real money and real users behind them. Today I build AI-powered products end to end, from the first conversation about the business problem to the infrastructure that keeps them running.

Portrait of Rohit Kushwaha

01 About me

A little more about me

I’m a software engineer from Kanpur, now based in Pune, and I build products end to end. My career started in full-stack web development — Ruby on Rails and React apps where I owned everything from the database schema to the last pixel — and that habit of owning the whole thing has stayed with me through every role since.

I studied for my MCA at Kanpur Institute of Technology and graduated as the topper of the batch. I joined W3villa Technologies in February 2023 and spent three and a half years there, growing from building features to leading whole products as Tech Lead — sitting with clients, shaping roadmaps, making the architecture calls and leading the engineers who shipped it.

Along the way AI became a big part of that work. I led multi-tenant agent platforms past 120,000 interactions a month, fine-tuned models for dependable tool calling, rebuilt search around hybrid retrieval, and cut LLM costs by more than a third. That stretch earned me the company’s Rising Star award.

Today I’m a Senior AI Engineer at Academian in Pune. I work on the infrastructure that AI products stand on: distributed vector search holding over twenty billion embeddings, model serving that does three times the work in a quarter of the memory, and multilingual pipelines that turn days of work into minutes.

I like a problem best when it starts as a business question and ends as a system in production. I care about latency, cost and reliability as much as accuracy — and I’d rather measure something than argue about it.

Outside work I build in the open. OpenDynamicGGUF is an open-source optimiser that picks the right bit-width for every tensor of an LLM, and Niyam is a multi-tenant hiring platform; there’s more on GitHub and in my projects.

I write about engineering on Medium, speak at developer events — most recently Google I/O Extended Kanpur, where I talked to more than 3,000 developers — and went back to my college to speak with students about careers in AI. I also attended the Microsoft Fabric & Azure AI Foundry meetup with the Pune Developers Community, and finished in the top three nationally in Unacademy’s Nextlevel backend challenge.

If you’re building something that needs to ship, scale and make business sense, I’d be glad to hear about it — say hello or find me on LinkedIn.

02 The story

From full-stack to systems at scale

  1. Where it started

    Full-stack, from the database up

    I learned by shipping whole web products on my own in Ruby on Rails and React — owning everything from the schema to the last pixel.

  2. Real users, real stakes

    Platforms that couldn't go down

    Then came products where payments, concurrency and uptime weren’t optional. Systems I’ve worked on have served 150,000+ users and processed $3.5M+ in payments a month.

  3. 2023 – 2026 · W3villa

    Leading whole products

    As Tech Lead I owned full-stack products for enterprise clients — shaping what we built with them, making the architecture calls and leading the team that delivered it.

  4. 2026 – now · Academian

    Systems at serious scale

    Today I build the platforms AI products stand on — distributed vector search, model serving and multilingual pipelines running in production.

03 Experience

Where I've worked

  1. Senior AI Engineer at Academian Inc.

    Jun 2026 – Present

    Pune, India

    Distributed vector search, model serving and multilingual agent pipelines for enterprise products.

    • Architected a distributed Vector-as-a-Service platform on Amazon EKS managing 20B+ embeddings (HNSW, vector quantisation, EC2 Spot) — 40% lower search latency, 3× indexing throughput, 99.9% availability.
    • Built a production serving stack with Unsloth dynamic quantisation and SGLang/Triton on SageMaker — 3× inference throughput and 75% smaller memory footprint, with zero-downtime rolling updates.
    • Designed a multilingual agent pipeline (translate → evaluate → refine → human review) with translation memory and glossary-aware prompting — turnaround from days to minutes, 70% less manual review.
    • Engineered an evaluation-first platform for large RAG systems using DSPy, ragas and LLM judges with regression datasets and Prometheus/Grafana monitoring.
  2. Tech Lead at W3villa Technologies

    Feb 2023 – Jun 2026

    Noida, India

    Owned full-stack products for enterprise clients — HRMS, CRM and LMS — from architecture to delivery and the team.

    • Led full-stack product engineering across Python/FastAPI, Ruby on Rails and React/Next.js — owning architecture, delivery plans, code review and client communication, and mentoring the engineers who shipped it.
    • Led multi-tenant agent platforms (LangGraph, CrewAI, AutoGen, MCP, A2A) scaling to 120,000+ monthly interactions at 98% intent accuracy and a sub-2s SLA.
    • Built hybrid semantic search over 1M+ enterprise records — 48% better discoverability and 31% fewer support tickets; on another system, latency fell from 420ms to 273ms with 28% better relevance.
    • Fine-tuned Gemma / FunctionGemma with LoRA and QLoRA for schema-constrained tool calling — 34% more deterministic tool selection, 41% less hallucination.
    • Built vLLM / TensorRT serving on EKS and Fargate with INT8/INT4, semantic caching and intelligent routing — 3× throughput, 38% lower overall LLM spend.
    • Engineered real-time voice pipelines (Whisper.cpp → LLM → ElevenLabs over WebRTC/LiveKit) holding sub-300ms conversational latency.
    • Set up LLMOps (LangSmith, Arize Phoenix, Prometheus/Grafana) and a version-controlled prompt management system, reducing prompt-related incidents by 60%.

04 Toolbox

What I work with

AI & Generative AI
GPT-4o, Claude, Gemini, Llama 3, Gemma, DeepSeek, Mistral · LoRA, QLoRA, PEFT, RLHF, SFT · HuggingFace, PyTorch
Agentic AI
LangGraph, LangChain, CrewAI, AutoGen, Google ADK, MCP, A2A, Semantic Kernel, ReAct, tool calling
RAG & Retrieval
LlamaIndex, Qdrant, FAISS, Weaviate, pgvector, Neo4j, HNSW, BM25, hybrid search, Cohere Rerank
Inference & Serving
vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, Triton, Ray Serve, ONNX, CUDA, GGUF, INT8/INT4
Backend & Data
FastAPI, Ruby on Rails, Node.js, PostgreSQL, Redis, Elasticsearch, Celery, Kafka, Spark, Airflow, REST, GraphQL
Frontend
React, Next.js, TypeScript, Tailwind CSS, Vite
Cloud & DevOps
AWS (Bedrock, SageMaker, EKS, ECS, Lambda, Fargate), Terraform, Docker, Kubernetes, Helm, GitHub Actions, ArgoCD
Classical ML
XGBoost, LightGBM, scikit-learn, Pandas, NumPy, Isolation Forest, LSTM, drift detection
LLMOps & Observability
LangSmith, Arize Phoenix, W&B, DSPy, ragas, DeepEval, Prometheus, Grafana, OpenTelemetry
AI Security & Governance
Guardrails, OWASP LLM Top 10, prompt-injection defence, PII redaction, RBAC, SOC 2, GDPR
Languages
Python, TypeScript, JavaScript, Ruby, C++, Bash, SQL

05 Milestones

Recognition, education & community

All milestones

Master of Computer Applications (MCA) · Kanpur Institute of Technology, Kanpur · 2021 – 2023 · CGPA 8.75 / 10