Mushood
Hanif

RoleSenior AI Engineer
CompanyHaga  (Jun '26 - Present)

Not what a model outputs - how the system decides, executes, and holds under load.

Optimizing: Residuals • Not: Roles
profile
MLOPS
GENAI
PyTorch
LLM
RAG
GCP
MLOPS
GENAI
PyTorch
LLM
RAG
GCP
MLOPS
GENAI
PyTorch
LLM
RAG
GCP
MLOPS
GENAI
PyTorch
LLM
RAG
GCP

about

Inference is easy. Everything else isn't.

Honest where it matters. Available when it's hard.

As a Senior AI Engineer and Founder of Haga, I specialize in building enterprise agentic systems, LLM  fine-tuning, and high-throughput inference architecture. My engineering focus centers on what surrounds generative models: multi-agent state orchestration, low-latency streaming pipelines, and GPU memory optimization. Rather than relying on simple API wrappers, I architect production AI infrastructure — from on-prem QLoRA  model tuning to sub-200ms  real-time decision systems.

Across 7 years at Afiniti and Confiz, I architected a 6-agent LangGraph document processing pipeline cutting manual turnaround by 92% (3 days to <4 hours), deployed streaming fraud detection handling 1.2M+ daily transactions at <180ms p95 latency, and scaled bilingual RAG engines using async FastAPI and FAISS vector search. Today at Haga, I am advancing physical AI by building the independent trust and physics-consistency verification layer for robot learning policies and generative world models.

Inference as a System

Most teams ship inference as a function call. The real questions - p95 latency10x load, what happens when a backend goes down - are architecture questions. I answer them before the first model goes live.

Physics-Informed Scientific ML

Data-driven physics models aren't data problems - they're structure problems. Ignoring governing equations forces the model to rediscover physics from data it may never have enough of. Embedding PDEs into the objective is what makes sparse data sufficient.

What I don't do

  • I don't ship AI wrappers dressed as products. Core API calls with a nice UI aren't systems.
  • I don't build for its own sake. The system has to earn what it costs to run.
  • I don't take off-the-shelf work. If the implementation is a Google search away, I'm not the right person.

experience & focus

Trajectory & Impact.

Building the independent trust, benchmarking, and physics-consistency verification layer for robot learning policies and generative world models.

  • Architected haga-core: Mass & friction degradation curves on MuJoCo / Robosuite environments.
  • Engineered physical violation detection (teleportation, hovering, impulse) across CogVideoX & CoTracker datasets.
  • Shipped public benchmark evidence browser (haga-web) and automated data room for investor diligence.
PythonJAXPyTorchMuJoCoNext.jsBun

impact & metrics

Proof, not promises.

stack

What I run in production.

Profiled under load. Not just imported.

PyTorch

PyTorch

Deep Learning

Primary framework for neural networks, fine-tuning & custom model pipelines.

Hugging Face

Hugging Face

LLM & NLP

Transformers, PEFT adapters, tokenizers & open-source model ecosystem.

LangGraph & LangChain

LangGraph & LangChain

Agentic Systems

Multi-agent state machines, graph-based routing & tool orchestration.

scikit-learn

scikit-learn

Machine Learning

Gradient boosting, classification, evaluation & active learning loops.

Python

Python

Core Language

Primary language for production AI infrastructure, microservices & MLOps.

TypeScript

TypeScript

Core Language

Type-safe application engineering, web interfaces & full-stack tooling.

SQL

SQL

Data Querying

Complex relational queries, analytical window functions & indexing.

Bash

Bash

Scripting

Linux shell automation, deployment scripts & cluster administration.

Git

Git

Version Control

Distributed version control, release branching & code review pipelines.

Linux

Linux

OS & Kernel

Enterprise Linux administration, process management & GPU driver setups.

FastAPI

FastAPI

Async Framework

High-performance REST & streaming endpoints handling 95+ req/s at <400ms p95.

Celery

Celery

Task Queue

Distributed asynchronous worker queues & background task execution.

Redis

Redis

In-Memory Store

Caching, pub/sub messaging, rate limiting & session state.

PostgreSQL

PostgreSQL

Relational DB

Enterprise relational database persistence & ACID-compliant transactions.

gRPC

gRPC

Inter-Service RPC

Low-latency protocol buffer RPC for high-concurrency microservices.

Bun & Next.js

Bun & Next.js

Web & JS Runtime

Ultra-fast JavaScript/TypeScript runtime & modern React server components.

FAISS & Vector Search

FAISS & Vector Search

Vector Search

HNSW dense vector indexing cutting retrieval latency from 340ms to 60ms.

Apache Airflow

Apache Airflow

Workflow Orchestration

Automated ETL pipeline DAGs, data movement & scheduled data processing.

pandas

pandas

Data Processing

Data manipulation, preprocessing, active learning & evaluation analysis.

Docker

Docker

Containerization

Containerized service packaging, multi-stage builds & environment isolation.

Kubernetes

Kubernetes

Orchestration

Production container orchestration, AKS & GKE cluster scaling.

Microsoft Azure

Microsoft Azure

Cloud Platform

Azure ML Studio, GPU streaming microservices & production AKS clusters.

Amazon Web Services

Amazon Web Services

Cloud Platform

EC2 GPU instances, SageMaker model training & cloud infrastructure.

Google Cloud (GCP)

Google Cloud (GCP)

Cloud Platform

GKE Kubernetes clusters, Vertex AI & Cloud Run serverless services.

get in touch

Hard problems welcome.

Heads-down building right now - not looking for roles. But if you've got a hard problem, a wild idea, or just want to talk shop about LLMs, distributed systems, scientific ML, or why this site is unreasonably over-engineered for a portfolio, I'm always up for that.

Logo© 2026 Mushood Hanif. All rights reserved.