Skip to content
JS
~/bengaluru — zsh

$

Jaspreet Singh

I build agentic LLM systems that hold up in production.

Applied AI Engineer/8+ years/Bangalore, India

Applied AI Engineer with 8+ years building production ML and large-scale distributed data systems. Ships agentic LLM applications end to end — multi-agent orchestration on LangGraph, RAG grounded in knowledge graphs, human-in-the-loop guardrails, and evaluation harnesses — on top of AWS streaming and batch pipelines processing 10M+ records daily. Research background in neural information retrieval and dense ranking.

Shipped, in production

daily alerts triaged by autonomous agents

0k+

daily alerts triaged by autonomous agents

Apple

reduction in mean time to resolution

0%

reduction in mean time to resolution

Apple

records/day through streaming pipelines

0M+

records/day through streaming pipelines

Block Scholes

QPS served by real-time analytics

0K+

QPS served by real-time analytics

TikTok Live

queries/day indexed for firmwide search

0M+

queries/day indexed for firmwide search

Goldman Sachs

Live demo — runs in your browser

An incident agent you can actually watch think

Three real-shaped incidents through the investigation graph I built at Apple: retrieve grounded evidence, propose a cause, verify it against telemetry, and stop at a human gate before touching anything. One of the three is wrong on purpose.

SEV-2

p99 latency on checkout-api crossed 2.4s (threshold 800ms) for 5m

service: checkout-api

0/25 steps

The clean path: retrieve, hypothesise, verify against live telemetry, propose, get approval.

refutedapproverejectingest_alertPENDINGtriagePENDINGretrievePENDINGhypothesizePENDINGverifyPENDINGdraft_rcaPENDINGhuman_gatePENDINGremediatePENDINGescalatePENDING

Press play to run the investigation.

Human-in-the-loop gate

Every consequential action passes through here. Accept and reject decisions are the evaluation set.

Agent precision

no decisions yet

Runs entirely in your browser — no model call, no backend, no network. The traces are authored to mirror the shape of the real system; service names and numbers are invented. space play/pause · step · 1–3 scenario.

Selected work

Systems, and what they cost to get right

Every number below is on the résumé. These pages are the part the résumé has no room for: the constraint, the trade-off, and the thing I would do differently.

Toolkit

What I reach for

Generative & Applied AI

LLMsRetrieval-Augmented Generation (RAG)GraphRAGAgentic WorkflowsMulti-Agent OrchestrationLangGraphPrompt Engineering & OptimizationKnowledge GraphsLong-Term Agent MemoryHuman-in-the-Loop (HITL) ReviewGuardrailsEvaluation FrameworksTransformers/BERTPyTorch

Languages

PythonJavaGoC++SQLJavaScript

Data & Cloud

AWS (Kinesis, S3, Lambda, EC2)KafkaPySparkParquetRedisElasticsearchKibanaDockerKubernetesTerraformAnsibleSpring BootLinux

AI-Assisted Development

Claude CodeCursorGitHub Copilot

Notes

Things worth writing down

Currently

Building applied-AI systems at Apple, and open to Applied AI and AI Engineer roles where the model is the easy part.

If you are hiring for agentic systems, retrieval, or the data infrastructure underneath them — quantitative finance included — I would like to hear about it.

LinkedInRésumé (PDF)