Production over demos
If it can't survive real traffic, it isn't done. Everything I ship is dockerized, monitored, and built to be handed over.
Taimour Abdul Karim · Data Scientist · AI Engineer
LLM evaluation, agentic systems, and production GenAI for teams from London to California. I measure models before I trust them: red-teaming, hallucination benchmarks, judge calibration, regression gates. Currently building Bryge.io at Datality while pursuing an MSc in AI at LUMS.

Shipped for Datality / Bryge.io · Qult Technologies · BornGreat · LUMS
I'm a data scientist and AI engineer. Three years in, I built Bryge.io end to end: you point it at a Postgres database you already own, and a tool-calling agent on AWS Bedrock answers questions in plain English, returning the SQL and the chart that prove the answer. My specialty is the part most teams skip: evaluation. I red-team my own apps, audit the LLM judges everyone else trusts, measure where a model's confidence stops matching its correctness, and put regression gates in CI so prompt changes can't silently break production. MSc in AI from LUMS; 3,000+ GitHub stars.
If it can't survive real traffic, it isn't done. Everything I ship is dockerized, monitored, and built to be handed over.
LLMs earn trust by citing sources. My RAG systems constrain every answer to retrieved context. No confident hallucinations.
“+20% accuracy” means a benchmark, not a feeling. Improvements get numbers, baselines, and reproducible runs.
Research is easy. Production is the test, and these systems passed it.
01 · Flagship · 2026 · Capstone · All 8 milestones built and measured
The Closed Retraining Loop
Every LLM agent in production degrades or overspends, and the fix is always a human: someone reads the traces, notices the pattern, rewrites the prompt, swaps the model. That human is the bottleneck and the reason most agent projects stall. Nobody ships the loop that does it without them, because closing it means solving the hard half first, an objective quality signal you can promote on.
System notes
AWS Bedrock · FAISS + BM25 · FastAPI · Next.js · MLX / torch on-device · SQLite tracing
Don't take my word for it. The code is public, and 2,400+ developers starred it.
487 followers · 74 public repositories · contributing since Jul 2021
Jan 2024 – Present
London (Remote)
Fullstack AI Engineer
Aug 2023 – Dec 2024
Lahore
AI/ML Engineer
Sep 2022 – Feb 2023
California (Remote)
Data Scientist
Jun 2021 – Aug 2021
Lahore
Python Developer (Intern)
MSc. in Artificial Intelligence
Lahore University of Management Sciences (LUMS)
Sep 2024 – Jun 2026
BSc. in Data Science
National University of Computing and Emerging Sciences
Sep 2020 – Jun 2024
Open to AI/ML engineering roles, remote or hybrid
or +92 326 1127700 · usually replies within a day