Research
I work on natural language processing and large language models, with a focus on how they are trained, fine-tuned, evaluated, and benchmarked. My broader interests include multi-agent systems, alignment and hallucination, and AI4Science. Across projects, I build benchmarks and evaluation pipelines, design multi-agent LLM systems, and test methods that make models more reliable and aligned.
Current project
Socio-linguistic modeling to understand the long-term dynamics of news engagement in online media
The project combines natural language processing, large language models, and causal inference to understand how people change their political views and how they approach social interactions on platforms such as Reddit.
So far, I have collected roughly 30,000 public comments and built the preprocessing pipeline, including PII scrubbing and filtering. I also built an LLM annotation system with typed schemas that benchmarks state-of-the-art models against human-coded gold standards on agreement, precision, and recall.
Benchmarking and evaluation
Benchmarks shape what we believe a model can do. I build benchmarks and evaluation pipelines, often together with domain experts, to measure where frontier models succeed and where they fail.
Artificial Intelligence and Law, 2026
IslamicLegalBench
A benchmark of how well LLMs know and reason about Islamic law. It spans 14 task types, from factual recall to advanced jurisprudential reasoning, and draws on legal texts covering more than 1,200 years and several schools of thought. We evaluated GPT-5, Gemini 2.5, Claude 4, DeepSeek R1, Grok 4, and Llama 4 with LLM-as-a-judge scoring, expert human annotation, and zero-shot, few-shot, and chain-of-thought prompting.
NeurIPS 2025 Workshop
Can LLMs Write Faithfully?
An agent-based framework for evaluating religious content written by LLMs. A quantitative agent scores individual essays and a qualitative agent compares responses from several models; both check citations with search tools and return structured analyses for human evaluators. Testing frontier and domain-specific models showed clear gaps in citation accuracy and in faith-sensitive writing.
Alignment and bias
I study how to measure bias in frontier models and how to align their answers so they are more inclusive and reliable.
JAIR, 2026
WorldView-Bench
A benchmark of 175 questions in 7 categories for evaluating global cultural perspectives in LLMs. In our multiplex approach, perspective-specific agents answer in parallel and a multiplex agent combines their views. This raised inclusivity from 13% to 94% and reduced negative sentiment to 2.4%.
IEEE EDUCON 2025
Auditing frontier LLMs for cultural bias
An audit of cultural bias in GPT-4, Claude 3.5, Llama 3.1 and 3.2, and Mistral 7B on educational topics. Bias-aware prompt calibration, controlled generation, and multi-agent systems raised inclusivity rates from 3.25% to 98%.
Multi-agent systems
I build multi-agent LLM systems in which coordinator and specialist agents split complex judgments into smaller checks, and I compare their verdicts with those of human experts.
IEEE EDUCON 2025
MASC, a multi-agent co-pilot for senior design projects
A co-pilot that helps assess engineering senior design projects. Eight specialist agents evaluate problem formulation, system complexity, ethics, risk, and methodology, alongside NLP measures such as lexical cohesion and clause density. Their combined assessment matched expert faculty evaluations 89% of the time.
Preprint, 2025
SLR-GPT: can agents judge systematic reviews like humans?
Twenty-seven LLM agents, organized into PRISMA-inspired deliberative societies, evaluate a systematic literature review from a single PDF. The system combines few-shot prompting, tool use, vision-language models, and arXiv access, and agreed with domain experts 84% of the time on reviews in medicine, AI, and AR/VR. This is joint work with Weill Cornell Medicine and Qatar University.
Earlier work
Before focusing on LLMs, I worked on human-computer interaction, language technology for a low-resource language, and 3D computer vision.
Immersive learning with VR
At Virtuality Labs, ITU, I led a team of junior researchers in designing and running Pakistan’s first VR classroom field experiment, comparing VR, face-to-face, and Zoom teaching with more than 100 first-year CS students. The study was published at ECIS 2025.
Urdu NLP for mental health
At the CSaLT Lab, LUMS, with Imperial College London, I fine-tuned and integrated Whisper, BERT, MuRIL, and GPT-2 on a custom Urdu dataset to build an Urdu-language psychotherapy chatbot.
3D reconstruction
At the Intelligent Machines Lab, ITU, in a joint project with ETRI (South Korea), I built a synthetic 4D dataset of more than 80 indoor scenes (about 2 TB) in Unity3D and benchmarked SfM, COLMAP, NeRF, and 3D Gaussian Splatting.
All papers, with links to code and data, are on the Publications page, and my full record is in my CV.