Muhammad Fazeel

Data scientist and machine learning engineer, Grenoble, France

Generative AI and machine learning, built to hold up outside the benchmark.

I work across data science and deep learning: LLM agents that plan a task, write the Python for it and run it; vision models adapted with few labels; and pipelines over images, signals and sequencing data. The through line is making moderate-size models useful through careful adaptation and honest evaluation rather than through scale. Clinical data is where I have gone deepest, because it punishes anything that only works on a benchmark.

Muhammad Fazeel

01  About

What I work on

My work splits between two things that are usually kept apart: building with generative models, and making predictive models dependable. On the generative side I build agents that take a stated objective, reason about how to reach it, write and execute the Python that does the work, and then check their own result. I do the same kind of work professionally on the agentic capabilities of a large language model product used by more than 500 million people.

On the predictive side I spend my time on the parts that decide whether a model is useful in practice: adapting it with few labels, building the data and annotation workflows around it, and designing evaluation that survives contact with real cases. I have done this on MRI, whole-slide histopathology, ECG and sequencing data, and on vision-language models, where I tested where a general model stops being dependable for localisation.

02  Experience

Research and work

  1. July 2024 to present

    Python Data Scientist and Analyst

    Turing, remote, independent contractor

    Optimising the agentic capabilities of a large language model product with a user base above 500 million, and earlier, structuring large-scale datasets to establish its performance metrics.

  2. March 2026 to August 2026

    Research Intern

    GIPSA-Lab (CNRS), Inria Grenoble Rhone-Alpes and CHU Grenoble Alpes

    Master's thesis on automatic quantification of skeletal muscle in pediatric MRI. Measured how far SegmentAnyMuscle, a foundation model pretrained on adult MRI, transfers to children, and characterised its failure modes along acquisition, anatomy and pathology on a clinical cohort of nine children.

    • Benchmarked three annotation-efficient adaptation routes: parameter-efficient fine-tuning with a frozen Vision Transformer, adapters and a mixture-of-experts head; a 2-D nnU-Net trained from scratch, reaching Dice 0.94 on a held-out severe-pathology thigh; and self-supervised masked-image pretraining on 123 unlabelled series.
    • Built a human-in-the-loop workflow in which clinicians corrected predictions instead of redrawing them, producing 285 expert-reviewed slices across 30 volumes from a previously unannotated archive.
    • Tested MedGemma as a controller for pixel-level annotation, through brush-tool control and function calling, to find where a general medical VLM stops being reliable for localisation.
    • Redesigned the evaluation protocol for empty reference masks, replacing Dice with a muscle-to-total-tissue ratio and explicit false-positive accounting.

    PyTorch · nnU-Net v2 · NVIDIA H100 · GRICAD HPC

  3. Nov 2022 to Aug 2024

    AI Researcher

    Khyber Medical University, Peshawar

    Deep learning on gigapixel whole-slide histopathology to reduce subjectivity in cancer grading, under a grant co-funded by the Higher Education Commission Pakistan and the British Council UK. In parallel, GPU-accelerated sequencing pipelines with NVIDIA Clara Parabricks to align, call and annotate variants and prioritise candidate oral-cancer genes.

  4. Nov 2021 to May 2022

    Research Intern

    AI in Healthcare Lab, National Center of Artificial Intelligence, Peshawar

    Classification of atrial fibrillation, bradycardia and tachycardia from ECG, and sound source separation and denoising over spectrograms with generative adversarial networks.

03  Publications

Peer-reviewed work

  1. 2025

    Exploring the mutational spectrum of key kinase genes PIK3CA, BRAF, EGFR, ALK and ROS1 in oral squamous cell carcinoma

    F. Nawab, W. Naeem, S. Fatima, A. Ali, A. T. Khalil, A. Mehmood, M. Fazeel, H. Ahmad, M. Alorini, M. Khan, I. A. Khan, M. Irfan, S. A. Khurram

    doi:10.1186/s12885-025-14609-8
  2. 2025

    Profiling genetic mutations in the DNA damage repair genes of oral squamous cell carcinoma patients from Pakistan

    W. Naeem, F. Nawab, M. T. Sarwar, A. T. Khalil, D. A. Gaber, H. Ahmad, M. Fazeel, M. Alorini, I. A. Khan, M. Irfan, M. Khan, S. A. Khurram, A. Ali

    doi:10.1038/s41598-025-91700-x
  3. 2024

    Single-channel speech enhancement using colored spectrograms

    S. Gul, M. S. Khan, M. Fazeel

    doi:10.1016/j.csl.2024.101626

04  Projects

Selected projects

Generative AI and agents

SpectraWeaver, an image-processing agent

A Streamlit application in which Gemini 2.5 Pro turns a stated preprocessing objective into executable Python: it inspects each image, writes and runs the processing code locally, visualises the result and judges whether the objective was met. Conversations branch from any earlier message, so alternative pipelines can be compared without losing the history. Code at github.com/fazeel15/SpectraWeaver.

Digital pathology

Interpretable Broders' grading of oral cancer

Detection and classification of squamous cells in histopathology images with object-detection networks, producing case-level grades backed by evidence a pathologist can inspect rather than a single opaque score.

Physiological signals

Arrhythmia classification on PhysioNet/CinC 2020

ECG preprocessing, handling of class imbalance, and evaluation of multi-class cardiac arrhythmia models with metrics suited to the imbalance.

Genomics

Sequencing pipeline for oral-cancer gene discovery

Alignment, variant calling and annotation of raw FASTQ data with NVIDIA Clara Parabricks, then downstream analysis and visualisation to prioritise candidate genes.

05  Skills

Tools and methods

Machine learning

  • Generative AI
  • LLM agents
  • Code generation and execution
  • Deep learning
  • Computer vision
  • Vision Transformers
  • Foundation model adaptation
  • Parameter-efficient fine-tuning
  • Self-supervised pretraining
  • Medical image segmentation
  • Vision-language models
  • Agentic function calling
  • LLM optimisation
  • GANs
  • Evaluation and metric design

Frameworks and tools

  • Python
  • PyTorch
  • TensorFlow
  • Streamlit
  • Gemini API
  • NumPy and scikit-image
  • nnU-Net v2
  • NVIDIA Clara Parabricks
  • C++
  • MATLAB
  • Docker
  • Git
  • LaTeX
  • GPU and HPC clusters

Data domains

  • Clinical MRI cohorts
  • Whole-slide histopathology
  • ECG signals
  • Illumina sequencing data
  • Clinician annotation workflows

06  Contact

Get in touch

For data science and applied AI roles, research collaboration and PhD positions, or anything about the work above.

khanfazeel15@gmail.com