Aditya Rawat

ML systems/Computer vision/LLM engineering/Aachen, Germany

I train segmentation models on data nobody has labelled yet, and build LLM systems that mark their own uncertainty instead of asserting through it.

Raw confocal microscopy volume.A
Raw inputWhat the microscope gives you.
Hand-annotated ground truth labels.B
Ground truthHand-annotated in napari, one label per bouton.
MicroSAM prediction on held-out volume.C
PredictionFine-tuned MicroSAM, never saw this volume.
Measureda real number from a real runShippeddeployed and usable now, in production or hostedPrototyperuns end to end on real input, still being ironed outPendingnot measured yet, and says soNot shippedbuilt, measured, left out on purpose

Selected work

Flagship & supporting projects

Segmentation models trained and benchmarked on real 3D scientific data, and LLM systems that run in production without me watching them.

Computer visionMeasured11/2025 – present · Master's thesis, RWTH Aachen

Microglomeruli Segmentation

Instance segmentation of synaptic boutons in Drosophila brain microscopy, and the tool that ships it to the lab

Raw confocal microscopy volume of the Drosophila mushroom body calyx, rotating in 3D.
Raw input
Hand-annotated ground-truth instance labels for the same volume, each bouton a distinct colour.
Ground truth
Predicted instance segmentation from the fine-tuned MicroSAM model on a held-out volume.
Prediction (held out)

Fine-tuned MicroSAM (vit_l_lm) on a held-out volume. Each colour is one predicted bouton instance.

A fully reproducible deep-learning pipeline that segments and quantifies synaptic boutons in confocal Z-stacks of the Drosophila mushroom body calyx, plus BoutonViewer: the napari desktop tool that puts the model in the hands of the biologists it was built for. Research and delivery as one project, because a model nobody can run is not a result.

  • Fine-tuned MicroSAM (vit_l_lm) with custom post-processing on a hand-annotated 3D dataset built from scratch in napari
  • Benchmarked four SOTA 3D instance-segmentation backbones head to head: Cellpose 3D, nnU-Net v2, StarDist 3D and SwinUNETR
  • Trained across three preprocessing variants (raw, difference-of-Gaussians, PSF-deconvolved) and evaluated with instance-matching metrics that account for physical voxel volume, not pixel counts
  • BoutonViewer reports per-bouton volume and surface area in µm³ / µm², auto-derives voxel calibration from acquisition type, and lets a biologist click away a false positive without touching code
  • Final checkpoints published on Hugging Face for reuse
LLM systemsShipped06/2026 – present · Solo project

TechDrishti

An agentic AI workflow for Hindi tech journalism, running unattended every day on GitHub Actions

A fully autonomous newsroom. Every morning it collects English tech news, decides for itself what is genuine news versus a job listing or SEO spam, calls out to its own research agent for anything a keyword search cannot answer, and writes an original Hindi article per story through a multi-stage LLM pipeline. Not machine translation. Zero servers, zero hosting cost.

  • Agentic multi-stage plan-execute pipeline across sarvam-30b/105b, engineered around real token-starvation and hallucination failure modes found in production, not anticipated in design
  • Cheap-model triage filters non-news before the expensive model ever sees it, which is what makes the economics work at roughly a cent per run
  • Persistent entity cache with a 45-day TTL and sense disambiguation, so the same name is never researched twice
  • Sentence-embedding clustering merges duplicate coverage of one story across 8 RSS feeds and GitHub trending
  • Every fix documented as claim, root cause, fix, verified on real articles, rather than a green test
LLM systemsShipped06/2026 – present · Solo project, spun out of TechDrishti

Deep Research Agent

An autonomous ReAct agent that plans its own searches and enforces grounding in code, not in a prompt

An autonomous ReAct agent that plans its own searches, reasons over what each result adds, and decides for itself when the evidence is enough — then returns a cited report with every comparison claim verified in code, not just promised in a prompt. It began as a free-search extension for TechDrishti and grew into a standalone, four-mode system: ask, quick grounded answer, research an article, and an automated Hindi article writer. Live and usable on Hugging Face Spaces.

  • Code-enforced iteration budget (MAX_ITERATIONS = 8): when it is exhausted the search tools are forcibly removed, but calculate survives so the report can still finish
  • A grounding check flags any comparison claim whose named target was not found in retrieved sources, because two rounds of increasingly explicit 'don't hallucinate' prompting failed reproducibly
  • AST-based calculator, structured tool errors, and evidence-tool separation, so a fabricated number cannot launder itself into looking grounded
  • Measured against TechDrishti integration and consciously left out of production. The case study does not spin that as a win
  • Free keyless retrieval layer (DuckDuckGo + Google News RSS + page/PDF fetch) instead of a metered search API

Also built

Computer vision
10/2020 – 05/2021 · Bachelor's project

Realtime Human Detection & Counting

YOLOv3-based distance enforcement during COVID-19

Led a team of 3 building a real-time human detection system to help enforce COVID-19 distancing restrictions, combining object detection with geometric distance estimation.

  • Real-time human detection using YOLOv3
  • Perspective transformation to convert image-space distances into real-world distances between individuals
Computer visionShipped
2025 – present · Thesis delivery tool

BoutonViewer

The napari desktop tool that puts the thesis segmentation model in biologists' hands

The delivery half of the thesis, packaged as its own tool: a napari app that runs the bouton-segmentation pipeline on a confocal or Airyscan stack, reports per-bouton volume and surface area in µm³ / µm², and lets a biologist delete a false positive without touching code. A step-by-step usage manual lives on the repo's GitHub wiki.

  • One-file-in, table-out workflow for non-programmers, with prediction caching so a display change never re-runs inference
  • Usage manual and model/data notes shipped alongside the tool, including where the model should not be trusted
Data & process mining
04/2024 – 07/2025 · Interdisciplinary lab course with Celonis

Cohort Discovery & Analysis Web App

Cohort discovery on event data — a proof-of-concept built for Celonis

A proof-of-concept for Celonis: an event-data application that segments event logs into behavioural cohorts by implementing a cohort-discovery research paper — data clustering over event sequences, then interactive analysis of the cohorts it finds. A light LLM assist sits on the analysis options, but the work is event-data mining, not an LLM system.

  • Full-stack app in React + Flask, tested with PyTest, for filtering, visualizing and comparing discovered cohorts
  • Bi-weekly Agile sprints, test-driven development. Grade: 1.7

Where the two tracks meet

Vision + LLM · case study

A live camera assistant that watches while your hands are busy, grown out of the paper I presented for my master's seminar. In the literature this is a training problem. Built without GPUs, a dataset, or anything to fine-tune, it has to become an architecture problem instead — which is the whole of what follows.

Vision + LLMPrototype07/2026 – present · Solo project · v2

Chitragupta

If you cannot train the decision to speak, stop asking a model to make it on every frame.

VideoLLM-online asks the harder of the two questions — not what is in this frame, but when a model watching a live stream should speak at all — and it answers it with a trained streaming head, a cached frame history, and purpose-built instruction data. Modern open-weights VLMs solve the other half outright: they read small print off a phone frame in a way a pooled CLIP vector never could.

The two halves do not compose. Without a GPU you reach those weights through a hosted endpoint, one independent call at a time — so you inherit the better eyes and lose every mechanism the paper used to earn its silence. Each of those jobs had to move somewhere cheaper, and the speak-or-stay-quiet decision moved furthest: out of the model entirely, into arithmetic over a document that costs nothing to check.

0.01 s
a question's time-to-wire, mid-tick — it was 2.00 s
0 tokens
to decide the user is owed something
1 + 1
model calls on an idle tick, and no more
18 min
of real traffic, read line by line
PHONE CAMERAa frame, on an intervalunchanged, or flat —never leaves the browserDEEPINFRAQwen3-VL-30B-A3Bnever reasons · never decidesa plain-text captionthe pixels stop here — nothing below this line has ever seen an imageTHE WORLD DOCUMENTprimary state — it survives restarts, and it is what every prompt is built fromgoal · tasks · proposed plan · expectations · find listnarrative · environment facts · recent captionsreads · writesreads onlySTAGE 1 · BOOKKEEPINGDeepSeek v4-flash15 tools · text only, never a pixelits prose is discardedTRIGGERS0 tokenspure arithmetic over the documentthe only thing that starts a sentenceScored on one thing: is the documentnow accurate? It is never asked to weighthat against whether to speak.an event — or nothingSTAGE 2 · THE SPEECH DECISION“does the user need to hearsomething?”no tools · no document · one questionpoliteness gate · 90 s[URGENT] bypasses itSPOKEN ALOUDon-device speech synthesisAn idle tick costs one vision calland one reasoning call. Stage 2 isskipped entirely when nothing happened.
Read the two arrows out of the document. One goes to a model, which reads and writes and costs a call. The other goes to arithmetic, which only reads and costs nothing — and it is the arithmetic, not the model, that decides a sentence is owed. Inverting those two is the whole difference between this and the version before it.

Skills

Languages, frameworks & tools

Languages

PythonTypeScriptJavaC++

Computer Vision

PyTorchMicroSAM / SAMnnU-Net v2SwinUNETRCellpose 3DMONAIYOLOOpenCVnapari3D instance segmentation

LLM & Agents

Agentic tool useLLM pipeline designGrounding & hallucination mitigationPrompt / token-budget engineeringsentence-transformersChainlitVision-language modelsEvals

ML & Data

TransformersFoundation modelsScikit-learnPandasMatplotlibSeaborn

Infra & Tools

GitDockerGitHub ActionsFastAPIHugging FacePostmanCelonis

Background

Experience & Education

Experience

11/2025 – 09/2026

Master's Thesis — Computer Vision Research

Software Engineering Group, RWTH Aachen University

Building a fully reproducible deep-learning pipeline to segment and quantify synaptic markers in confocal microscopy of the Drosophila mushroom body calyx — fine-tuning MicroSAM and benchmarking it against Cellpose 3D, nnU-Net v2, and SwinUNETR. Supervised by Prof. Dr. Abigail Morrison, with Prof. Dr.-Ing. Johannes Stegmaier as second examiner.

07/2024 – 05/2025

Student Assistant — Computer Vision Research

Institut für Industrieofenbau und Wärmetechnik, Aachen

Researched monocular metrology methods for measuring hot heel depth in an Electric Arc Furnace. Built Python/OpenCV tools for industrial computer vision applications and supported PhD candidates with data collection and analysis.

06/2021 – 03/2023

Quality Engineer

Larsen & Toubro Infotech, Hyderabad

Delivered an automated testing framework (Python + Selenium) from proof-of-concept to production, improving testing efficiency by 50%. API testing with Postman; trained junior engineers on testing methodology.

06/2020 – 07/2020

Intern

Ramrao Adik Institute of Technology, Navi Mumbai

Managed a team of 5 building a mobile test-taking app in Flutter/Firebase, including GPS-based authentication and UI design end to end.

Education

10/2023 – 09/2026 (expected)

M.Sc. Data Science

RWTH Aachen University

Focus: Computer Vision, Data Science, Machine Learning. Grades — Machine Learning: 1.7, Business Process Intelligence: 1.7, Advanced Process Mining: 1.7, Intro to Data Science: 2.1.

09/2017 – 05/2021

B.Sc. Computer Engineering

University of Mumbai

Final grade 8.42/10. Focus: Software Development, Software Engineering, Machine Learning.

Certifications19 verified Coursera credentials — DeepLearning.AI, Stanford, Google, Michigan

Get in touch

Open to AI / ML engineering and computer vision roles

Based in Aachen, Germany, graduating September 2026. Reach out about a role, a project, or anything on this page you want to argue with.