I’m a software and machine-learning engineer and researcher at Iowa State University, pursuing a B.S. in Data Science and an M.S. in Mechanical Engineering, and conducting research at the Translational AI Center (TrAC) with Prof. Soumik Sarkar.
My M.S. thesis builds agentic vision-language systems for cyber-agriculture; I co-first-authored SAGE, accepted at a CVPR 2026 workshop. I build end-to-end systems — from web apps and data pipelines to large-scale LLM/VLM evaluation and reproducible benchmarks — spanning agriculture, healthcare, and finance.
Three questions organise most of what I do:
Representative work is highlighted with an orange badge. Click any title for the paper. Conference tiers shown in parentheses are CORE rankings; emerging venues not yet rated by CORE are marked top-tier by reputation.
I work on agentic vision-language systems, large language models and multi-agent systems, and rigorous LLM/VLM evaluation, with applications across agriculture, healthcare, and finance. Representative work is highlighted. Click any title for the paper. Conference tiers shown in parentheses are CORE rankings; emerging venues not yet rated by CORE are marked top-tier by reputation.
Federated learning across clients running different frozen vision encoders, whose coordinate systems cannot be averaged or aligned. Each client sends only the Gram matrix of its class prototypes — a coordinate-invariant summary — and the server recovers a global quotient geometry in a single round, with no alignment maps and no representation training. Across six benchmarks under heterogeneous CNN and transformer encoders it gave the best accuracy of all methods, +9.4 points over no collaboration, at one to two orders of magnitude lower communication.
The first Mixture-of-Recursions framework for gene- and pathway-level omics learning: biological structure drives embedding smoothing, a structural attention bias, and a graph-aware router that sets each token’s recursion depth. Across eight single-cell and multi-omics benchmarks under one five-fold protocol it improved macro-F1 by 8.2 points over the strongest biology-agnostic MoR baseline while using 75% fewer parameters and up to 58% fewer FLOPs than a non-recursive Transformer.
A training-free framework that audits whether frozen pathology foundation models encode biologically meaningful signal, pairing histology with spatial transcriptomics (HEST-1k, 240 samples across breast, skin, and brain) and using TabPFN as a standardized probe. Across five encoders, UNI led pathway decodability at 0.303, but organ-level shift cost 62% of the pathway signal and section identity stayed linearly decodable up to 214× chance — while benign augmentations retained ≥95% performance.
AgriMorph treats a field robot’s physical shape as a decision made while it works, not a fixed hardware property. Separate latent models of the terrain and of the robot’s own state let a learned world model roll out candidate futures under each available body, so shapes are compared before slip, stall or crop contact happens and changed only when the predicted gain beats the transformation cost. An ensemble-disagreement confidence gate blocks shape changes made on weak sensing, and one shape-conditioned controller carries over to every body without retraining.
A walk-forward backtesting framework benchmarks classical, neural, and world-model predictors on a global equity dataset, isolating when engineered features outperform raw price inputs across model families.
An autonomous, symptom-grounded vision-language agent for crop-disease diagnosis, built on the largest plant-disease image–symptom dataset to date (335 crops, 1,251 classes, ~839K images). Stepwise diagnostic reasoning raised accuracy by +16.2 points and generalizes to unseen crops with no retraining.
Selected for poster presentation at the AI Scientist Summer Workshop hosted by Microsoft Research (Aug 4, 2026): an autonomous, symptom-grounded vision-language agent for crop-disease diagnosis, evaluated on the largest plant-disease image–symptom dataset to date with stepwise diagnostic reasoning that generalizes to unseen crops.
The first unified benchmark for pathway-guided cancer therapy-response modeling (BINN, GraphPath, PATH), over 2,622 TCGA patients across 5 cohorts; the best model reached 0.92 AUROC on prostate targeted-therapy prediction at 11% class prevalence.
A cross-year benchmark for rice-disease recognition on Bangladeshi field images: training on 2021–2025 and testing only on held-out 2026 data cost 43.91 macro-F1 points across all eight model/strategy configurations, with only 48.28% of validation performance retained. The better fine-tuning strategy proved architecture-dependent, and the most temporally robust model was not the one validation ranked first.
A multi-agent iterative-refinement system (generators DeepSeek R1 + Med-PaLM; evaluators LLaMA 3.1 + Phi-4) aligned to AMA ethics and a five-tier safety assessment. Cut ethical violations by 89% with a 92% risk-downgrade rate over 900 clinical queries.
A reinforcement-learning intervention engine that models intervention timing and selection as a sequential decision problem for syndemic health management.
A first-authored analysis interpreting transformer architectures through physics and dynamical-systems theory to explain their internal computation dynamics.
An agentic, pathway-grounded tree-of-thought inference pipeline that structures LLM reasoning over biological pathway knowledge for mechanistic therapeutic reasoning in oncology.
Benchmarked LLaMA on multi-class healthcare test-result prediction, assessing reliability for responsible medical-AI deployment.
Baseline benchmarks and reproducible metrics for protein structural-ensemble prediction in computational structural biology.
An LLM-powered tool to estimate and analyze carbon footprint from operational data.
A privacy-preserving federated-learning framework for multi-modal 2D/3D dental imaging with cross-modal prompt alignment, hierarchical optimization, and Byzantine-resilient aggregation with differential privacy.
The poster prototype of the multi-agent medical-LLM safety framework (DeepSeek R1 / Med-PaLM with LLaMA 3.1 / Phi-4 evaluators), later extended into the arXiv study above.
AgEval, a 12-task plant-stress phenotyping benchmark for zero-/few-shot in-context learning of VLMs (Claude, GPT, Gemini, LLaVA). Best-model F1 improved 46.24% → 73.37% (8-shot); strategic example selection added +15.38% F1.
A GPT-4o pipeline over 430 Shark Tank pitches quantifying how market orientation and brand storytelling drive investor decisions, with expert annotation and statistical validation for a peer-reviewed B2B marketing study.
Built with real-world industry partners.
| February 2026 | Discovery and curiosity come together in undergraduate research — LAS News, Iowa State University. Featured for undergraduate research at the Translational AI Center on explainable AI for persuasive digital content, mentored by Soumik Sarkar and Priyanka Jayashankar. |
| October 2025 | Award winners at the 22nd Annual Norman Borlaug Lectureship Poster Competition — Global Resource Systems, Iowa State University. Named among Iowa State students recognized at the competition, where I won first place in the undergraduate division. |