B.Sc. Computer Engineering, Sharif University of Technology
Tehran, Iran
I work on inference and approximation under uncertainty: when the thing you care about cannot be computed exactly — a posterior, a value function, an expectation over an intractable distribution — you compute an approximation instead, and the question that interests me is what licenses trusting it.
Steering a language model toward a reward implicitly defines a tilted target distribution, and the samplers used in practice do not draw from it — they draw from something nearby, and the gap is rarely measured. This work samples from the tilted target exactly in distribution, so the remaining cost is compute rather than a silent bias. Manuscript in preparation.
Completed and presented an empirical implementation of a theory-only ICML 2024 result on offline RL under unobserved confounding. At full benchmark scale the confounding-blind baseline beat the corrected estimator in every cell of the sweep — identification is not estimation. I drove the implementation.
Reading in batched bandits, privacy-constrained linear bandits, and misspecified kernelized bandit optimization. No project or manuscript yet.
End-to-end empirical implementation of model-based RL for confounded POMDPs, with exact-oracle checks and a documented pessimism failure mode.
Bayesian type filtering and potential-based reward shaping for a multi-stage war of attrition with incomplete information.
Dense + sparse retrieval fused by RRF, with cross-encoder reranking.
Chambolle-Pock and ADMM written from their update rules, checked against a solver rather than delegated to one.
A backtest of mine returned several hundred percent. This is the study of whether any of the reasons I gave for it hold up. They do not.
Attention U-Net variants for segmentation, and Soft Actor-Critic built from scratch in PyTorch.
A version control system written from scratch in C — staging, commits, branches, checkout, revert, tags, and a pre-commit hook framework. No libraries.
~190M IMDb rows into PostgreSQL via a SKIP LOCKED queue, then eight queries indexed with real before/after measurements.
Head TA for Signals & Systems (Dr. Manzuri). Also Stochastic Processes and Probability & Statistics (Dr. Najafi), Machine Learning and Probability & Statistics (Dr. Sharifi-Zarchi), Linear Algebra twice (Dr. Rabiee & Dr. Ramezani), Logic Design (Dr. Hessabi & Dr. Arshadi), and Data Structures & Algorithms (Dr. Abam).
An online Physics Olympiad school, started as a small study group, for students who have no selective high school near them. I teach the mechanics, electromagnetics and laboratory tracks and write problem sets for the national selection rounds.
GPA 19.38 / 20.0, ranked 14th of 186 — the departmental average for this cohort is 17.24. Stochastic Processes and Reinforcement Learning taken as graduate courses.
7th nationwide among more than 10,000 participants, and 2nd nationwide in the experimental examination. Admitted to Sharif through the olympiad medalist track.
An IYPT-format research tournament under the World Federation of Physics Competitions. Judged team defenses of open-ended physics research problems.