Soheil Sayah Varg

Soheil Sayah Varg

B.Sc. Computer Engineering, Sharif University of Technology
Tehran, Iran

I work on inference and approximation under uncertainty: when the thing you care about cannot be computed exactly — a posterior, a value function, an expectation over an intractable distribution — you compute an approximation instead, and the question that interests me is what licenses trusting it.

Large Language ModelsBanditsMachine Learning ML TheoryReinforcement LearningRL Theory OptimizationProbability & StatisticsStochastic Processes

01Research

Exact controlled generation from tilted autoregressive targets

Data Science & Machine Learning Laboratory, Sharif University of Technology
December 2025 – present · Supervised by Dr. Ali Rostami (Ph.D. Computer Science, UC Irvine, 2024)

Steering a language model toward a reward implicitly defines a tilted target distribution, and the samplers used in practice do not draw from it — they draw from something nearby, and the gap is rarely measured. This work samples from the tilted target exactly in distribution, so the remaining cost is compute rather than a silent bias. Manuscript in preparation.

Model-based reinforcement learning in confounded POMDPs

Sharif University of Technology · February 2026 – August 2026 · team of three

Completed and presented an empirical implementation of a theory-only ICML 2024 result on offline RL under unobserved confounding. At full benchmark scale the confounding-blind baseline beat the corrected estimator in every cell of the sweep — identification is not estimation. I drove the implementation.

Bandit literature review

Under the guidance of Dr. Amir Najafi · May 2026 – present

Reading in batched bandits, privacy-constrained linear bandits, and misspecified kernelized bandit optimization. No project or manuscript yet.

02Selected projects

model-based-rl-confounded-pomdps

End-to-end empirical implementation of model-based RL for confounded POMDPs, with exact-oracle checks and a documented pessimism failure mode.

war-of-attrition-rl

Bayesian type filtering and potential-based reward shaping for a multi-stage war of attrition with incomplete information.

modern-information-retrieval

Dense + sparse retrieval fused by RRF, with cross-encoder reranking.

convex-optimization-algorithms

Chambolle-Pock and ADMM written from their update rules, checked against a solver rather than delegated to one.

market-predictability-study

A backtest of mine returned several hundred percent. This is the study of whether any of the reasons I gave for it hold up. They do not.

segmentation-and-deep-rl

Attention U-Net variants for segmentation, and Soft Actor-Critic built from scratch in PyTorch.

neogit

A version control system written from scratch in C — staging, commits, branches, checkout, revert, tags, and a pre-commit hook framework. No libraries.

imdb-postgresql-optimization

~190M IMDb rows into PostgreSQL via a SKIP LOCKED queue, then eight queries indexed with real before/after measurements.

03Teaching

Teaching Assistant, Sharif University of Technology

2024 – 2026 · ten course-semesters across five terms

Head TA for Signals & Systems (Dr. Manzuri). Also Stochastic Processes and Probability & Statistics (Dr. Najafi), Machine Learning and Probability & Statistics (Dr. Sharifi-Zarchi), Linear Algebra twice (Dr. Rabiee & Dr. Ramezani), Logic Design (Dr. Hessabi & Dr. Arshadi), and Data Structures & Algorithms (Dr. Abam).

Co-founder, Resonance

August 2023 – present

An online Physics Olympiad school, started as a small study group, for students who have no selective high school near them. I teach the mechanics, electromagnetics and laboratory tracks and write problem sets for the national selection rounds.

04Education and honors

B.Sc. Computer Engineering, Sharif University of Technology

September 2023 – July 2027 (expected)

GPA 19.38 / 20.0, ranked 14th of 186 — the departmental average for this cohort is 17.24. Stochastic Processes and Reinforcement Learning taken as graduate courses.

Gold Medal, Iranian National Physics Olympiad — 2023

7th nationwide among more than 10,000 participants, and 2nd nationwide in the experimental examination. Admitted to Sharif through the olympiad medalist track.

Juror, 13th Persian Young Naturalists’ Tournament — 2025

An IYPT-format research tournament under the World Federation of Physics Competitions. Judged team defenses of open-ended physics research problems.