Writing
Writeups of projects and finings, plus some miscellaneous notes. Mostly class projects from my time at Berkeley but I will try to write more about random stuff soon.
Technical posts
2026
- Offline RL: SAC+BC, IQL, and flow Q-learning on OGBench Three ways to stay close to the data while maximizing return, compared on a manipulation task and a navigation task. Most of the interesting differences are in how sensitive each one is to its one knob.
- A RAG system for the Berkeley EECS website, on a CPU with 4 GB of RAM Crawling ~15K eecs.berkeley.edu pages, writing a 138-question QA set, and building dense retrieval → cross-encoder reranking → Llama-3.1-8B. Most of the gains came from the corpus and retrieval, not the model.
- Ideal flow machines: where does a flow model's creativity come from?● Interactive I trained a flow matching UNet on MNIST, then reimplemented Kamb & Ganguli's analytic score machines (IS, LS, ELS, bbELS) as velocity fields and ran them from the same noise. The ideal flow memorizes. Locality is what lets the trained model do anything else.
- From DQN to SAC, and what the temperature is actually doing● Interactive Double DQN on CartPole, LunarLander, and MsPacman from pixels, then SAC with auto-tuned temperature and clipped double-Q on HalfCheetah and Hopper. Including the plot that contradicted what I wrote in my homework.
- Policy gradients, one variance reduction trick at a time● Interactive REINFORCE on CartPole, HalfCheetah, LunarLander, and InvertedPendulum, adding reward-to-go, advantage normalization, a learned baseline, and GAE one at a time, and watching what each one does to the learning curve.
- Facial keypoints: regression vs. transfer learning vs. heatmaps Three ways to predict 68 facial landmarks: a CNN regressing coordinates, pretrained ResNet-18 and DINOv2 backbones, and a U-Net predicting Gaussian heatmaps. Plus an honest look at why my best validation model wasn't my best test model.
- Push-T, or why an MSE policy won't commit● Interactive Behavior cloning on Push-T with an MSE action-chunking policy and a flow matching policy. The flow policy wins 0.82 to 0.67, and the gap is the conditional mean.
Notes
2026
- RL for VLAs Notes from Seohong Park's lectures in Sergey Levine's VLA seminar.
- clanker My 3D-printed SO-101 arm.
- Noise Optimization for Collapse Recovery in Trained Diffusion Models Quick notes about a cool paper.
CS 180 projects · Fall 2025
2025
- Diffusion and Flow Matching Sampling, CFG, SDEdit, inpainting, and visual anagrams with DeepFloyd IF, then training a flow-matching UNet on MNIST.
- Neural Radiance Fields Calibrating my own camera, fitting a 2D neural field, then a NeRF from scratch: ray sampling, volume rendering, novel views.
- Autostitching Photo Mosaics Homographies, projective warping, and fully automatic stitching with Harris corners, ANMS, and RANSAC.
- Filters and Frequencies Convolution, unsharp masking, hybrid images, and Laplacian-pyramid blending (the oraple).
- Coloring the Prokudin-Gorskii photo collection Aligning the three glass-plate exposures with NCC and an image pyramid.
Also: course reviews · 2024 study streams · RSS