Abraham Yeung

Abraham Yeung

Math & CS at Stanford. Reinforcement learning, post-training, and AI safety.

status: mid post-training  ·  reward signal: curiosity  ·  KL penalty: sleep

📈 training history📊 evals🤖 ask

🤖policy

About me

I'm a junior at Stanford studying Mathematics and Computer Science. I work on reinforcement learning and post-training, and on what those methods do to a model's safety: whether its reasoning stays honest, and which training choices quietly trade oversight for capability. This fall I'm joining Redwood Research part-time as an AI safety research intern, after a summer on the Data Platform team at Databricks.

I grew up in Hong Kong and went to Eton College. I sing baritone with the Mendicants, Stanford's oldest a cappella group. I speak English, Cantonese, and Mandarin natively, plus some German and beginner Japanese.

🥕reward model

What I optimize for

∇gradient updates

Research

Three papers are accepted at NeurIPS 2026 workshops: two at InterpScience and one at MATH-AI. Titles, summaries, and PDFs are on the papers page.

What Happens to a Monitor's Accuracy When You Train Against It

Abraham Yeung, Anagha Ramaswamy · InterpScience workshop, NeurIPS 2026 · PDF

A linear probe on a model's hidden states predicts whether its next answer will be correct, with held-out AUROC 0.982. Used as the reward for reinforcement learning, a probe of the same representation drives true accuracy down by 31 percentage points, and the probe's own accuracy does not register the change for about 40 training steps. One statistic, the share of answers above a threshold fixed at the start, registers it early and without ground truth.

How Small Can We Go? Calibrating Activation-Cache Compression for SAE Training

Compressing the activation cache used to train a sparse autoencoder changes which features it learns, but so does perturbing one value in a million by the smallest fp32 step. The paper calibrates each site between retraining on identical bytes and a reseed, then measures codecs against that: naive int4 loses more features than a reseed, and standardizing channels first brings int4 back above the floor at the same file size.

Coding Agents for Coding Theory

Five weeks of an LLM coding agent on open cells of the quaternary edit-metric code table, under a protocol that required verifiers to be written and tested before any search. Searching over codes invariant under a prescribed symmetry group improved thirteen published lower bounds, including E4(6,3) from 114, a bound standing since 2012, to 120. The paper gives equal space to how the work went wrong.

📝rollouts

Projects

Research Frontier Mineragents · in progress

A resumable agent that mines recent ML papers, walks the citation graph to find proposed extensions nobody has attempted, and ranks the open problems. SQLite plus vector search, content-addressed caching, best-first frontier.

Prediction Markets Agentmarkets · top 5 of 100+

Agent that evaluates live Polymarket prices and flags mispriced contracts from incoming information and microstructure signals. Finalist at the NVIDIA, Vercel, and Brex hackathon.

Multi-Agent Causal Modelingagents · Bridgewater AI Hackathon

Selected as 1 of 24. A multi-agent system that maps causal graphs for policy-impact scenarios by decomposing macro questions into testable sub-claims.

Maestroai / music · TreeHacks runner-up

Multimodal music coach that analyzes instrument audio and video with NVIDIA vision and audio models. Devpost

gigabpesystems · Rust · PyPI

BPE tokenizer trainer in Rust with Python bindings. A 32k vocabulary on 12.9 GB of FineWeb in 38 s against 257 s for HuggingFace tokenizers, with every merge byte-identical and a fraction of the memory. GitHub

Transformer from scratchsystems · CS 336

A language model built end to end in PyTorch: tokenizer, training loop, FlashAttention as Triton kernels, distributed data parallel with optimizer sharding, profiling, scaling laws, and SFT plus GRPO post-training.

AQI Forecastingtime series · CS 229

LSTM, GNN, and CNN models over 545k+ observations for spatiotemporal air quality forecasting, with robustness testing under distribution shift.

📈training history

Experience

📊evals

Honors

🤖inference

Ask my post-trained self

A small open-weights model with a short bio in its context. Ask it about my work. It can be wrong, so trust the sections above over it.