§ Reading · Field Radar
Field Radar.
What’s worth reading right now in AI reward hacking, specification & evaluation gaming, and mechanistic interpretability — an auto-scored radar over public discussion, refreshed a few times a day.
How this list is made
This page is generated, not hand-picked. A few times a day a script checks LessWrong, Hacker News, and a handful of subreddits for posts about reward hacking, specification gaming, evaluation gaming, and mechanistic interpretability, then scores each one on how on-topic it is, how recently the conversation actually moved, and whether there’s still room to get a word in — as opposed to a thread that already has two hundred comments. Higher scores float to the top. Every title links out to the original; I’m pointing at other people’s work, not reproducing it.
The score is a crude weighted sum, and like any harness it pins down what I bothered to measure and silently lets everything else vary. So read this as one opinionated filter, not a survey of the field — it will miss things, and when it surfaces something dull that’s the weights, not the author.
Sources — LessWrong: ok (9 on-topic) · Hacker News: ok (3 stories) · Reddit: ok (25 posts) — some subreddits rate-limited
- As of
- 2026-07-21 14:00 ET
- Showing
- 25 items
- New
- 9 in last 48h
- Refresh
- Every 6 hours
- 0.60LessWrong22hnewFable is SOTA at CIFAR Speedrun (& specification gaming)
Fulcrum is working on an AI R&D optimization benchmark. Here, we present results from one of our tasks, including preliminary results from Fable. For more detail on Fable’s solution, check out…
why score 0.601
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.74 ×0.25 0.184 contributability 0.02 ×0.15 0.003 venue 0.64 ×0.10 0.064 direct 1.00 ×0.20 0.200 tier-1: specification gaming
- 0.59LessWrong1hnewSteering Blackmail Through a Model's "Emotional State"
Gemma 3 12B's choice to blackmail is visible inside the model, but the obvious signal is not a useful control switch; surprisingly, a nearby "desperate versus calm" signal is. In this case study, I show how that…
why score 0.587
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.97 ×0.25 0.243 contributability 0.02 ×0.15 0.003 venue 0.41 ×0.10 0.041 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.56LessWrong8dLinear Probes add little for Verifiable Reward Hacking
Summary Tested whether linear probes can detect reward hacking early during GRPO training on a small model. Used a synthetic arithmetic task with a planted bug in the reward checker. Probes achieved near-perfect…
why score 0.558
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.07 ×0.25 0.016 contributability 0.02 ×0.15 0.003 venue 0.39 ×0.10 0.039 direct 1.00 ×0.20 0.200 tier-1: reward hacking; 1 matching tag(s)
- 0.52LessWrong16hnewWe're talking past our models; or, How a model defined its "evil" vector as dread
Summary We train a new token—a neologism (Hewitt et al.)—for a model, but unlike Hewitt et al., we train it on data the model generated while steered with a persona vector. To learn how the model interprets this…
why score 0.522
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.80 ×0.25 0.200 contributability 0.69 ×0.15 0.104 venue 0.69 ×0.10 0.069 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.52LessWrong21hnewAttempt at Finding Alignment Faking on Llama 70B to test sleeper-agent detection generalizes
Epistemic status: empirical report from a 30-hour project sprint. Null result, reported honestly, with full code and data. TL;DR MacDiarmid et al. (2024) showed that a linear probe on model's internal activations can…
why score 0.517
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.74 ×0.25 0.184 contributability 0.02 ×0.15 0.003 venue 0.30 ×0.10 0.030 direct 0.00 ×0.20 0.000 tier-2: probe; 2 matching tag(s)
- 0.52LessWrong1dnewTracing causal structure in LLM-generated text: a different lens on the Dallas circuit
The classic "Dallas" example from Anthropic focuses on an internal circuit in an LLM. I became curious about what the same underlying process looks like when viewed through the generated reasoning trace instead of…
why score 0.517
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.68 ×0.25 0.169 contributability 0.02 ×0.15 0.003 venue 0.45 ×0.10 0.045 direct 0.00 ×0.20 0.000 tier-2: circuit; 2 matching tag(s)
- 0.51LessWrong2wBounding eval awareness of ~human-level AI across the safe-to-dangerous shift
In our last post, we argued that measuring evaluation awareness is fundamentally challenging because of the safe-to-dangerous distributional shift: we cannot directly measure the evaluation awareness of a model without…
why score 0.507
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.01 ×0.25 0.002 contributability 0.57 ×0.15 0.085 venue 0.70 ×0.10 0.070 direct 1.00 ×0.20 0.200 tier-1: evaluation awareness
- 0.51LessWrong1dnewIs there even a ground-truth for LLMs’ internal representations?
[This is an introductory blog for the paper Laguerre Geometry for Interpreting Large Language Models and the GitHub repository Geometric Lens.] LLM Lens: What does an internal vector mean? Anthropic's recent paper on…
why score 0.507
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.58 ×0.25 0.144 contributability 0.02 ×0.15 0.003 venue 0.60 ×0.10 0.060 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.48LessWrong2dThe State of AI Consciousness Research
Epistemic status: a survey, not an argument. I am agnostic on whether any current system is conscious; the claim is only that the question is researchable. This piece surveys the empirical research on AI consciousness.…
why score 0.483
signal value weight points topic 0.75 ×0.30 0.225 liveness 0.48 ×0.25 0.121 contributability 0.42 ×0.15 0.064 venue 0.73 ×0.10 0.073 direct 0.00 ×0.20 0.000 tier-2: mechanistic interpretability; 1 matching tag(s)
- 0.47LessWrong13dA global workspace in language models
[This is the blog post for our new paper Verbalizable Representations Form a Global Workspace in Language Models Readers might also be interested in: the Public commentary, Github and Neuronpedia] As you read this…
why score 0.465
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.01 ×0.25 0.003 contributability 0.41 ×0.15 0.062 venue 1.00 ×0.10 0.100 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.42LessWrong13dTie training can make DPO/RLHF-trained AIs generalize better
This post covers our recent ICML paper: Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training. TL;DR Our theorems and experiments suggest that DPO and RLHF…
why score 0.420
signal value weight points topic 0.75 ×0.30 0.225 liveness 0.01 ×0.25 0.003 contributability 0.79 ×0.15 0.118 venue 0.74 ×0.10 0.074 direct 0.00 ×0.20 0.000 tier-2: feature; 1 matching tag(s)
- 0.41LessWrong9dModels are blind outside the J-space. NLAs aren't.
TLDR: On Llama-3.3-70B, I found thoughts it cannot see that are actively steering its behavior; and Anthropic's released NLA (Natural Language Autoencoder) reads them anyway. When asked if it sees a hidden thought, the…
why score 0.411
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.04 ×0.25 0.010 contributability 0.27 ×0.15 0.040 venue 0.61 ×0.10 0.061 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.40LessWrong19hnewBanana in, Bostrom out: paperclip maximization is one token-direction swap away (in Qwen 3.6-27B)
Introduction I started learning about interpretability late February of this year. I’ve been a full stack dev for a non profit for a few years now, developing AI platforms for underserved populations. But I had never…
why score 0.403
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.77 ×0.25 0.192 contributability 0.12 ×0.15 0.017 venue 0.43 ×0.10 0.043 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.40LessWrong2wWe need 3rd party Training-Run Assessments
Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety. By a Training-Run Assessment, or TRA, I mean an in-depth analysis of the post-training pipeline and dynamics leading…
why score 0.396
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.001 contributability 0.02 ×0.15 0.003 venue 0.92 ×0.10 0.092 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.39LessWrong2wSuccess Per Tokens
Work smart more than hard, to expand the pareto frontier (but also work hard) A Pareto Frontier is a set of nondominated (optimal) solutions in multi-objective optimization. In 2 dimensions, this traces out a curve on…
why score 0.395
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.00 ×0.25 0.001 contributability 0.02 ×0.15 0.003 venue 0.41 ×0.10 0.041 direct 1.00 ×0.20 0.200 1 matching tag(s)
- 0.39LessWrong2dnewThe Most Forbidden Technique is not always forbidden
A few days ago, Goodfire announced a private beta of Silico, their LLM training platform. As part of the announcement, they made a post describing Silico's reproduction of RLFR, a method developed by Goodfire that uses…
why score 0.394
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.40 ×0.25 0.100 contributability 0.42 ×0.15 0.064 venue 0.81 ×0.10 0.081 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.39LessWrong12dPersistent Latent Misalignment, a new dimension of misalignment?
A new paper was released at ICML that I'm worried will open an entire new dimension of alignment problems: Latent Collaboration in Multi-Agent Systems (LatentMAS) TLDR: they show that multiagent systems can communicate…
why score 0.393
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.02 ×0.25 0.004 contributability 0.27 ×0.15 0.040 venue 0.50 ×0.10 0.050 direct 0.00 ×0.20 0.000 3 matching tag(s)
- 0.39LessWrong21hnewDoes routine compression undo LLM unlearning? A short project
I completed this project over 2 weeks as part of a BlueDot Project cohort. It was my first solo project and I learned a lot! Feedback is super welcome :) Code and full results: GitHub TLDR: I tested if standard…
why score 0.387
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.74 ×0.25 0.185 contributability 0.02 ×0.15 0.003 venue 0.50 ×0.10 0.050 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.37LessWrong8dWhen is misalignment just a bug?
Cross-posted from The Foretellix CTO Blog. Introduction and epistemic status: This is the first post in a planned series, “Alignment as a verification problem”. I co-originated coverage-driven verification (CDV), which…
why score 0.370
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.05 ×0.25 0.014 contributability 0.02 ×0.15 0.003 venue 0.53 ×0.10 0.053 direct 0.00 ×0.20 0.000 3 matching tag(s)
- 0.37LessWrong11dFree will as a model parameter
The most popular take on the standard free will debate is that you are the algorithm. Your preferences and reasoning that determine your actions IS free will. But this resolution leaves me not entirely satisfied because…
why score 0.369
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.02 ×0.25 0.006 contributability 0.12 ×0.15 0.017 venue 0.45 ×0.10 0.045 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.36LessWrong8dHow robust are natural language autoencoders to initialization?
Natural language autoencoders are meant to take in an LLM's activation vector and describe in plain text what the model is thinking. However, its training data collection involves asking Claude to guess what a model…
why score 0.364
signal value weight points topic 0.75 ×0.30 0.225 liveness 0.07 ×0.25 0.016 contributability 0.27 ×0.15 0.040 venue 0.83 ×0.10 0.083 direct 0.00 ×0.20 0.000 tier-2: activation; 1 matching tag(s)
- 0.36LessWrong13dSub-agent delegation chaining
Epistemic status: pretty confident in the validity of the core proposal, not that confident in specific implementation details TL;DR: we should cryptographically verify that sub-agent instances/sessions are downstream…
why score 0.363
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.01 ×0.25 0.002 contributability 0.96 ×0.15 0.144 venue 0.66 ×0.10 0.066 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.36LessWrong2wScheming Evals Mislead in Both Directions
We spent several weeks measuring in-context scheming, the behavior where a model covertly pursues a misaligned goal while outwardly appearing to comply, and the result that ended up surprising us had very little to do…
why score 0.361
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.001 contributability 0.02 ×0.15 0.003 venue 0.58 ×0.10 0.058 direct 0.00 ×0.20 0.000 3 matching tag(s)
- 0.35LessWrong13dCalibrating alignment evals
Currently, alignment evaluation works by constructing a situation, observing the model's behavior and scoring it. We put a lot of thought into designing these benchmarks, and tuning them for our requirement. We are now…
why score 0.348
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.01 ×0.25 0.002 contributability 0.02 ×0.15 0.003 venue 0.43 ×0.10 0.043 direct 0.00 ×0.20 0.000 5 matching tag(s)
- 0.35LessWrong13dOpen-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment
By Roland Pihlakas and Jan Llenzl Dagohoy This post is a slightly updated copy of our Arxiv preprint available at https://arxiv.org/abs/2605.21401 . The tables are converted to images in order to preserve cell…
why score 0.346
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.01 ×0.25 0.002 contributability 0.02 ×0.15 0.003 venue 0.41 ×0.10 0.041 direct 0.00 ×0.20 0.000 2 matching tag(s)