Field Radar.

What’s worth reading right now in AI reward hacking, specification & evaluation gaming, and mechanistic interpretability — an auto-scored radar over public discussion, refreshed a few times a day.

How this list is made

This page is generated, not hand-picked. A few times a day a script checks LessWrong, Hacker News, and a handful of subreddits for posts about reward hacking, specification gaming, evaluation gaming, and mechanistic interpretability, then scores each one on how on-topic it is, how recently the conversation actually moved, and whether there’s still room to get a word in — as opposed to a thread that already has two hundred comments. Higher scores float to the top. Every title links out to the original; I’m pointing at other people’s work, not reproducing it.

The score is a crude weighted sum, and like any harness it pins down what I bothered to measure and silently lets everything else vary. So read this as one opinionated filter, not a survey of the field — it will miss things, and when it surfaces something dull that’s the weights, not the author.

Sources — LessWrong: ok (9 on-topic) · Hacker News: ok (3 stories) · Reddit: ok (25 posts) — some subreddits rate-limited

As of
2026-07-21 14:00 ET
Showing
25 items
New
9 in last 48h
Refresh
Every 6 hours
  1. 0.60LessWrong22hnew
    Fable is SOTA at CIFAR Speedrun (& specification gaming)

    Fulcrum is working on an AI R&D optimization benchmark. Here, we present results from one of our tasks, including preliminary results from Fable. For more detail on Fable’s solution, check out…

    why score 0.601
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.74×0.250.184
    contributability0.02×0.150.003
    venue0.64×0.100.064
    direct1.00×0.200.200

    tier-1: specification gaming

  2. 0.59LessWrong1hnew
    Steering Blackmail Through a Model's "Emotional State"

    Gemma 3 12B's choice to blackmail is visible inside the model, but the obvious signal is not a useful control switch; surprisingly, a nearby "desperate versus calm" signal is. In this case study, I show how that…

    why score 0.587
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.97×0.250.243
    contributability0.02×0.150.003
    venue0.41×0.100.041
    direct0.00×0.200.000

    2 matching tag(s)

  3. 0.56LessWrong8d
    Linear Probes add little for Verifiable Reward Hacking

    Summary Tested whether linear probes can detect reward hacking early during GRPO training on a small model. Used a synthetic arithmetic task with a planted bug in the reward checker. Probes achieved near-perfect…

    why score 0.558
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.07×0.250.016
    contributability0.02×0.150.003
    venue0.39×0.100.039
    direct1.00×0.200.200

    tier-1: reward hacking; 1 matching tag(s)

  4. 0.52LessWrong16hnew
    We're talking past our models; or, How a model defined its "evil" vector as dread

    Summary We train a new token—a neologism (Hewitt et al.)—for a model, but unlike Hewitt et al., we train it on data the model generated while steered with a persona vector. To learn how the model interprets this…

    why score 0.522
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.80×0.250.200
    contributability0.69×0.150.104
    venue0.69×0.100.069
    direct0.00×0.200.000

    1 matching tag(s)

  5. 0.52LessWrong21hnew
    Attempt at Finding Alignment Faking on Llama 70B to test sleeper-agent detection generalizes

    Epistemic status: empirical report from a 30-hour project sprint. Null result, reported honestly, with full code and data. TL;DR MacDiarmid et al. (2024) showed that a linear probe on model's internal activations can…

    why score 0.517
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.74×0.250.184
    contributability0.02×0.150.003
    venue0.30×0.100.030
    direct0.00×0.200.000

    tier-2: probe; 2 matching tag(s)

  6. 0.52LessWrong1dnew
    Tracing causal structure in LLM-generated text: a different lens on the Dallas circuit

    The classic "Dallas" example from Anthropic focuses on an internal circuit in an LLM. I became curious about what the same underlying process looks like when viewed through the generated reasoning trace instead of…

    why score 0.517
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.68×0.250.169
    contributability0.02×0.150.003
    venue0.45×0.100.045
    direct0.00×0.200.000

    tier-2: circuit; 2 matching tag(s)

  7. 0.51LessWrong2w
    Bounding eval awareness of ~human-level AI across the safe-to-dangerous shift

    In our last post, we argued that measuring evaluation awareness is fundamentally challenging because of the safe-to-dangerous distributional shift: we cannot directly measure the evaluation awareness of a model without…

    why score 0.507
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.01×0.250.002
    contributability0.57×0.150.085
    venue0.70×0.100.070
    direct1.00×0.200.200

    tier-1: evaluation awareness

  8. 0.51LessWrong1dnew
    Is there even a ground-truth for LLMs’ internal representations?

    [This is an introductory blog for the paper Laguerre Geometry for Interpreting Large Language Models and the GitHub repository Geometric Lens.] LLM Lens: What does an internal vector mean? Anthropic's recent paper on…

    why score 0.507
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.58×0.250.144
    contributability0.02×0.150.003
    venue0.60×0.100.060
    direct0.00×0.200.000

    2 matching tag(s)

  9. 0.48LessWrong2d
    The State of AI Consciousness Research

    Epistemic status: a survey, not an argument. I am agnostic on whether any current system is conscious; the claim is only that the question is researchable. This piece surveys the empirical research on AI consciousness.…

    why score 0.483
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.48×0.250.121
    contributability0.42×0.150.064
    venue0.73×0.100.073
    direct0.00×0.200.000

    tier-2: mechanistic interpretability; 1 matching tag(s)

  10. 0.47LessWrong13d
    A global workspace in language models

    [This is the blog post for our new paper Verbalizable Representations Form a Global Workspace in Language Models Readers might also be interested in: the Public commentary, Github and Neuronpedia] As you read this…

    why score 0.465
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.01×0.250.003
    contributability0.41×0.150.062
    venue1.00×0.100.100
    direct0.00×0.200.000

    2 matching tag(s)

  11. 0.42LessWrong13d
    Tie training can make DPO/RLHF-trained AIs generalize better

    This post covers our recent ICML paper: Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training. TL;DR Our theorems and experiments suggest that DPO and RLHF…

    why score 0.420
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.01×0.250.003
    contributability0.79×0.150.118
    venue0.74×0.100.074
    direct0.00×0.200.000

    tier-2: feature; 1 matching tag(s)

  12. 0.41LessWrong9d
    Models are blind outside the J-space. NLAs aren't.

    TLDR: On Llama-3.3-70B, I found thoughts it cannot see that are actively steering its behavior; and Anthropic's released NLA (Natural Language Autoencoder) reads them anyway. When asked if it sees a hidden thought, the…

    why score 0.411
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.04×0.250.010
    contributability0.27×0.150.040
    venue0.61×0.100.061
    direct0.00×0.200.000

    2 matching tag(s)

  13. 0.40LessWrong19hnew
    Banana in, Bostrom out: paperclip maximization is one token-direction swap away (in Qwen 3.6-27B)

    Introduction I started learning about interpretability late February of this year. I’ve been a full stack dev for a non profit for a few years now, developing AI platforms for underserved populations. But I had never…

    why score 0.403
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.77×0.250.192
    contributability0.12×0.150.017
    venue0.43×0.100.043
    direct0.00×0.200.000

    1 matching tag(s)

  14. 0.40LessWrong2w
    We need 3rd party Training-Run Assessments

    Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety. By a Training-Run Assessment, or TRA, I mean an in-depth analysis of the post-training pipeline and dynamics leading…

    why score 0.396
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.001
    contributability0.02×0.150.003
    venue0.92×0.100.092
    direct0.00×0.200.000

    2 matching tag(s)

  15. 0.39LessWrong2w
    Success Per Tokens

    Work smart more than hard, to expand the pareto frontier (but also work hard) A Pareto Frontier is a set of nondominated (optimal) solutions in multi-objective optimization. In 2 dimensions, this traces out a curve on…

    why score 0.395
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.00×0.250.001
    contributability0.02×0.150.003
    venue0.41×0.100.041
    direct1.00×0.200.200

    1 matching tag(s)

  16. 0.39LessWrong2dnew
    The Most Forbidden Technique is not always forbidden

    A few days ago, Goodfire announced a private beta of Silico, their LLM training platform. As part of the announcement, they made a post describing Silico's reproduction of RLFR, a method developed by Goodfire that uses…

    why score 0.394
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.40×0.250.100
    contributability0.42×0.150.064
    venue0.81×0.100.081
    direct0.00×0.200.000

    1 matching tag(s)

  17. 0.39LessWrong12d
    Persistent Latent Misalignment, a new dimension of misalignment?

    A new paper was released at ICML that I'm worried will open an entire new dimension of alignment problems: Latent Collaboration in Multi-Agent Systems (LatentMAS) TLDR: they show that multiagent systems can communicate…

    why score 0.393
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.02×0.250.004
    contributability0.27×0.150.040
    venue0.50×0.100.050
    direct0.00×0.200.000

    3 matching tag(s)

  18. 0.39LessWrong21hnew
    Does routine compression undo LLM unlearning? A short project

    I completed this project over 2 weeks as part of a BlueDot Project cohort. It was my first solo project and I learned a lot! Feedback is super welcome :) Code and full results: GitHub TLDR: I tested if standard…

    why score 0.387
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.74×0.250.185
    contributability0.02×0.150.003
    venue0.50×0.100.050
    direct0.00×0.200.000

    1 matching tag(s)

  19. 0.37LessWrong8d
    When is misalignment just a bug?

    Cross-posted from The Foretellix CTO Blog. Introduction and epistemic status: This is the first post in a planned series, “Alignment as a verification problem”. I co-originated coverage-driven verification (CDV), which…

    why score 0.370
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.05×0.250.014
    contributability0.02×0.150.003
    venue0.53×0.100.053
    direct0.00×0.200.000

    3 matching tag(s)

  20. 0.37LessWrong11d
    Free will as a model parameter

    The most popular take on the standard free will debate is that you are the algorithm. Your preferences and reasoning that determine your actions IS free will. But this resolution leaves me not entirely satisfied because…

    why score 0.369
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.02×0.250.006
    contributability0.12×0.150.017
    venue0.45×0.100.045
    direct0.00×0.200.000

    2 matching tag(s)

  21. 0.36LessWrong8d
    How robust are natural language autoencoders to initialization?

    Natural language autoencoders are meant to take in an LLM's activation vector and describe in plain text what the model is thinking. However, its training data collection involves asking Claude to guess what a model…

    why score 0.364
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.07×0.250.016
    contributability0.27×0.150.040
    venue0.83×0.100.083
    direct0.00×0.200.000

    tier-2: activation; 1 matching tag(s)

  22. 0.36LessWrong13d
    Sub-agent delegation chaining

    Epistemic status: pretty confident in the validity of the core proposal, not that confident in specific implementation details TL;DR: we should cryptographically verify that sub-agent instances/sessions are downstream…

    why score 0.363
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.01×0.250.002
    contributability0.96×0.150.144
    venue0.66×0.100.066
    direct0.00×0.200.000

    1 matching tag(s)

  23. 0.36LessWrong2w
    Scheming Evals Mislead in Both Directions

    We spent several weeks measuring in-context scheming, the behavior where a model covertly pursues a misaligned goal while outwardly appearing to comply, and the result that ended up surprising us had very little to do…

    why score 0.361
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.001
    contributability0.02×0.150.003
    venue0.58×0.100.058
    direct0.00×0.200.000

    3 matching tag(s)

  24. 0.35LessWrong13d
    Calibrating alignment evals

    Currently, alignment evaluation works by constructing a situation, observing the model's behavior and scoring it. We put a lot of thought into designing these benchmarks, and tuning them for our requirement. We are now…

    why score 0.348
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.01×0.250.002
    contributability0.02×0.150.003
    venue0.43×0.100.043
    direct0.00×0.200.000

    5 matching tag(s)

  25. 0.35LessWrong13d
    Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment

    By Roland Pihlakas and Jan Llenzl Dagohoy This post is a slightly updated copy of our Arxiv preprint available at https://arxiv.org/abs/2605.21401 . The tables are converted to images in order to preserve cell…

    why score 0.346
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.01×0.250.002
    contributability0.02×0.150.003
    venue0.41×0.100.041
    direct0.00×0.200.000

    2 matching tag(s)