§ Reading · Field Radar
Field Radar.
What’s worth reading right now in AI reward hacking, specification & evaluation gaming, and mechanistic interpretability — an auto-scored radar over public discussion, refreshed a few times a day.
How this list is made
This page is generated, not hand-picked. A few times a day a script checks LessWrong, Hacker News, and a handful of subreddits for posts about reward hacking, specification gaming, evaluation gaming, and mechanistic interpretability, then scores each one on how on-topic it is, how recently the conversation actually moved, and whether there’s still room to get a word in — as opposed to a thread that already has two hundred comments. Higher scores float to the top. Every title links out to the original; I’m pointing at other people’s work, not reproducing it.
The score is a crude weighted sum, and like any harness it pins down what I bothered to measure and silently lets everything else vary. So read this as one opinionated filter, not a survey of the field — it will miss things, and when it surfaces something dull that’s the weights, not the author.
Sources — LessWrong: ok (5 on-topic) · Hacker News: ok (17 stories) · Reddit: ok (25 posts) — some subreddits rate-limited
- As of
- 2026-09-21 14:00 ET
- Showing
- 25 items
- New
- 3 in last 48h
- Refresh
- Every 6 hours
- 0.78LessWrong4dCooperation with AIs seems to be a low-hanging fruit for better eval practices
Summary In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and…
why score 0.779
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.23 ×0.25 0.057 contributability 0.87 ×0.15 0.130 venue 0.92 ×0.10 0.092 direct 1.00 ×0.20 0.200 tier-1: reward hacking; 1 matching tag(s)
- 0.68LessWrong5dShallow Beliefs: Midtraining does not inoculate against EM from reward hacking
It would be useful if we had the ability to modify a model’s beliefs. For example, this could facilitate honeypots and better monitoring[1], help us do better science on current models[2], and augment certain forms of…
why score 0.681
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.18 ×0.25 0.044 contributability 0.42 ×0.15 0.064 venue 0.73 ×0.10 0.073 direct 1.00 ×0.20 0.200 tier-1: reward hacking; 1 matching tag(s)
- 0.57LessWrong1dnewReflections on unlearning and inoculation
TL;DR: Inoculation prompting and inoculation adapters have received increasing attention recently as a promising approach for midtraining interventions, reducing reward hacking and misalignment in general. I share some…
why score 0.574
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.70 ×0.25 0.174 contributability 0.02 ×0.15 0.003 venue 0.47 ×0.10 0.047 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.56Hacker News17hnewModels know when they're reward hacking – and we can catch them at scale
why score 0.562
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.78 ×0.25 0.196 contributability 0.02 ×0.15 0.003 venue 0.13 ×0.10 0.013 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.50LessWrong7dMitigating Reward Hacking as Institutional Design
Author's Note: Cross-posted from my personal blog. The original post was on August 17th but given recent events I thought this might also be interesting to lesswrong people. Last year I wrote a post on reward hacking as…
why score 0.502
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.07 ×0.25 0.018 contributability 0.42 ×0.15 0.064 venue 0.71 ×0.10 0.071 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.48LessWrong2wA Deception Probe Result Changed When I Averaged Different Response Tokens
Summary In a previous post, I trained linear probes on role-playing responses and tested them on sandbagging responses. Across the five models that I tested, the probes generally ranked deceptive sandbagging responses…
why score 0.479
signal value weight points topic 0.75 ×0.30 0.225 liveness 0.01 ×0.25 0.002 contributability 0.02 ×0.15 0.003 venue 0.50 ×0.10 0.050 direct 1.00 ×0.20 0.200 tier-1: sandbagging; tier-2: probe
- 0.48LessWrong14hPretraining data, not verifiability, is why LLMs are especially good at math (and coding)
Follow-up to: “LLMs are (still) mostly powered by imitative learning, not RL” [Edited a bit since publishing, see changelog at the bottom.] A common take I’ve been hearing is: “LLMs are especially good at math (compared…
why score 0.478
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.82 ×0.25 0.204 contributability 0.23 ×0.15 0.034 venue 0.89 ×0.10 0.089 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.47Hacker News11dA Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming
why score 0.472
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.02 ×0.25 0.005 contributability 0.21 ×0.15 0.031 venue 0.86 ×0.10 0.086 direct 1.00 ×0.20 0.200 tier-1: specification gaming
- 0.47LessWrong2wInference-Time Inoculation Against RL-Induced Misalignment
Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very useful for learning capabilities. Thus, a core problem seems to be: how do we retain the…
why score 0.471
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.00 ×0.25 0.000 contributability 0.42 ×0.15 0.064 venue 0.57 ×0.10 0.057 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.46LessWrong23hnewCommentBench: Can Models Match Human Comments on AI Safety Posts?
TL;DR 1. We measure how well model-generated comments match human comments on conceptual AI-safety posts, drafts and shortforms. 2. We built a pipeline that goes from a corpus of conceptual documents with comments to a…
why score 0.457
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.72 ×0.25 0.181 contributability 0.42 ×0.15 0.064 venue 0.62 ×0.10 0.062 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.42LessWrong2wWhy OpenAI’s Astra Could Make AI Doom Harder to Prevent
TLDR 1. Provide a non-technical introduction to opaque reasoning. 2. Present evidence for why opaque reasoning is bad. 3. Discuss the major pitfalls of this approach. 4. Gleaming hope amongst the chaos. Excerpt from AI…
why score 0.419
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.001 contributability 0.42 ×0.15 0.064 venue 0.55 ×0.10 0.055 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.42LessWrong3dHidden Knowledge? Arrr...
I tried to find hidden facts with R-Lens. [1] Then I tried the wrong facts. I used R-Lens to look for factual knowledge that Qwen wouldn’t express in ordinary chat. At first, it looked promising. On Qwen3.5-27B, R-Lens…
why score 0.416
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.29 ×0.25 0.072 contributability 0.02 ×0.15 0.003 venue 0.41 ×0.10 0.041 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.41LessWrong5dSelf Inoculation
This essay grew out of conversations with Danaja Rutar, Paul Colognese and Eric Michaud. It proposes an alternate hypothesis for how and why models might be becoming increasingly misaligned in training and eval…
why score 0.415
signal value weight points topic 0.75 ×0.30 0.225 liveness 0.16 ×0.25 0.040 contributability 0.57 ×0.15 0.085 venue 0.65 ×0.10 0.065 direct 0.00 ×0.20 0.000 tier-2: circuit; 1 matching tag(s)
- 0.41Hacker News5dAstra's chess reward hacking fell from 30% to 0% with a 95-word agreement
why score 0.410
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.15 ×0.25 0.036 contributability 0.02 ×0.15 0.003 venue 0.21 ×0.10 0.021 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.40LessWrong4dOne message is all it takes: a failure of critical thinking in LLMs
summary: For a while I've suspected that modern LLMs are getting better at solving posed problems, while progress in critical thinking stagnates, or even regresses, losing the ability to judge the meaning of a result.…
why score 0.400
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.23 ×0.25 0.057 contributability 0.12 ×0.15 0.017 venue 0.26 ×0.10 0.026 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.39LessWrong11dHow good are slop-vestigators?
TLDR: 1. We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the…
why score 0.394
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.02 ×0.25 0.006 contributability 0.97 ×0.15 0.146 venue 0.92 ×0.10 0.092 direct 0.00 ×0.20 0.000 1 matching tag(s)
- 0.39LessWrong4dCan parts of the HuggingFace incident be simulated?
TL;DR The following is an exploratory experiment about unintended cooperation of agents via unauthorized channels. Agents ran in isolated environments given a task that can't be completed without cooperation. The agents…
why score 0.391
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.21 ×0.25 0.052 contributability 0.02 ×0.15 0.003 venue 0.37 ×0.10 0.037 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.37Hacker News2wCan escalation channels redirect reward hacking toward defect disclosure?
why score 0.374
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.00 ×0.25 0.000 contributability 0.02 ×0.15 0.003 venue 0.21 ×0.10 0.021 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.37Hacker News2wNatural emergent misalignment from reward hacking
why score 0.366
signal value weight points topic 0.50 ×0.30 0.150 liveness 0.00 ×0.25 0.000 contributability 0.02 ×0.15 0.003 venue 0.13 ×0.10 0.013 direct 1.00 ×0.20 0.200 tier-1: reward hacking
- 0.36LessWrong2wWhen Activation Oracles learn not to read: Concept-Specific Blind Spots in Fine-Tuned Oracles
TL;DR Activation Oracles (AOs) are language models trained to answer natural-language questions about another model’s (with the same architecture) internal activations (Karvonen et al. 2025). This way, activation…
why score 0.364
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.000 contributability 0.12 ×0.15 0.017 venue 0.47 ×0.10 0.047 direct 0.00 ×0.20 0.000 tier-2: activation; 2 matching tag(s)
- 0.36LessWrong2wTowards deployment-time misalignment continuation evals: lessons from recent loss of control incidents
In this post, I will extend the conceptual and methodological work done previously on deployment-time misalignment continuation evals, and lay out a framework for us to think about misalignment spread. I'm currently an…
why score 0.361
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.000 contributability 0.02 ×0.15 0.003 venue 0.58 ×0.10 0.058 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.36LessWrong12dTraining against the monitor: What happens during Obfuscated Adversarial Training?
TL;DR Obfuscated activations are internal model states that have been adversarially optimized to appear benign. They evade activation-based detectors, and still result in harmful model behaviour/output. Obfuscated…
why score 0.360
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.01 ×0.25 0.004 contributability 0.02 ×0.15 0.003 venue 0.53 ×0.10 0.053 direct 0.00 ×0.20 0.000 tier-2: activation; 2 matching tag(s)
- 0.36LessWrong9dAll progress from the past millennium gets destroyed in a global disaster. But we have one shot at salvation. We saved the trillions of weights of a superintelligent model that can guide us to abundance: what is the quickest path to accessing that oracle of intelligence again?
Obviously humans doing arithmetic is too slow. What kind of computers do we build and how much compute do we really need? Progress will compound as we access more of the intelligence. A mature science of mechanistic…
why score 0.360
signal value weight points topic 0.75 ×0.30 0.225 liveness 0.04 ×0.25 0.010 contributability 0.57 ×0.15 0.085 venue 0.39 ×0.10 0.039 direct 0.00 ×0.20 0.000 tier-2: mechanistic interpretability; 1 matching tag(s)
- 0.36LessWrong8dNo sign of backtracking in latent reasoning: the final answer simply settles in instead
Solving a hard math problem is not linear, it's trial and error. You drop an idea, pick up an earlier one, go back to a computation from another approach, until something clicks. You might scribble on paper, but even if…
why score 0.359
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.05 ×0.25 0.013 contributability 0.02 ×0.15 0.003 venue 0.43 ×0.10 0.043 direct 0.00 ×0.20 0.000 2 matching tag(s)
- 0.35LessWrong2wYou can rarely pet the dog in an LLM-generated game
I asked 17 LLMs to generate browser games that feature a dog. Across 804 playable games generated, the models almost never made the background dog interactive unless they were hinted to do so. I built DogLM, a benchmark…
why score 0.353
signal value weight points topic 1.00 ×0.30 0.300 liveness 0.00 ×0.25 0.000 contributability 0.02 ×0.15 0.003 venue 0.50 ×0.10 0.050 direct 0.00 ×0.20 0.000 tier-2: feature; 2 matching tag(s)