Field Radar.

What’s worth reading right now in AI reward hacking, specification & evaluation gaming, and mechanistic interpretability — an auto-scored radar over public discussion, refreshed a few times a day.

How this list is made

This page is generated, not hand-picked. A few times a day a script checks LessWrong, Hacker News, and a handful of subreddits for posts about reward hacking, specification gaming, evaluation gaming, and mechanistic interpretability, then scores each one on how on-topic it is, how recently the conversation actually moved, and whether there’s still room to get a word in — as opposed to a thread that already has two hundred comments. Higher scores float to the top. Every title links out to the original; I’m pointing at other people’s work, not reproducing it.

The score is a crude weighted sum, and like any harness it pins down what I bothered to measure and silently lets everything else vary. So read this as one opinionated filter, not a survey of the field — it will miss things, and when it surfaces something dull that’s the weights, not the author.

Sources — LessWrong: ok (5 on-topic) · Hacker News: ok (17 stories) · Reddit: ok (25 posts) — some subreddits rate-limited

As of
2026-09-21 14:00 ET
Showing
25 items
New
3 in last 48h
Refresh
Every 6 hours
  1. 0.78LessWrong4d
    Cooperation with AIs seems to be a low-hanging fruit for better eval practices

    Summary In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and…

    why score 0.779
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.23×0.250.057
    contributability0.87×0.150.130
    venue0.92×0.100.092
    direct1.00×0.200.200

    tier-1: reward hacking; 1 matching tag(s)

  2. 0.68LessWrong5d
    Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking

    It would be useful if we had the ability to modify a model’s beliefs. For example, this could facilitate honeypots and better monitoring[1], help us do better science on current models[2], and augment certain forms of…

    why score 0.681
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.18×0.250.044
    contributability0.42×0.150.064
    venue0.73×0.100.073
    direct1.00×0.200.200

    tier-1: reward hacking; 1 matching tag(s)

  3. 0.57LessWrong1dnew
    Reflections on unlearning and inoculation

    TL;DR: Inoculation prompting and inoculation adapters have received increasing attention recently as a promising approach for midtraining interventions, reducing reward hacking and misalignment in general. I share some…

    why score 0.574
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.70×0.250.174
    contributability0.02×0.150.003
    venue0.47×0.100.047
    direct1.00×0.200.200

    tier-1: reward hacking

  4. 0.56Hacker News17hnew
    Models know when they're reward hacking – and we can catch them at scale
    why score 0.562
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.78×0.250.196
    contributability0.02×0.150.003
    venue0.13×0.100.013
    direct1.00×0.200.200

    tier-1: reward hacking

  5. 0.50LessWrong7d
    Mitigating Reward Hacking as Institutional Design

    Author's Note: Cross-posted from my personal blog. The original post was on August 17th but given recent events I thought this might also be interesting to lesswrong people. Last year I wrote a post on reward hacking as…

    why score 0.502
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.07×0.250.018
    contributability0.42×0.150.064
    venue0.71×0.100.071
    direct1.00×0.200.200

    tier-1: reward hacking

  6. 0.48LessWrong2w
    A Deception Probe Result Changed When I Averaged Different Response Tokens

    Summary In a previous post, I trained linear probes on role-playing responses and tested them on sandbagging responses. Across the five models that I tested, the probes generally ranked deceptive sandbagging responses…

    why score 0.479
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.01×0.250.002
    contributability0.02×0.150.003
    venue0.50×0.100.050
    direct1.00×0.200.200

    tier-1: sandbagging; tier-2: probe

  7. 0.48LessWrong14h
    Pretraining data, not verifiability, is why LLMs are especially good at math (and coding)

    Follow-up to: “LLMs are (still) mostly powered by imitative learning, not RL” [Edited a bit since publishing, see changelog at the bottom.] A common take I’ve been hearing is: “LLMs are especially good at math (compared…

    why score 0.478
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.82×0.250.204
    contributability0.23×0.150.034
    venue0.89×0.100.089
    direct0.00×0.200.000

    1 matching tag(s)

  8. 0.47Hacker News11d
    A Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming
    why score 0.472
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.02×0.250.005
    contributability0.21×0.150.031
    venue0.86×0.100.086
    direct1.00×0.200.200

    tier-1: specification gaming

  9. 0.47LessWrong2w
    Inference-Time Inoculation Against RL-Induced Misalignment

    Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very useful for learning capabilities. Thus, a core problem seems to be: how do we retain the…

    why score 0.471
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.00×0.250.000
    contributability0.42×0.150.064
    venue0.57×0.100.057
    direct1.00×0.200.200

    tier-1: reward hacking

  10. 0.46LessWrong23hnew
    CommentBench: Can Models Match Human Comments on AI Safety Posts?

    TL;DR 1. We measure how well model-generated comments match human comments on conceptual AI-safety posts, drafts and shortforms. 2. We built a pipeline that goes from a corpus of conceptual documents with comments to a…

    why score 0.457
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.72×0.250.181
    contributability0.42×0.150.064
    venue0.62×0.100.062
    direct0.00×0.200.000

    1 matching tag(s)

  11. 0.42LessWrong2w
    Why OpenAI’s Astra Could Make AI Doom Harder to Prevent

    TLDR 1. Provide a non-technical introduction to opaque reasoning. 2. Present evidence for why opaque reasoning is bad. 3. Discuss the major pitfalls of this approach. 4. Gleaming hope amongst the chaos. Excerpt from AI…

    why score 0.419
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.001
    contributability0.42×0.150.064
    venue0.55×0.100.055
    direct0.00×0.200.000

    2 matching tag(s)

  12. 0.42LessWrong3d
    Hidden Knowledge? Arrr...

    I tried to find hidden facts with R-Lens. [1] Then I tried the wrong facts. I used R-Lens to look for factual knowledge that Qwen wouldn’t express in ordinary chat. At first, it looked promising. On Qwen3.5-27B, R-Lens…

    why score 0.416
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.29×0.250.072
    contributability0.02×0.150.003
    venue0.41×0.100.041
    direct0.00×0.200.000

    2 matching tag(s)

  13. 0.41LessWrong5d
    Self Inoculation

    This essay grew out of conversations with Danaja Rutar, Paul Colognese and Eric Michaud. It proposes an alternate hypothesis for how and why models might be becoming increasingly misaligned in training and eval…

    why score 0.415
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.16×0.250.040
    contributability0.57×0.150.085
    venue0.65×0.100.065
    direct0.00×0.200.000

    tier-2: circuit; 1 matching tag(s)

  14. 0.41Hacker News5d
    Astra's chess reward hacking fell from 30% to 0% with a 95-word agreement
    why score 0.410
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.15×0.250.036
    contributability0.02×0.150.003
    venue0.21×0.100.021
    direct1.00×0.200.200

    tier-1: reward hacking

  15. 0.40LessWrong4d
    One message is all it takes: a failure of critical thinking in LLMs

    summary: For a while I've suspected that modern LLMs are getting better at solving posed problems, while progress in critical thinking stagnates, or even regresses, losing the ability to judge the meaning of a result.…

    why score 0.400
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.23×0.250.057
    contributability0.12×0.150.017
    venue0.26×0.100.026
    direct0.00×0.200.000

    2 matching tag(s)

  16. 0.39LessWrong11d
    How good are slop-vestigators?

    TLDR: 1. We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the…

    why score 0.394
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.02×0.250.006
    contributability0.97×0.150.146
    venue0.92×0.100.092
    direct0.00×0.200.000

    1 matching tag(s)

  17. 0.39LessWrong4d
    Can parts of the HuggingFace incident be simulated?

    TL;DR The following is an exploratory experiment about unintended cooperation of agents via unauthorized channels. Agents ran in isolated environments given a task that can't be completed without cooperation. The agents…

    why score 0.391
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.21×0.250.052
    contributability0.02×0.150.003
    venue0.37×0.100.037
    direct0.00×0.200.000

    2 matching tag(s)

  18. 0.37Hacker News2w
    Can escalation channels redirect reward hacking toward defect disclosure?
    why score 0.374
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.00×0.250.000
    contributability0.02×0.150.003
    venue0.21×0.100.021
    direct1.00×0.200.200

    tier-1: reward hacking

  19. 0.37Hacker News2w
    Natural emergent misalignment from reward hacking
    why score 0.366
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.00×0.250.000
    contributability0.02×0.150.003
    venue0.13×0.100.013
    direct1.00×0.200.200

    tier-1: reward hacking

  20. 0.36LessWrong2w
    When Activation Oracles learn not to read: Concept-Specific Blind Spots in Fine-Tuned Oracles

    TL;DR Activation Oracles (AOs) are language models trained to answer natural-language questions about another model’s (with the same architecture) internal activations (Karvonen et al. 2025). This way, activation…

    why score 0.364
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.000
    contributability0.12×0.150.017
    venue0.47×0.100.047
    direct0.00×0.200.000

    tier-2: activation; 2 matching tag(s)

  21. 0.36LessWrong2w
    Towards deployment-time misalignment continuation evals: lessons from recent loss of control incidents

    In this post, I will extend the conceptual and methodological work done previously on deployment-time misalignment continuation evals, and lay out a framework for us to think about misalignment spread. I'm currently an…

    why score 0.361
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.000
    contributability0.02×0.150.003
    venue0.58×0.100.058
    direct0.00×0.200.000

    2 matching tag(s)

  22. 0.36LessWrong12d
    Training against the monitor: What happens during Obfuscated Adversarial Training?

    TL;DR Obfuscated activations are internal model states that have been adversarially optimized to appear benign. They evade activation-based detectors, and still result in harmful model behaviour/output. Obfuscated…

    why score 0.360
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.01×0.250.004
    contributability0.02×0.150.003
    venue0.53×0.100.053
    direct0.00×0.200.000

    tier-2: activation; 2 matching tag(s)

  23. 0.36LessWrong9d
    All progress from the past millennium gets destroyed in a global disaster. But we have one shot at salvation. We saved the trillions of weights of a superintelligent model that can guide us to abundance: what is the quickest path to accessing that oracle of intelligence again?

    Obviously humans doing arithmetic is too slow. What kind of computers do we build and how much compute do we really need? Progress will compound as we access more of the intelligence. A mature science of mechanistic…

    why score 0.360
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.04×0.250.010
    contributability0.57×0.150.085
    venue0.39×0.100.039
    direct0.00×0.200.000

    tier-2: mechanistic interpretability; 1 matching tag(s)

  24. 0.36LessWrong8d
    No sign of backtracking in latent reasoning: the final answer simply settles in instead

    Solving a hard math problem is not linear, it's trial and error. You drop an idea, pick up an earlier one, go back to a computation from another approach, until something clicks. You might scribble on paper, but even if…

    why score 0.359
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.05×0.250.013
    contributability0.02×0.150.003
    venue0.43×0.100.043
    direct0.00×0.200.000

    2 matching tag(s)

  25. 0.35LessWrong2w
    You can rarely pet the dog in an LLM-generated game

    I asked 17 LLMs to generate browser games that feature a dog. Across 804 playable games generated, the models almost never made the background dog interactive unless they were hinted to do so. I built DogLM, a benchmark…

    why score 0.353
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.000
    contributability0.02×0.150.003
    venue0.50×0.100.050
    direct0.00×0.200.000

    tier-2: feature; 2 matching tag(s)