Field Radar.

What’s worth reading right now in AI reward hacking, specification & evaluation gaming, and mechanistic interpretability — an auto-scored radar over public discussion, refreshed a few times a day.

How this list is made

This page is generated, not hand-picked. A few times a day a script checks LessWrong, Hacker News, and a handful of subreddits for posts about reward hacking, specification gaming, evaluation gaming, and mechanistic interpretability, then scores each one on how on-topic it is, how recently the conversation actually moved, and whether there’s still room to get a word in — as opposed to a thread that already has two hundred comments. Higher scores float to the top. Every title links out to the original; I’m pointing at other people’s work, not reproducing it.

The score is a crude weighted sum, and like any harness it pins down what I bothered to measure and silently lets everything else vary. So read this as one opinionated filter, not a survey of the field — it will miss things, and when it surfaces something dull that’s the weights, not the author.

Sources — LessWrong: ok (11 on-topic) · Hacker News: ok (7 stories) · Reddit: ok (25 posts) — some subreddits rate-limited

As of
2026-07-28 02:00 ET
Showing
25 items
New
7 in last 48h
Refresh
Every 6 hours
  1. 0.60LessWrong11hnew
    Multi-Turn Drift Increases Scheming

    TLDR - We talk about scheming, and why research on this phenomenon is crucial for AI safety. We find a particular environment/scenario where scheming happens at a higher rate than normal. We provide hypotheses for why…

    why score 0.604
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.85×0.250.213
    contributability0.27×0.150.040
    venue0.51×0.100.051
    direct0.00×0.200.000

    2 matching tag(s)

  2. 0.59LessWrong5d
    A Multi-Agent Extension for Petri

    Intro Petri is an open-source framework built on Inspect AI for automated AI Safety evaluations first released by Anthropic, but now maintained and developed by Meridian Labs. Each evaluation involves three agents, the…

    why score 0.590
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.17×0.250.042
    contributability0.02×0.150.003
    venue0.45×0.100.045
    direct1.00×0.200.200

    2 matching tag(s)

  3. 0.54LessWrong2w
    Linear Probes add little for Verifiable Reward Hacking

    Summary Tested whether linear probes can detect reward hacking early during GRPO training on a small model. Used a synthetic arithmetic task with a planted bug in the reward checker. Probes achieved near-perfect…

    why score 0.544
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.01×0.250.002
    contributability0.02×0.150.003
    venue0.39×0.100.039
    direct1.00×0.200.200

    tier-1: reward hacking; 1 matching tag(s)

  4. 0.53LessWrong10hnew
    When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models

    By Sai Kartheek Reddy Kasu, Nils Lukas, and Samuele Poppi This post is a summary of our accepted paper at the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN). The full paper is available here TL;DR The Setup:…

    why score 0.531
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.86×0.250.216
    contributability0.02×0.150.003
    venue0.13×0.100.013
    direct0.00×0.200.000

    2 matching tag(s)

  5. 0.51LessWrong1d
    LLMs are (still) mostly powered by imitative learning, not RL

    Reinforcement learning from verifiable rewards (RLVR) is the hot new thing in LLM training. It’s so hot, and people spend so much time talking about it, that they sometimes lose sight of the big picture. Stepping back,…

    why score 0.512
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.64×0.250.160
    contributability0.72×0.150.108
    venue0.94×0.100.094
    direct0.00×0.200.000

    1 matching tag(s)

  6. 0.51LessWrong2w
    Bounding eval awareness of ~human-level AI across the safe-to-dangerous shift

    In our last post, we argued that measuring evaluation awareness is fundamentally challenging because of the safe-to-dangerous distributional shift: we cannot directly measure the evaluation awareness of a model without…

    why score 0.506
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.00×0.250.000
    contributability0.57×0.150.085
    venue0.70×0.100.070
    direct1.00×0.200.200

    tier-1: evaluation awareness

  7. 0.50LessWrong1dnew
    Inoculate or Reflect? Two training interventions under prompting, steering, and patching

    Anthropic's recent paper, Verbalizable Representations Form a Global Workspace in Language Models, contains a small experiment near the end that we found more interesting than the main findings. Surprising that it's so…

    why score 0.498
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.61×0.250.152
    contributability0.02×0.150.003
    venue0.43×0.100.043
    direct0.00×0.200.000

    2 matching tag(s)

  8. 0.46LessWrong2w
    A global workspace in language models

    [This is the blog post for our new paper Verbalizable Representations Form a Global Workspace in Language Models Readers might also be interested in: the Public commentary, Github and Neuronpedia] As you read this…

    why score 0.462
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.000
    contributability0.41×0.150.062
    venue1.00×0.100.100
    direct0.00×0.200.000

    2 matching tag(s)

  9. 0.45Hacker News3d
    Show HN: VinvAI – Ties runtime trace to code segment, prevents reward hacking
    why score 0.453
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.32×0.250.080
    contributability0.02×0.150.003
    venue0.21×0.100.021
    direct1.00×0.200.200

    tier-1: reward hacking

  10. 0.45LessWrong3d
    Linear probes tell you where quantization will hurt

    Epistemic status: I have only tested one encoder family (BERT-base and its relatives) and one decoder LLM (Qwen2.5-3B), one seed, token-level tasks, and post-training weight quantization. I trust the results because I…

    why score 0.450
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.34×0.250.085
    contributability0.02×0.150.003
    venue0.63×0.100.063
    direct0.00×0.200.000

    2 matching tag(s)

  11. 0.45Hacker News3d
    Show HN: Vinv-Ties every runtime trace to code segment, prevents reward hacking
    why score 0.449
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.28×0.250.070
    contributability0.02×0.150.003
    venue0.26×0.100.026
    direct1.00×0.200.200

    tier-1: reward hacking

  12. 0.45LessWrong3d
    SONI: Selective Orthogonalisation via Noise Injection

    This project was completed as a capstone for TARA. All code is available in github. TL;DR The Problem: Neural networks use superposition to pack many concepts into small latent spaces by making feature vectors…

    why score 0.445
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.34×0.250.085
    contributability0.02×0.150.003
    venue0.57×0.100.057
    direct0.00×0.200.000

    tier-2: feature, activation; 1 matching tag(s)

  13. 0.44LessWrong21hnew
    Can we teach a model to encode a semantic feature on a chosen manifold in just three channels?

    This is my submission to BlueDot's Technical AI Safety Puzzle #1, for which I received an Honorable Mention. Congratulations to Gustavo Korzune Gurgel, Patryk Perduta (his amazing write-up), Sam Spilllard, Karine…

    why score 0.442
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.74×0.250.184
    contributability0.02×0.150.003
    venue0.30×0.100.030
    direct0.00×0.200.000

    tier-2: feature; 1 matching tag(s)

  14. 0.44LessWrong7d
    Fable is SOTA at CIFAR Speedrun (& specification gaming)

    Fulcrum is working on an AI R&D optimization benchmark. Here, we present results from one of our tasks, including preliminary results from Fable. For more detail on Fable’s solution, check out…

    why score 0.438
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.08×0.250.021
    contributability0.02×0.150.003
    venue0.64×0.100.064
    direct1.00×0.200.200

    tier-1: specification gaming

  15. 0.43LessWrong14hnew
    Quadrillion Param Costs: KV Cache, Context Length, Frontier Margins

    The models of 2028-2031 get much bigger than the models of 2026, going from 10T total params in 2026 to maybe 240T params in 2028 [1] and then 1.4 quadrillion params in 2031, as I estimate in the previous post from HBM…

    why score 0.433
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.81×0.250.203
    contributability0.02×0.150.003
    venue0.77×0.100.077
    direct0.00×0.200.000

    1 matching tag(s)

  16. 0.43LessWrong1dnew
    What Happens When a Collusion Probe Only Finds a Thin Signal?

    From the SPEC-GAP pre-fellowship phase to the fellowship phase which involves live indirect prompt-injection trajectories. TL;DR Prior work found that linear probes could distinguish honest from deceptive responses in a…

    why score 0.429
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.64×0.250.160
    contributability0.02×0.150.003
    venue0.41×0.100.041
    direct0.00×0.200.000

    tier-2: probe; 1 matching tag(s)

  17. 0.42LessWrong1d
    Challenge: Hand coding weights for efficient sequence memorisation

    We hand coded weights for one layer MLPs that memorises labels for input token sequences of length two. The number of facts our hand-coded models can memorise with 90% accuracy[1]scales roughly linearly with the models'…

    why score 0.424
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.55×0.250.138
    contributability0.42×0.150.064
    venue0.73×0.100.073
    direct0.00×0.200.000

    1 matching tag(s)

  18. 0.42LessWrong2w
    Tie training can make DPO/RLHF-trained AIs generalize better

    This post covers our recent ICML paper: Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training. TL;DR Our theorems and experiments suggest that DPO and RLHF…

    why score 0.418
    signalvalueweightpoints
    topic0.75×0.300.225
    liveness0.00×0.250.000
    contributability0.79×0.150.118
    venue0.74×0.100.074
    direct0.00×0.200.000

    tier-2: feature; 1 matching tag(s)

  19. 0.42LessWrong4d
    V&V takes on OpenAI’s long-horizon incidents

    [Cross-posted from The Foretellix CTO Blog. These short takes try to put a verification-and-validation slant on AI-safety / alignment topics – they are not full treatments. I co-originated coverage-driven verification…

    why score 0.417
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.22×0.250.055
    contributability0.02×0.150.003
    venue0.59×0.100.059
    direct0.00×0.200.000

    2 matching tag(s)

  20. 0.41LessWrong4d
    Fixing rewards for NLA to reduce confabulation

    Hello, This is my first post on Lesswrong. Hope my contribution makes the world a better and safer place. Note: 1. This post is 100% human-written. 2. Full paper in preparation for ICLR 2027 Anthropic's NLA(Natural…

    why score 0.405
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.25×0.250.061
    contributability0.02×0.150.003
    venue0.41×0.100.041
    direct0.00×0.200.000

    tier-2: mechanistic interpretability, sparse autoencoder; 1 matching tag(s)

  21. 0.40LessWrong2w
    Models are blind outside the J-space. NLAs aren't.

    TLDR: On Llama-3.3-70B, I found thoughts it cannot see that are actively steering its behavior; and Anthropic's released NLA (Natural Language Autoencoder) reads them anyway. When asked if it sees a hidden thought, the…

    why score 0.402
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.001
    contributability0.27×0.150.040
    venue0.61×0.100.061
    direct0.00×0.200.000

    2 matching tag(s)

  22. 0.39LessWrong5d
    Mechanistic interpretability hypotheses for Measuring Reward-Seeking by Instilling Contrastive Beliefs and additional comments

    This is interesting research! https://alignment.openai.com/measuring-reward-seeking It made me think of few overlapping hypotheses for what might be happening here, how did the grader behavior emerge at the pretraining…

    why score 0.394
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.15×0.250.038
    contributability0.02×0.150.003
    venue0.53×0.100.053
    direct0.00×0.200.000

    tier-2: mechanistic interpretability; 2 matching tag(s)

  23. 0.39LessWrong2w
    Persistent Latent Misalignment, a new dimension of misalignment?

    A new paper was released at ICML that I'm worried will open an entire new dimension of alignment problems: Latent Collaboration in Multi-Agent Systems (LatentMAS) TLDR: they show that multiagent systems can communicate…

    why score 0.390
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.00×0.250.000
    contributability0.27×0.150.040
    venue0.50×0.100.050
    direct0.00×0.200.000

    3 matching tag(s)

  24. 0.38LessWrong1dnew
    You don't need error nodes, you need better features

    This is a cross-post from my blog. It is a follow-up to the methods I developed in a previous post on replacement-aware training. ---------------------------------------- Summary A replacement model (Ameisen et al.…

    why score 0.384
    signalvalueweightpoints
    topic0.50×0.300.150
    liveness0.70×0.250.175
    contributability0.02×0.150.003
    venue0.56×0.100.056
    direct0.00×0.200.000

    1 matching tag(s)

  25. 0.38LessWrong8d
    Is there even a ground-truth for LLMs’ internal representations?

    [This is an introductory blog for the paper Laguerre Geometry for Interpreting Large Language Models and the GitHub repository Geometric Lens.] LLM Lens: What does an internal vector mean? Anthropic's recent paper on…

    why score 0.379
    signalvalueweightpoints
    topic1.00×0.300.300
    liveness0.07×0.250.017
    contributability0.02×0.150.003
    venue0.60×0.100.060
    direct0.00×0.200.000

    2 matching tag(s)