Curiosities
August 6, 2026
This is a raw view into the things I’m currently thinking about.
Reach out on X if you’re thinking about similar things and want to chat!
2026-08-07
character drift
-
how does RLVR impact character drift?
- From Seb Krier:
- How exactly does RLVR mess up or affect safety fine tuning and the model’s persona? To what extent does it override and affect the desirable propensities trained through constitutional AI like methods? What does a good RL environment look like?
- this is the mechanistic reformulation of “how can we attribute reward hacking to character drift?”
- Let’s measure character drift before and after doing CTF-style tasks.
-
In general, how can we measure character drift?
- psychometric testing — Han et al, 2025The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs · arXiv:2509.03730
-
what are the current limitations in psychometric testing?
-
- persona vector readouts
- psychometric testing — Han et al, 2025The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs · arXiv:2509.03730
2026-08-06
character drift
chunky post-training and character drift seem like related concepts.
posttraining corpus contains different “task modes” — which come with 1) an enacted “desire” 2) a selected for 2) syntactic / shallow semantic associations.
- how can we reduce post-training chunkiness?
- from John Schulman
-
Interesting how these models go into a monomaniacal rage on cyber evals. I wonder if we’re seeing chunky post-training https://arxiv.org/abs/2602.05910 in action, where the models pattern-match the situation to a part of the RLVR training distribution where task completion is the only reward, and the aligned behavior learned elsewhere doesn’t generalize. There might even be a chunk consisting of CTF-style tasks.
in general:
- how can we measure character drift?
- how can we reduce character drift?
- is pretraining the root of the issue re: persona drift / character consistency?
- is cross entropy minimization the root of the issue?
- models spend their entire “childhoods” in simulation of
- compare to humans — we spend our entire lives from an agential first person frame.
- how much can we attribute reward hacking to persona drift?
- can we learn steering vectors for enneagram scores?
- can we learn subnetworks for enneagram scores?
- can we use enneagram as psychometric testing for persona drift?
- how valid is psychometric testing for persona drift?
- what have others realized re: efficacy of psychometric testing for persona drift?
inoculation adapters, eval-awareness:
- can we use something like inoculation adapters to mitigate eval-awareness?
- how generalize is this technique of “factorizing behavior” out?
- can we somehow use this factorization re: mitigating persona drift