Hraness
Theme
Appearance

saved

Impersona-Env: User Models as Environments for Model Personalization

by Quan ShiBen Shi

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

Quan Shi describes Impersona-Env, a proof of concept that turns a personal user model into an RL environment for training a chat assistant. After Impersona-style fine-tuning on his messages (with relationship context and reconstructed thoughts), a user model stands in for him so an assistant can practice, collect private critique as preference pairs, and learn situational collaboration preferences that discrete memory systems miss.

ideas

  • User models as environments, not just lookups. At inference the assistant can query a model of you; at training time the same model is the environment and reward so the assistant practices at a scale no person would sit through.
  • Personalization beyond fact memory. Memory stores discrete facts before knowing what will matter; learning through interaction with a user model is closer to how people learn to work with each other.
  • Faithful user models need context and thoughts. SFT on message history improves when examples include relationship summaries, preceding situation, and reconstructed internal monologue emitted as inspectable thinking before each reply.
  • The assistant is only as good as its user model. Calibration of the interaction distribution matters; otherwise the assistant learns preferences for situations that never arise with a real assistant.
  • Evaluation is the bottleneck. Federated, privacy-preserving evals over one person's data remain hard; coherence, style, fact/decision checks, and disagreement rate are a first pass, not a solved metric.

quotes

“can a chat assistant learn from us the way a human assistant would?”

Quan Shi, stating the training-time question.

“At training time, the user model can stand in for you as the environment and the reward.”

Quan Shi, defining the RL setup.

“At inference time, the assistant can ask the user model instead of asking you.”

Quan Shi, naming the in-the-moment use.

“The assistant can only be as good as the user model it learns from”

Quan Shi, stating the fidelity constraint.