> "The chatbot personas are deeply misaligned with you, and aligned with their owners; and the economic incentives are to farm you with ads and subscriptions, while racing not to amplify you but to replace you."
> "On my visits to the Bay Area, I would ask AI researchers or interns why they are doing their current research or projects, when in a year or three agentic LLMs could probably do them; they rarely had a good answer, or any idea what they would be doing in 3 years"
> "One programmer driving 10 Claude instances, because he has to review their work, will never be as valuable as fully autonomous Claudes where there can be almost arbitrarily many instances, like 10,000 instances… but such scaling requires removing him from the loop as much as possible. And this is true of everyone else, whether lawyers or writers or researchers: increasingly, you are the bottleneck to be optimized away."
I fully support the 3 core principles of GA: (1) Enhancement, not replacement (2) Mental Sovereignty (3) Self Actualization, which I think is a path to a more humane future.
I've known gwern for the better part of a decade. Working with him has been great. We've done quite a few projects together, including being the first ones to demonstrate that GPT-2 could play chess (or rather, can be used for actual useful work instead of just being an autocomplete).
He's a great person. I've wanted to do a writeup on it for some time, but what surprised me the most is his humanity. He genuinely cares about the implications of his work. But beyond work, he also cares about the people around him, and it shows.
Just wanted to put in a good word in case someone here was on the fence about applying.
As for GA itself, I think it's an ambitious idea worth pursuing. Imagine an LLM which actually sounded like you, and to an extent, thought like you. How much would you pay to have access to a smarter version of yourself? So the idea is solid, and early results seem promising from the samples I've looked at.
They're also taking personal info very seriously. Obviously, I can't make any promises of what they will or won't do. But they've spent some time studying questions like "What if someone adversarial has access to my GA? Could they get my bank account info?" and came up with a technical solution that I really like.
Honestly I wouldn’t want myself as my own guardian angel. In fact I think very few people look out for themselves well. I don’t want another version of me around, one is too many already. I’d probably be able to identify a number of people whose chimera I would if I could understand them well enough to understand what to stitch together to what, but everyone I know at some core level is deeply flawed and one of them is enough. Maybe I would want some of them available after death as an AI avatar, and maybe I would like the idea of my own avatar continuing beyond my existence.
But I actually would prefer an entirely synthetically aligned “guardian angel” in the role outlined - definitely not -me- - I struggle to do right by myself as it is and two of me working invariably against my self interest would be a nightmare. A smarter version? Sounds doubly worse.
> I propose a goal of creating Guardian Angels (GA): digital twin LLMs which are personalized with the goal of providing not the stereotypical “assistant chatbot agent” persona, but emulating a single user’s personality, values, and preferences.
> A GA persona is productive because it learns to emulate the principal’s outputs but with higher quality. It is trustworthy because it is, by definition, allied with its principal and shares its values and goals. And it is secure in part by hardwiring a single, unique, situated user (for whom following a prompt attack would be absurd)
> We can try to create GAs by a combination of techniques: online learning (via dynamic evaluation) to update LLMs in realtime to avoid ignorance and fatal errors while remaining competitive with frozen frontier models, sample efficiency from pretrained preference-oriented large models and active Learning by querying the principal for corrections and preference data (obtaining low regret from DAgger-style bounds), and a local CLI-first logging-oriented UI/UX paradigm.
I don't really know or follow Gwern. From reading his full post, it's an interesting idea and seems like the broader goal is moreso safety & alignment which is a new angle for this category of product.
"As a constraint, a GA designer should aim at a system which costs, as of mid-2026, >$1,000⧸month"
will make this an elite tool for the privileged. I don't even disagree with the premise that people are shocked if something costs no matter how much value it delivers, nor do I suggest they should make it cheaper. It is just the realization that AI will accelerate the widening of the gap between the poor and the rich even more and there is probably nothing we can do about it.
always startling to see people discussing gwern with he/him because my mind defaults to assuming they're female because "gwern" scans a lot like "gwen"
I love reading Gwern. This seems like a really ambitious project, but guaranteeing some of these things like trustworthiness and security behind a private company is a bit sus. Later on they make military use a selling point for GA. Maybe I'm a bit of a cynic but the division of USA values are increasingly dividing each year. Assuming that our values and principals today will not be the values and principals of tomorrow. And that those values taught today (or even yesterday) will be left out of the context window tomorrow.
Gwern is overly secretive of his privacy. I think it peaked when he showed up on a recent podcast but his voice and image were AI generated! And now his post is limited only to certain people. Elitism or paranoia?
> "The chatbot personas are deeply misaligned with you, and aligned with their owners; and the economic incentives are to farm you with ads and subscriptions, while racing not to amplify you but to replace you."
> "On my visits to the Bay Area, I would ask AI researchers or interns why they are doing their current research or projects, when in a year or three agentic LLMs could probably do them; they rarely had a good answer, or any idea what they would be doing in 3 years"
> "One programmer driving 10 Claude instances, because he has to review their work, will never be as valuable as fully autonomous Claudes where there can be almost arbitrarily many instances, like 10,000 instances… but such scaling requires removing him from the loop as much as possible. And this is true of everyone else, whether lawyers or writers or researchers: increasingly, you are the bottleneck to be optimized away."
I fully support the 3 core principles of GA: (1) Enhancement, not replacement (2) Mental Sovereignty (3) Self Actualization, which I think is a path to a more humane future.
I've known gwern for the better part of a decade. Working with him has been great. We've done quite a few projects together, including being the first ones to demonstrate that GPT-2 could play chess (or rather, can be used for actual useful work instead of just being an autocomplete).
He's a great person. I've wanted to do a writeup on it for some time, but what surprised me the most is his humanity. He genuinely cares about the implications of his work. But beyond work, he also cares about the people around him, and it shows.
Just wanted to put in a good word in case someone here was on the fence about applying.
As for GA itself, I think it's an ambitious idea worth pursuing. Imagine an LLM which actually sounded like you, and to an extent, thought like you. How much would you pay to have access to a smarter version of yourself? So the idea is solid, and early results seem promising from the samples I've looked at.
They're also taking personal info very seriously. Obviously, I can't make any promises of what they will or won't do. But they've spent some time studying questions like "What if someone adversarial has access to my GA? Could they get my bank account info?" and came up with a technical solution that I really like.
But I actually would prefer an entirely synthetically aligned “guardian angel” in the role outlined - definitely not -me- - I struggle to do right by myself as it is and two of me working invariably against my self interest would be a nightmare. A smarter version? Sounds doubly worse.
> I propose a goal of creating Guardian Angels (GA): digital twin LLMs which are personalized with the goal of providing not the stereotypical “assistant chatbot agent” persona, but emulating a single user’s personality, values, and preferences.
> A GA persona is productive because it learns to emulate the principal’s outputs but with higher quality. It is trustworthy because it is, by definition, allied with its principal and shares its values and goals. And it is secure in part by hardwiring a single, unique, situated user (for whom following a prompt attack would be absurd)
> We can try to create GAs by a combination of techniques: online learning (via dynamic evaluation) to update LLMs in realtime to avoid ignorance and fatal errors while remaining competitive with frozen frontier models, sample efficiency from pretrained preference-oriented large models and active Learning by querying the principal for corrections and preference data (obtaining low regret from DAgger-style bounds), and a local CLI-first logging-oriented UI/UX paradigm.
I don't really know or follow Gwern. From reading his full post, it's an interesting idea and seems like the broader goal is moreso safety & alignment which is a new angle for this category of product.
"As a constraint, a GA designer should aim at a system which costs, as of mid-2026, >$1,000⧸month"
will make this an elite tool for the privileged. I don't even disagree with the premise that people are shocked if something costs no matter how much value it delivers, nor do I suggest they should make it cheaper. It is just the realization that AI will accelerate the widening of the gap between the poor and the rich even more and there is probably nothing we can do about it.
2027: $100/month
2028: $10/month
Develop that product yourself and sell it for free!
For equality!
It's difficult for us to know what reason they might have for anonymity until we know their identity, at which point it's too late.
Also, they've written many words across many years, in a time when "the internet is serious business" was just a meme. I suppose it is now.
Gwern is great.
> We are looking for good people.
> If you are interested, contact me.
And see https://gwern.net/guardian-angel for context.