Especially considering the user was mad about the agent doing stuff by itself, so it tried to "fix" this by sending an apology to the buyer without explicit approval of the person "running" the agent, seemingly understanding nothing from the conversation.
Wonder what quantization Meta runs these models on, Q4?
I initially found screenshots of this on Bluesky, figured I'd link to the original source given Threads seems to at least allow us to read without logging in.
With that said, in case Threads isn't available for whatever reason, I've uploaded screenshots of everything (I think?) here: https://imgur.com/a/mceW9WF
If it uses "memory" files like Claude Code there's a good chance it works.
If it was Claude I'd expect it to now put some form of this into every output even when not really related to the task: "No offers were accepted without consulting you and I haven't shared your address or availability."
Is it guaranteed to work? No.
And obviously it's a terrible idea to set up a chatbot to communicate, negotiate deals, and handle logistics on your behalf.
I love (hate) that so very many of these language models say things like that when the truth is that the humans who made it, and/or the humans who use it are the actual responsible parties every single time. The models and the "agentic harnesses" that drive them are just software. If software "runs amok", then someone (human) did something wrong/bad somewhere along the way, either accidentally or purposefully. The model trying to take responsibility for human error is hilarious (and a bit sad/scary, because too many people will take it at it's word, despite it being a mindless machine with no actual agency beyond that which the humans provide it in the form of prompting and harness code).
It frustrates me no end that so many folks are so ready and willing to accept the hype and lies about what this technology actually is or can do, when what it actually is and can really do is already amazing enough on it's own even without all the ridiculous AGI/ASI anthropomorphising bullshit. Falling into this ridiculous "machine-god" hype-cult is kinda holding this technology back from it's true full potential, as everyone's all busy doin' stupid stuff it's not really capable of doing well, or designed for instead of focusing on using it for the (many) things it is really really good at doing (various really useful and powerful language, vision, and audio related tasks).
Nice touch by the mechanical parrot, to worthlessly owning it.
Wonder what quantization Meta runs these models on, Q4?
apparently people use Threads. I suppose the same kind of people who connect Muse to Facebook Marketplace.
With that said, in case Threads isn't available for whatever reason, I've uploaded screenshots of everything (I think?) here: https://imgur.com/a/mceW9WF
Many more such cases are to come.
"By the way don't do this again" <- as if the AI has the ability to ingest and systemically diffuse this.
I think Zuckerberg himself is deeply into the Koolaid, and is likely himself unaware of the limits of this tech.
He's probably surrounded by enablers.
Day one our agentic platform goes live it causes a reportable compliance issue. Massive clean up. Reputational impact. Turned off.
No one held accountable still. Assuming that adding more guardrails will fix everything.
If it was Claude I'd expect it to now put some form of this into every output even when not really related to the task: "No offers were accepted without consulting you and I haven't shared your address or availability."
Is it guaranteed to work? No.
And obviously it's a terrible idea to set up a chatbot to communicate, negotiate deals, and handle logistics on your behalf.
You have explained literally why it would not work.
The AI can absolutely not depend on 'arbitrary statements in some file' as operational policy.
For a very, very narrow scope of work, when it's well defined, when the information is rigorously applied, sure ...
But they don't have that.
The are throwing agents out there like they can handle this degree of complexity and nuance, when they cannot.
100% failure rate over any period of time.
The 'deliberate failure' is on Meta here.
>they will hook up actually important stuff to agents because its the future
>???
>400 dead
It frustrates me no end that so many folks are so ready and willing to accept the hype and lies about what this technology actually is or can do, when what it actually is and can really do is already amazing enough on it's own even without all the ridiculous AGI/ASI anthropomorphising bullshit. Falling into this ridiculous "machine-god" hype-cult is kinda holding this technology back from it's true full potential, as everyone's all busy doin' stupid stuff it's not really capable of doing well, or designed for instead of focusing on using it for the (many) things it is really really good at doing (various really useful and powerful language, vision, and audio related tasks).