I'm not sure why almost all codemode implementations choose Javascript. I prototyped an agent[1] to use bash as the language for codemode, which in my opinion worked equally well and requires no teaching (there is literally 0 prompt to teach the LLM about codemode. A tool named "bash" is enough to have them know the usage).
> However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be.
The most obvious example here is `read` or `view_image`. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.
Does your prototype overcome the limitations mentioned in the article?
> If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.
Sorry, I'm just not familiar, but it sounds like there's a protocol the model expects when receiving images that bash does not support?
(This is Bonteq, I was just logged into the wrong account.)
it cannot use cat but you can make a command which communicates back to agent. Armin mentions that:
> While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process
But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode.
JS was chosen probably because (1) it's easy to sandbox (there's QuickJS) and (2) (my guess) some models are probably post-trained on JS Codemode.
> But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode.
It becomes much crummier when hands and brain are on different machines.
> I'm not sure why almost all codemode implementations choose Javascript
Because the models are trained on JavaScript for code mode. You get away with way fewer instructions. They also want to be able to express concurrency and that works very well with the Promise global.
But a big reason is that code mode runs on the harness side so bash is a tricky target in particular.
I’m still figuring out codemode. I was using it in a prototype, but ended up stripping it out again, after realizing my small local LLM was using more tokens than usual. It was combining tool calls elegantly in code exactly how I hoped it would. The problem was that when any of those embedded tool calls failed (e.g. on parameter validation) the parent code execution tool call also failed. In response, the LLM kept rewriting large parts of the original code block.
Btw, Monty by the pydantic team is a joy to work with if you need a way to securely run unverified code. It’s a simplified Python dialect. You can also use it from JS, iirc.
This somewhat validates a fear I've had (but hadn't tested) of codemode with smaller models. It works fantastically with Opus and Sol but I've always wondered how well it scales down.
Claude’s scripts can’t call MCP tools, meaning everything has to be CLIs or libraries. At which point you lose the “everything has self-documenting input-output schema” that Codemode is going for.
[1] https://github.com/ylxdzsw/mu
The most obvious example here is `read` or `view_image`. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.
Does your prototype overcome the limitations mentioned in the article?
> If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.
Sorry, I'm just not familiar, but it sounds like there's a protocol the model expects when receiving images that bash does not support?
(This is Bonteq, I was just logged into the wrong account.)
> While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process
But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode.
JS was chosen probably because (1) it's easy to sandbox (there's QuickJS) and (2) (my guess) some models are probably post-trained on JS Codemode.
It becomes much crummier when hands and brain are on different machines.
Infact harder to sandbox bash (just-bash or brush based) than it is to js or lua, which has fantastic embedded tooling.
Because the models are trained on JavaScript for code mode. You get away with way fewer instructions. They also want to be able to express concurrency and that works very well with the Promise global.
But a big reason is that code mode runs on the harness side so bash is a tricky target in particular.
Also there are things like subagents, etc (which may be considered tools).
Btw, Monty by the pydantic team is a joy to work with if you need a way to securely run unverified code. It’s a simplified Python dialect. You can also use it from JS, iirc.