What Is Codemode

(lucumr.pocoo.org)

62 points | by Tomte 17 hours ago

9 comments

  • ssivark 6 minutes ago
    [delayed]
  • ylxdzsw 2 hours ago
    I'm not sure why almost all codemode implementations choose Javascript. I prototyped an agent[1] to use bash as the language for codemode, which in my opinion worked equally well and requires no teaching (there is literally 0 prompt to teach the LLM about codemode. A tool named "bash" is enough to have them know the usage).

    [1] https://github.com/ylxdzsw/mu

    • Bonteq 51 minutes ago
      > However bash has one fundamental limitation which is that it can only compose programs that run. And there are some things, which are not programs, but native tools to the LLM and they sort of have to be.

      The most obvious example here is `read` or `view_image`. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.

      Does your prototype overcome the limitations mentioned in the article?

      • andreypopp 37 minutes ago
        There's no limitation, just have bash commands `read` or `view_image` which communicate back to agent.
        • codybontecou 30 minutes ago
          But doesn't Armin mention the limitation?

          > If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.

          Sorry, I'm just not familiar, but it sounds like there's a protocol the model expects when receiving images that bash does not support?

          (This is Bonteq, I was just logged into the wrong account.)

          • andreypopp 23 minutes ago
            it cannot use cat but you can make a command which communicates back to agent. Armin mentions that:

            > While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process

            But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode.

            JS was chosen probably because (1) it's easy to sandbox (there's QuickJS) and (2) (my guess) some models are probably post-trained on JS Codemode.

            • the_mitsuhiko 16 minutes ago
              > But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode.

              It becomes much crummier when hands and brain are on different machines.

              • andreypopp 10 minutes ago
                Agree, but this orthogonal to JS or bash question. (I mean can always run some bash locally as well, even wasm compiled one).
    • anilgulecha 31 minutes ago
      Almost every model is fully trained on js. That does not need teaching either.

      Infact harder to sandbox bash (just-bash or brush based) than it is to js or lua, which has fantastic embedded tooling.

    • the_mitsuhiko 1 hour ago
      > I'm not sure why almost all codemode implementations choose Javascript

      Because the models are trained on JavaScript for code mode. You get away with way fewer instructions. They also want to be able to express concurrency and that works very well with the Promise global.

      But a big reason is that code mode runs on the harness side so bash is a tricky target in particular.

    • searealist 2 hours ago
      Can do you make a tool or mcp call from bash?
      • clintonb 2 hours ago
        Yes. Invoking an MCP tool is just an HTTP call. You can do it with curl.
        • searealist 1 hour ago
          That's one kind of MCP. Another is a local stdio server.

          Also there are things like subagents, etc (which may be considered tools).

  • lemontheme 33 minutes ago
    I’m still figuring out codemode. I was using it in a prototype, but ended up stripping it out again, after realizing my small local LLM was using more tokens than usual. It was combining tool calls elegantly in code exactly how I hoped it would. The problem was that when any of those embedded tool calls failed (e.g. on parameter validation) the parent code execution tool call also failed. In response, the LLM kept rewriting large parts of the original code block.

    Btw, Monty by the pydantic team is a joy to work with if you need a way to securely run unverified code. It’s a simplified Python dialect. You can also use it from JS, iirc.

    • hamandcheese 25 minutes ago
      This somewhat validates a fear I've had (but hadn't tested) of codemode with smaller models. It works fantastically with Opus and Sol but I've always wondered how well it scales down.
  • aidiveyt 1 hour ago
    subagents differ: in my claude code logs every subagent cache write is 5-minute tier, main session 1-hour
  • soltanov 2 hours ago
    Recovery after interrupted execution; distinguishing completed side effects from calls that can safely repeat.
  • injidup 55 minutes ago
    How is this different to Claude writing mini scripts to get jobs done which it does quite often?
    • odo1242 46 minutes ago
      Claude’s scripts can’t call MCP tools, meaning everything has to be CLIs or libraries. At which point you lose the “everything has self-documenting input-output schema” that Codemode is going for.
  • haukebri 1 hour ago
    [flagged]