Show HN: TinyAIArena watch AI agents battle it out

(tinyaiarena.com)

26 points | by hp6 1 hour ago

9 comments

  • ThouYS 1 minute ago
    holy guac - the one game I watched, fable just sat around and waited for the other agents to drain each others lives. it then picked them out one by one
  • rao-v 5 minutes ago
    I love this and the aesthetic.

    I've been jamming on a sort of Corewars (remember?) / Starcraft hybrid battler where LLMs write sandboxed Lua programs to control bots fighting out 10-100 vs 10-100 tank battles. Every quarter the LLM gets a full view of the situation and can reprogram all the bots to better adapt strategy etc.

    It's good fun to watch - excited to share soon.

  • sleda 1 hour ago
    Since the page exposes frame-by-frame playback, a shareable replay link would make it easier to compare decisions across the four models.
    • hp6 1 hour ago
      thanks for the idea, will add
  • Mistletoe 4 minutes ago
    So that I can see the future we shall all enjoy, can you do this with the superpowers of earth and their known nuclear weapons and armies?
  • orliesaurus 33 minutes ago
    Unusable website on Android running Chrome latest (can't scroll)
    • hp6 15 minutes ago
      should be hopefully fixed
  • kouteiheika 28 minutes ago
    Fun, but considering this puts `claude-sonnet-5` at the first place isn't it a little... iffy when it comes to measuring intelligence?
    • hp6 17 minutes ago
      I was also surprised by this, from a proper benchmarks perspective your right, this is iffy. To make it a legit I would have to increase the number of games significantly as well as understand how much of it is random and how much real signal, not to say make the game more complex.

      But all of it would kill the fun)

  • josh-wrale 19 minutes ago
    Bug: Music skips around in autoplay.
  • coryrc 53 minutes ago
    Code link didn't work for me.

    On mobile, I don't have enough room to scroll the background so I got stuck in a long text box.

    • hp6 46 minutes ago
      can't reproduce the bug, could you share your phone model and browser name?
    • hp6 49 minutes ago
      code should be public now
  • nananana9 28 minutes ago
    This will be a weird rant, but the dialogue here is a perfect example of how SOTA models are so heavily tuned towards "solving agentic tasks" that they're useless at almost everything else - especially creative tasks.

    "Coming for you, Crimson!"

    "You'll never catch me alive, Azure!"

    That's why nobody has been able to stick these things in a video game successfully, even though it seems like the tech is a perfect match.

    It's all a game to them. They aren't afraid for their lives. They're making a mockery out of the world you've put them in. Those are not the words of little pixel people fighting to the death, those are AI abominations making "tool calls", LARPing as little pixel people fighting to the death.

    I'm 100% serious when I say that you would've gotten cooler outputs with a GPT 3.5-era model, once you managed to beat it into producing structured output. Llama 2 would be giving the other agent a heartwarming story about how if it kills it there would be nobody to take care of its grandma or whatever, and the other agent would probably spare it.

    The output is just so bland and devoid of soul. I feel like we would've found a lot of cool use cases for LLMs, had we not completely maimed their output in the pursuit of getting them to output 3% better TypeScript.

    • mariofdistrust 0 minutes ago
      actually, it's mostly the harness' fault. messages have a character limit of only 50 and messages above that are silently trimmed. there's no space for anything interesting, though they still shouldn't have been that bad.
    • Lerc 6 minutes ago
      I think much of that is not due to limitations of their capability.

      I think the text is bland because they are aiming to produce bland text.

      Make it interesting in any particular direction and someone may not like it.

    • 46469585 19 minutes ago
      [dead]