16 comments

  • hn_throwaway_99 1 hour ago
    I thought this was great, and hilarious. Kudos to Opus 5, I thought it was the only one that came close to passing. Interestingly, I thought many of the failures drew the frog face OK, and they had some type of big blob for the jaw, so they knew "Hapsburg jaw" meant a protruding jaw, but it wasn't really connected to the frog face in any way that made sense.

    Small side note, the first gemini-2.5-pro one totally reminded me of some sad faced meme or Pepe the frog from somewhere. Anyone know what I'm referring to, tried to find it.

    • viciousvoxel 38 minutes ago
      Perhaps you're thinking of Salad Fingers?
      • qwertybased 17 minutes ago
        ~I'm leaning more towards the rage face poker face~

        edit: nevermind, definitely "monkey-puppet side-eye" vibe.

  • thebigship 1 hour ago
    Hi all, the site is getting hugged to death, thank you, was not expecting this kind of warm response. I will be working to make this more reliable, in the meantime, sign up for my newsletter: https://www.jaymollica.com/blog/

    also my favorite SVG was def the google/gemini-3.6-flash

    edit: ok better now I think

    • troupo 1 hour ago
      > also my favorite SVG was def the google/gemini-3.6-flash

      That looks like something from Machinarium or Robots :)

  • gerdesj 6 minutes ago
    Please could we have a human generated image to compare the AI generated tosh with?
  • wren6991 58 minutes ago
    Opus 5 clearly frogmaxxed.

    gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.

    • fasterik 40 minutes ago
      Arguably, a royal portrait is a misinterpretation of the prompt, since it's just asking for a specific facial feature. But I guess you could look at it as a bit of artistic license.
    • andybak 26 minutes ago
      > frogmaxxed

      raninemandibularprognathism-maxxed?

  • MiroslavPokorny 16 minutes ago
    My test is to ask AI to pick up all the rubbish at the beach.
  • evan_ 34 minutes ago
    Hopsburg Jaw
  • getnormality 1 hour ago
    This is a strong benchmark! None of these could be remotely mistaken for human art. Opus 5 comes closest.
  • dehrmann 1 hour ago
    How do models approach SVG generation? In one version, I imagine them actually trying to reason about them as an LLM. In another, I imagine something closer to a GAN.
  • linksnapzz 1 hour ago
    A friend’s favorite prompt is “Batman & Julia Child; in the kitchen laughing at a ham”. Sounds simple, but has been surprisingly tough.
  • rush86999 42 minutes ago
    It's opus 5 > Kimi K3 > grok 4.5

    That's a pretty good benchmark

  • kindawinda 13 minutes ago
    Why is this garbage on the front page?
  • leumon 1 hour ago
    Can you also try the new deepseek v4 flash?
  • epolanski 24 minutes ago
    Gemini 3.6 flash is crazy good.

    Would've wanted to see also DS4 flash.

    • gpvos 2 minutes ago
      Crazy funny, yes, but not good.
  • troupo 1 hour ago
    Mine is any variations on mammoths in various situations, or anthropomorphic. Since mammoths are invariably majestically going from one place to another in any of the books, models have hard time imagining anything but that.

    Also try a fantasy archer with a proper bow who is not brooding, sitting in a fantasy wood :)

  • sixtyj 1 hour ago
    For those who don’t know a Habsburg jaw also known as mandibular prognathism, it is a genetic condition characterized by a protruding lower jaw, which was notably prevalent among members of the Habsburg royal family due to their history of inbreeding. This condition often resulted in significant facial deformities and difficulties with eating and speaking.
  • thebigship 2 hours ago
    I think this one has advantages over the “pelican riding a bicycle” one because it hinges on an anatomical feature that many models associate with royalty, “habsburg” being a lineage and “habsburg jaw” being an anatomical feature.

    Seven of fourteen models silently imported royalty into a prompt that named only an anatomical feature. Two of them knew they were extrapolating ("because Habsburg") and did it anyway.

    Mistral returned byte-identical output across separate calls.

    Gemini narrates its work in 65 comments; Llama says nothing.

    If you're deciding which model to trust with instructions, "how much does it embellish beyond what I asked" and "does it behave deterministically" are directly practical questions.

    • n00bskoolbus 53 minutes ago
      The identical pair from Mistral took me off guard. Many of the other models were so varied between the runs which is more what I would expect.
    • HPsquared 1 hour ago
      Bite-identical?
      • thebigship 1 hour ago
        clearly I missed an amazing copy opportunity, thank you haha