Ox-Alpha Is GLM?

(dejan.ai)

71 points | by jitbit 14 hours ago

13 comments

  • gvkhna 2 hours ago
    If it’s not zhipu then why is it returning errors that zhipu does for other models? Who else would return the exact same errors even if they took a lot of core infra like tokenizer from z?
    • e9 1 hour ago
      Someone could've trained model on top of GLM. Same way Cognition trained their SWE model on top of Kimi and Cursor did same with their Composer model.
      • gvkhna 1 hour ago
        While possible the amount of variation in serving infrastructure is unlikely to land with actually giving the exact same errors zhipu does.

        It feels like glm flash, and there was a report zhipu had secured a huge new cluster suggesting they have the capacity. My guess anyway.

        https://www.tomshardware.com/tech-industry/artificial-intell...

      • vitorgrs 43 minutes ago
        The reasoning levels are the same as GLM 5.3. GLM 5.3 is still not open...

        I believe it's GLM 5.3 Flash or Air.

        • weiran 32 minutes ago
          Reasoning levels are often just injected system prompts so not a great way to fingerprint models.
    • ggcr 35 minutes ago
      Ziphu has that many resources to be able to serve capacity for 1 quadrillion tokens per day on Nous portal? My bet is that it's a Composer model from Cursor running on xAI cluster, they already did a Composer based on Kimi-K2.5
      • LaurensBER 11 minutes ago
        There's three options here:

        - The provider has a massive amount of (unused) hardware. Google or Cursor seem most likely

        - The model is extremely efficient, beyond anything we've seen so far

        - Whomever made the model has improved the cache efficiency in such a way that it's very cheap to serve. See e.g Deepseeks or Xiaomi caching (pre-price increase)

        • re-thc 5 minutes ago
          Option 4: the claimed capacity is not true. Real world usage hasn’t reached anywhere close to it.
  • swiftcoder 33 minutes ago
    > How many words are in the previous message?

    Its amazing to me that providers haven't added any sort of masking of the prompt in the thinking traces to avoid prompt extraction via this sort of trivial attack

  • pijalu 16 minutes ago
    My bet: it's Google running a "new" model based on GLM
  • jerrythegerbil 2 hours ago
    As someone who uses NCD nearly every day, I have concerns about how it’s been used here.

    But while we’re “guessing”: Xiaomi MiMO

    • walrus01 1 hour ago
      Have also seen people guess it's a next version of Longcat, but I also think that's unlikely
  • tadkar 1 hour ago
    I wonder if the NCD metric says something about distillation too. Would you expect that a model that has been distilled/seen traces from other models would have a smaller NCD? It would be really interesting to see if this holds up and provides evidence of distillation or certainly evidence of model outputs being used in the training mix.
  • volf_ 13 hours ago
    GLM 5.3 and all previous models don't have a vision encoder and can only accept text. Ox-Alpha can accept video and images, so unless Z-ai added a pretty good vision encoder for this model, I don't think so.

    My money is on Moonshot and this being Kimi K3.5. The measured tps and latency is in-line with K3's tps and latency from Moonshot.

    MiniMax M3.5 is also possible (but the MiniiMax provider is a lot more performant than the lab behind ox-alpha, so less likely).

    • nylonstrung 11 hours ago
      It would be stranger to me that Kimi switched to GLM's tokenizer than that GLM added multimodal like Kimi and Deepseek both did recently
    • minimaxir 2 hours ago
      The other tell from the provider angle is capacity. Whoever is hosting Ox Alpha has a lot of capacity which narrows down a lot of the Chinese companies.
    • Bolwin 13 hours ago
      Glm had made vision models in the past. Look up GLM 5v.

      The only question now is if it's 5.3v, 5.4/5.5 or a dedicated flash/vision model

      • BoredomIsFun 1 hour ago
        GLM made pretty decent for that time small 9b vision model, GLM-4.1.
      • volf_ 12 hours ago
        Yeah. It could be. The Z.ai DC latency is still ~1.2s faster than whomever is serving this model.
    • Almondsetat 13 hours ago
      DeepSeek literally just came out with the vision-enabled version of Flash v4 which was purely text based. Why would GLM not be able to do the same thing?
      • volf_ 12 hours ago
        It's possible
    • TiredOfLife 36 minutes ago
      But Moonshot limited signups because they lacked compute.
  • mogili 2 hours ago
    It's not a good model tbh, got a bunch of things wrong that Opus corrected in my codebase.
    • gvkhna 36 minutes ago
      The harness is making a big difference, lackluster performance with pi but somehow very good performance with opencode. There’s some rl there for sure, for a smaller model it’s likely going to perform much better in a harness it understands the best.
    • petesergeant 1 hour ago
      Yet to find a model that cross-model review doesn’t find a bunch of things wrong with. I’m running simultaneous review with whichever of Grok4.6/GLM5.3/Fable/Sol didn’t write it, and each model tends to find items the others didn’t.
      • mogili 1 hour ago
        Wasn't just a review, it failed the task I gave and Opus completed the task
      • epolanski 1 hour ago
        If your changes are non trivial even the same model will loop over and over with the feedback.
  • xorgun 2 hours ago
    Dont rule out ssi
    • nullbio 2 hours ago
      That would be insanely disappointing.
  • try-working 45 minutes ago
    Nvidia
  • ChrisArchitect 13 hours ago
  • petesergeant 1 hour ago
    I think within 12 months we’re going to see a frontier (inc open models) that’s so good at almost all human-directed tasks that which model you use just won’t matter. Only differences that remain will be in deep research or very long-range tasks.
    • stingraycharles 1 hour ago
      People were saying this last year, and they’ll be saying the exact same thing next year. The goalpost keeps moving.
      • Tepix 1 hour ago
        It‘s already happening, people are using cheaper models because they are good enough
      • petesergeant 1 hour ago
        Someone else having been too early on a prediction has little bearing on my prediction. A year ago almost nobody was using open models as daily drivers, today they are. When I run out of Fable and Sol credits in a week, I switch to GLM5.3, and it's not quite there, but it's good enough for productive work.
  • behnamoh 2 hours ago
    [flagged]
    • walrus01 1 hour ago
      I wasn't aware that an inanimate piece of software run by a corporation can be 'doxxed'. Totally inaccurate use of the neologism.
    • minimaxir 2 hours ago
      You cannot "dox" an AI model.

      Given the traction the model has received, it is extremely newsworthy to know who's developing and hosting it.

      • behnamoh 2 hours ago
        my question is: how does that affect a company's strategy? it's not like management is gonna switch models soon as a new shiny one drops. entire workflows depend on specific models working the way they do; you can't just swap out models.
        • minimaxir 2 hours ago
          If it's a really really good model, then yes, people will switch as long as the price is right. Ox Alpha is looking to be a really really good model to the point that it competes with Fable/Sol, and will likely beat them on price.
        • volf_ 1 hour ago
          same business model as crack. best way to get people hooked is to make the first hits free.
    • JSR_FDED 1 hour ago
      You can learn from the variety of techniques they used to come to this conclusion.
    • tjwebbnorfolk 1 hour ago
      > so much time on your hands

      Yet here you are, reading AND commenting about it