Why we write our own C and C++ inference engines

(localai.io)

33 points | by eatonphil 2 days ago

6 comments

  • dennis16384 2 hours ago
    I had a similar success with Model2Vec static embedder and NER inference (both GGUF, compiled for WASM), ported to plain C from ONNX Runtime.

    Wasm size from 30Mb to 300kb and 1.5x speedup. It's definitely worth it for performance or distribution size.

  • scottcodie 2 hours ago
    I did took a native c++ approach when writing a relational transformers engine (RelativeDB). My journey was pytorch -> c++ -> Triton (lang). While C++ was more performant than Triton, I couldn't afford to optimize on every gpu. I just accepted the ~15% throughput loss for my cloud service, which honestly wasn't bad for the amount of flexibility I got out of it.

    But the cpp port of vllm looks great, that'd be great if you'll maintain that. I hit the same limitations with vllm.

  • stephbook 3 hours ago
    Should have started with writing your own blog posts.
    • lelanthran 53 minutes ago
      > Should have started with writing your own blog posts.

      While the page looks vibe-coded[1], the content itself does not have any AI tells. What are the tells you are seeing?

      [1] Too many sites I find on HN frontpage these days slow my PC to a crawl. I assume they are all using the same autogenerated HTML, Javascrip and CSS to make animated backgrounds :-( On this specific site scrolling is laggy.

    • pjmlp 24 minutes ago
      Same could be said for all that talk about having Claude do their work.
    • winter_blue 1 hour ago
      I found the post insightful and interesting. I'm not sure it was written with AI assistance, but even if it was, I don't see that as a reason to dismiss it. For what it's worth, I spend hours everyday reading AI output and summaries.
    • nnevatie 2 hours ago
      Came here to say the same. Really tiring to read these slop-infested posts, where everything has the “right shape”.
    • altmanaltman 2 hours ago
      I went through the post because of your comment but it really doesn't look like AI slop. Can you please share why you feel like its slop and not written by a human? I can also say "should have started writing your own comments" to you and its unfalsifiable. Blanket accusations with no proof is not a good move really.
      • nnevatie 1 hour ago
        The post is full of signs. Here's only a couple of examples:

        > The method, the measurements, and what it costs us.

        > That is the general shape of these wins.

        > Parity is the gate, speed is the follow-up

        I could go on and on, but you probably get the point. If you don't find anything funny with the above, you might have not been enough-exposed to slop.

        • wannabe44 55 minutes ago
          It's always hyping up something and throwing punch lines in every sentence. Normies love this shit.
          • nnevatie 50 minutes ago
            Yes, it’s basically business-as-usual but on speed.
  • piterrro 37 minutes ago
    Could this vllm port be faster to install? Im starting gpu machine multiple times a day and it takes 5 minutes to set vllm up. If Inise this port that time is minimized?
  • adithyassekhar 2 hours ago
    What you get: X is the A, Y is the B.
  • federicoTXTS 1 day ago
    [flagged]