8 comments

  • VulgarExigency 57 minutes ago
    > The first attempt at recovering the key was wrong in a very specific way; it produced a working key and the signature check passed, but a hash the binary computes as an integrity check didn't match. In my experience, most models would have called it done and left it at that, but Qwen 3.8 27B didn't do that. Instead, it highlighted the mismatch, went back to the drawing board, and kept going until the value matched byte for byte.

    This seems to be a pattern in the more recently released models that I think accounts for an increase in the quality of their work. They are very persistent in verifying that their work is actually correct, so even if they're not as "smart" as bigger models that get it right the first time, they have the ability to follow through to ensure that the work is actually done.

    • dataviz1000 32 minutes ago
      > They are very persistent in verifying that their work is actually correct

      I visualized this with Qwen 3 4B [0] and Sonnet [1] so that people can very easily grok what you mean by "the ability to follow through." What is key about the reasoning tokens is that they will form a pattern of verification tokens in the sequence which, has been shown, will develop in RL training purely without the supervised fine tuning (SFT) which is used to make the output human readable.

      [0]https://adamsohn.com/reasoning-grid/

      [1] https://adamsohn.com/lambda-variance/

    • catlifeonmars 33 minutes ago
      Your explanation makes sense, but I don’t know how you can make a sweeping generalization about more recently released models without qualifying it.

      This is says more about humans tendency to pattern match than anything else.

      X works better than Y only is only a useful observation if we are using the same X and Y in a similar context, with similar parameters. Kind of goes out the window without it and I think this is part of why people have such vastly different opinions about the same technologies. We’re all talking past each other.

    • criemen 32 minutes ago
      I believe this is part of the complaints of new models taking longer/requiring higher spend - they go the extra mile on verification, regardless of whether their change is correct already or not. So on problems that an earlier model one-shotted an answer to and did some lighter verification, the newer models might take longer to come back to the user due to running all the tests for your software they could find.
    • braiamp 49 minutes ago
      Well, it seems that Linus doesn't use those:

      > And this was a debug session from hell, enormously helped by an AI doing much of the grunt-work.

      > I'd like to call it my tireless helper, but the AI several times stated flat out that this was impossible and unsolvable and that we should just write a report about it.

      > I suspect those things have been trained by people who may not be quite as stubborn as I am.

      https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...

  • exceptione 49 minutes ago
    Local models would be even better if they did not ship with all the refusal shenanigans built-in. You can safely bet organized crime has access to the best models without these hoops, which makes the case that the average user (=non-criminal) should have access too. As I understood from an ex-Anthropic employee, some orgs got access to Mythos based on their high enough spending level, not on other grounds.

    Either we are in command over the software, or the corp is in command over us via the software. I can on a theoretical level understand the concerns, but either we ban all LLMs or we have a level playing field for everybody. Let's not forget: defense and offense are different sides of the same coin in software. I guess this wouldn't apply to bio weapons, but I am not in the know about that.

    • ninahaberl 23 minutes ago
      I’d expect these shenanigans to get much worse over time for the average Joe.

      Imagine a world where any random person can run a super-capable model on their own hardware with no limitations and no one to pull the plug.

      Information has always been power and those who already have power won't just allow everyone else having the same tools as them

    • TofuLover 16 minutes ago
      Completely coincidentally, we're just about to launch a service that does exactly this (API access to uncensored open models)! We have a waitlist at the moment but will be live very soon!

      https://violentdelights.ai

    • dantudor 35 minutes ago
      There are versions of Qwen3.8-27B that are unrestricted and available from hugging face.

      "It will comply with harmful, unethical, offensive, or illegal requests that the original Qwen3.8-27B would refuse. It has no meaningful built-in guardrails."

      • radlad 31 minutes ago
        > What makes this build different is the word before FP8: uncensored. We applied abliteration — orthogonalizing the refusal direction out of the residual stream — to remove the model's safety-alignment refusals. The result is a model that will comply with requests the original would refuse.

        Surely this has unintended side effects on output quality?

        • andsoitis 18 minutes ago
          > > What makes this build different is the word before FP8: uncensored. We applied abliteration — orthogonalizing the refusal direction out of the residual stream — to remove the model's safety-alignment refusals. The result is a model that will comply with requests the original would refuse.

          > Surely this has unintended side effects on output quality?

          What's your reasoning for why that follows?

        • DiabloD3 18 minutes ago
          It does depending on the technique.
        • miroljub 18 minutes ago
          A bit worse quality is a fine trade off when the alternative is no output (zero quality).
          • radlad 15 minutes ago
            On censored inputs only.
    • ramon156 21 minutes ago
      heretics and manual iterations get you very far to the point where i have ethical questions about whether this should be possible
    • binary132 43 minutes ago
      Ehh, it’s at least given as the excuse for gain-of-function bioweapon research
      • datsci_est_2015 26 minutes ago
        Digression, but this is the real Great Filter imo, not AI. I think technology advances to a point where it only takes one or two bad actors to type the right prompt to get a recipe for civilization-destroying bioweapons before you get anywhere near true AGI or anything relevant to the Kardashev scale. Biology is fragile.

        But not that that’s a good justification for hamstrung models. I think it’s just the inevitable endgame and it’s more sad than scary

        • jeremyjh 19 minutes ago
          I have the same concern. If it becomes possible to engineer Captain Trips with a budget in the low 8 digits it won’t really matter what else happens.
  • saidinesh5 57 minutes ago
    Lately I genuinely believe that the future will be large frontier models generating and updating inputs/skills for "good enough" local models to solve our daily problems.

    A lot of tasks which need a bit of intelligence don't really need that much compute. Just good enough documentation / skills, tool calling and a good enough local model.

    Not sure what exactly this means for all those data centers that are getting built... But exciting times.

    • mark_l_watson 0 minutes ago
      Yes! My main use of very strong models is in writing my own coding harnesses for small local models, tailored for my needs. I also use very strong models to get much smaller skill files and also writing tools for my harnesses.

      re: data centers: pump and dump. Wealthy investors will have made their money and walked away, and the corrupt democrat and republican politicians in Washington will, as usual, protect the interests of the ultra wealthy and leave the general public to pay for poor decisions. There will be a government bailout.

      Anyway, on a positive note, I am all in for small local models that are augmented by strong hosted models for specific tasks. Use technology to help people, not make billionaires even mire money.

    • Tepix 22 minutes ago
      With AI being more useful with access to more of your data, I can't see myself using cloud AI models for purposes such as personal assistants.

      Perhaps with differential privacy or confidential compute...

      But ideally these models run locally.

    • catlifeonmars 30 minutes ago
      What’s the fundamental difference between a frontier model and a local model anyway?
  • pi-victor 27 minutes ago
    i'm not good with paper work, in fact, i'm horrible with anything that's paperwork related. for the past few days, i ran this model on my rtx 4090 + rtx 3070 and told it to check all the bills, invoices, contracts for me and my small company. i used pi with llama and the pi-llama plugin. oh, boy - i hooked it to my email, told it to download all of the invoices and bills i had for both me and my company and organize them by company/date/ and then merge them with the ones i have locally. it did ocr, wrote scripts, organized everything neatly. i am now the most organized i've ever been in my life. Next: RAG on all the documents and bills i have. if you connect staan-search (there is a pi plugin for that) and ctx7 to this it almost does miracles. the downside is i have to sit next to my noisy threadripper as the magic happens and pay for the electricity, but that's about it, i'll gladly do that. and as i finished this paragraph, it also finished organizing all my personal documents on my san. i don't use the expression "game changer" easily, but it's hard to resist in this case. out of all the models i've used locally qwen3.8:27b blows everything out of the water.

    my setup

    # Logical CUDA0 = RTX 4090, logical CUDA1 = RTX 3070 export CUDA_VISIBLE_DEVICES=0,1

    cd ~/projects/misc/llama.cpp/

    exec ./build/bin/llama-server -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M --mmproj /xx/xx/xx/xx/xx/mmproj-Qwen3.8-27B-Q8_0.gguf --host 0.0.0.0 --port 8080 --jinja --parallel 1 --split-mode layer --tensor-split 6,1 --fit on -fa on -c 98304 -ctk q8_0 -ctv q8_0 --image-min-tokens 1024

    i load more on the 4090 because it's faster.

    usually the temp stays around 65 for both. utilization for 4090: 70-90% 3070: 30-50%. I get around 30-40 tk/s. if i offload more to the 4090 the tk/s goes up, but i stress the card too much and that thing now is worth its weight in gold.

    note: the pi-llama plugin needs a patch for pi to send the model vision capabilities, seems it doesn't work out of the box.

    • Tepix 24 minutes ago
      I hope you have backups.
  • topper00_raptor 39 minutes ago
    Why does the screenshot on your pi terminal shows opus-4.6-medium from your claude subscription ? Instead of Qwen ?
  • jchw 59 minutes ago
    I'd personally like to know more about what tools it used/wanted and the harness setup, because this sounds pretty cool. I have a dual Arc Pro B70 setup and currently get around 22 t/s which isn't great but isn't terrible either (it is at least less quantized.)

    I've seen GPT 5.6 Sol happily invoke objdump and even write jobs to run headlessly which Ghidra when trying to disassemble a binary.

    • trollbridge 50 minutes ago
      My M5 Pro gets around 12-15 (6 bit MTP), although I haven’t worked on optimising it at all yet.

      A nice thing about running locally is you can run an uncensored model and you don’t have to worry about TOS violations on your OpenAI account when you ask it to “reverse engineer this ancient router firmware and give me a licence key that will work on it”.

      • medler 27 minutes ago
        Qwen is very much censored. Just try asking it about Tiananmen or how to build a bomb. But it is nice that you can experiment with it locally without having to worry about your account getting nuked
  • RobertasTa 1 hour ago
    Nice writeup — and the 30-minute static-analysis marathon is exactly the kind of task where this model's reasoning earns its keep. Complementary datapoint: I spent this week measuring the thinking levels of the same model locally (27B Q4, Ollama). On a concurrency-bug prompt, thinking off gave a complete answer in ~36s, low/medium thought for a few thousand tokens and answered fine, but high/max spent the entire 16k token budget thinking and never produced an answer at all. So the effort knob is real — deep analysis like this article's job wants it high, but leaving it high for everyday coding just burns context.

    Two things surprised me along the way. The defaults differ by stack: llama.cpp's chat template defaults to xhigh while Ollama lands closer to medium, so how much Qwen "overthinks" partly depends on your runtime. And the model card recommends different sampling per mode (temp 1.0 thinking vs 0.7 non-thinking with presence penalty), which almost nobody adjusts when toggling thinking off.

    Funny detail: with thinking fully off, my agent harness compensated by just running more tool calls — and still landed the correct fix.