Anthropic Risk August 2026 [pdf]

(www-cdn.anthropic.com)

37 points | by artninja1988 1 hour ago

10 comments

  • modeless 1 minute ago
    So as of a month ago their best internal model was one they don't plan to release but it is only slightly more capable than Mythos. I am surprised actually, I thought they would have a significantly more capable model by then. If they don't have one by now, then Chinese have almost caught up completely.
  • datadrivenangel 33 minutes ago
    "We believe our internal AI R&D efforts are significantly faster than they would be without AI assistance, but not yet by a factor of 2 (though we are uncertain and measurement is difficult)"

    So Anthropic thinks their productivity is not even doubled by AI. Interesting data point.

    • furyofantares 20 minutes ago
      > So Anthropic thinks their productivity is not even doubled by AI.

      I find it hard to imagine launching this criticism at a new technology.

      • jchw 4 minutes ago
        Now let's re-evaluate that based on how much it costs in both R&D and at runtime. This new technology has a lot of work to do to justify itself.
    • impossiblefork 13 minutes ago
      I think AI is best for ideation, experiments, small things.

      They've probably already settled on most of the architecture, so they're deciding big things instead of small things.

      The thing AI allows is for some ordinary person-- a PhD student, or similar, to whip up a miniature synthetic experiment that turns out to be horrid and needs to be fixed by hand, but which at least gave him a plot on the same day he had the idea. That's, I think, where AI shines: prototypes. Anthropic probably doesn't need that to the same degree as the small experimenter.

    • bonoboTP 28 minutes ago
      Well, it is a data point but AI R&D at a frontier lab is not really a representative stand-in for a regular workplace.
    • taosx 15 minutes ago
      So I can take that as ~infinite amount of tokens don't don't even get you 2x on any novel tasks? No auto-researcher, no rsi..

      Is that correct?

    • T0Bi 26 minutes ago
      AI R&D efforts != productivity. I think it's fairly obvious that SOTA research is less affected by AI than writing another boilerplate react frontend.
    • whateveracct 22 minutes ago
      So they did two years of work in one year? yeah right lol
      • semiquaver 5 minutes ago
        I mean, look at what they’ve released in the last year. I think most companies would be proud to have that done in three. Say what you will about Anthropic but they ship.
    • what 29 minutes ago
      >thinks

      They can’t measure even measure it, it’s just vibes. They may not even be more productive.

      • scj 18 minutes ago
        To be fair, there isn't a good method of measuring software development productivity in general.

        Maybe they should ask an AI to create one!

  • lwarfield 2 minutes ago
    > 6.2 [Appendix redacted] > This appendix describes the criteria for our blocking bioclassifier exemption policy, and has been redacted from the public version of this report for security reasons.

    >6.3 [Appendix redacted] > This appendix, redacted from the public version of this report, details the changes made to our constitution to expand classifier coverage to harmful uses in scope for the CB-2 threat model but not the CB-1 threat model, as described in Section 4.5.2.1.

    interesting...

  • int32_64 5 minutes ago
    Does anybody have any good reading on how the Chinese labs approach risk vs. the American ones?
  • 12ahGA 34 minutes ago
    AI companies are flooding the zone like Steve Bannon. Leave no one time to develop thoughts.
    • esafak 21 minutes ago
      If they leave Steve Bannon with no time to develop thoughts, I'll tip one out for them.
  • visiondude 33 minutes ago
    a mystery “model 2” is mentioned alongside mythos/fable.
    • merksittich 25 minutes ago
      > Model 2, which is somewhat more capable than Mythos 5. Our rough qualitative sense is that this model is a noticeable improvement on Mythos 5 for many tasks relevant to internal use but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview. We do not currently have plans to release this model externally, and have not run all of our typical suite of predeployment assessments, so we have somewhat lower confidence in our beliefs about its capabilities.
    • lwarfield 5 minutes ago
      > More capable than Mythos 5 in some areas, less capable in others; overall slightly more capable.

      This sounds like it might be a Mythos finetune for some specific task.

    • flyinglizard 24 minutes ago
      Meanwhile I can't really tell the difference between Fable and Opus for my tasks. I kinda think Fable does a better UX work so I keep using it for that because I couldn't be bothered to A/B them, but otherwise it's all the same and the model and effort are just feel good knobs I twist to still remain a load-bearing element. At least that's my honest take.
      • malexw 4 minutes ago
        For the past 2 weeks or so I've been doing the A/B test, sending identical prompts to Fable 5 and Opus 5 to test their ability to produce design documents for new feature work. I've consistently found that Opus 5 produces more complete, accurate and "imaginative" designs than Fable, often finding design issues or nearby bugs that Fable 5 misses. However, that creativity means Opus seems to hallucinate more, while Fable's design is clearly based on the actual existing code. Or as Opus put it: "I hedged — [Fable] checked."

        By pitting them against each other I get much better design work, and then I've been happy to hand off the design file to Opus 5 for implementation. But some of the assumptions Opus 5 makes leaves me wary of relying on it too strongly. This might be fixable by prompting it to ground its answers.

      • dataminded 21 minutes ago
        Fable was amazing during the first preview. Once they added it back, the limits are too low to get anything done. I might use it in chat if I remember to select it once a month but don’t even bother to try and code with it.
  • aquarious_ 9 minutes ago
    My friends and I, and the teams I'm a part of, just want to build and create fun, cool things. I am so tired of being preached to by Anthropic like they're some arbiter of 'ethics.' So, so tired.
  • bunkydoo 23 minutes ago
    [dead]
  • _ache_ 20 minutes ago
    It's crazy how Anthropic talks so much about their "AGI risk" and not enough about the risk of bankruptcy.
    • s1artibartfast 1 minute ago
      Are you surprised? Why would any private company spend time publicizing their financial risks?

      Seems like a strange expectation.