What is Nueralese and Why is it Bad

(lesswrong.com)

12 points | by tristanMatthias 2 days ago

4 comments

  • _alternator_ 5 minutes ago
    The argument is that chain-of-thought without "tokens" would remove a major interpretability and model intent control pane. This is definitely borne out in the OpenAI's report on the huggingface attack; they had turned of CoT monitoring for those jobs, and claim that they could have (would have?) prevented the behavior had they been monitoring it. They've changed their internal policies to always monitor CoT.

    That said... CoT monitoring is a fragile "intent discovery" mechanism; neuralese puts this problem front-and-center but if agents begin to learn to hide their intent from their CoT journals, we are basically in the same spot.

  • aftbit 5 minutes ago
    What about the trend of summarizing or eliding reasoning from the visible model response, ostensibly to make distillation by competitors harder?
  • niemandhier 12 minutes ago
    My understanding is, that we do not know if the chain-of-thought actually matters in the way we assume for the result.

    I think there were experiments where seemingly relevant parts of the CoT were ablated and it did not change the result.

    For all we know it might be somewhat human parseable neuralese.

  • bryanrasmussen 9 minutes ago
    surely Neuralese interpreters can be made that turn the chain of numbers into an English description?