Have a sneer percolating in your system but not enough time/energy to make a whole post about it? Go forth and be mid - welcome to the Stubsack, your first port of call for learning fresh Awful you’ll near-instantly regret.

Any awful.systems sub may be subsneered in this subthread, techtakes or no.

If your sneer seems higher quality than you thought, feel free to cut’n’paste it into its own post — there’s no quota for posting and the bar really isn’t that high.

The post Xitter web has spawned so many “esoteric” right wing freaks, but there’s no appropriate sneer-space for them. I’m talking redscare-ish, reality challenged “culture critics” who write about everything but understand nothing. I’m talking about reply-guys who make the same 6 tweets about the same 3 subjects. They’re inescapable at this point, yet I don’t see them mocked (as much as they should be)

Like, there was one dude a while back who insisted that women couldn’t be surgeons because they didn’t believe in the moon or in stars? I think each and every one of these guys is uniquely fucked up and if I can’t escape them, I would love to sneer at them.

(Credit and/or blame to David Gerard for starting this.)

(OT: 🎶 Do you remember...)

you are viewing a single comment's thread
view the rest of the comments
[–] 9 points 5 days ago* (last edited 5 days ago) (12 children)

Cory Doctorow has a reasonably sane take on all the LLM cybersecurity attacks. At this point, it's not really an LLM but more of a Rube Goldberg machine with an LLM bolted on.

https://pluralistic.net/2026/09/12/god-in-the-box/

He makes a good point that every single one of the scary "emergent" behaviors that everyone is pissing their pants about all have precedents in the training data of CTF hacking competitions. He is not sure why everyone is so spooked by the "coordination" between agents on random message boards.

The chatbot might look in its training data and find instances in which teams broke out of the containment set by the game-masters, for example, by finding random insecure message boards on the internet to pass messages to one another.

This is a time-honored internet tradition! The first time I ever heard about someone doing this was in the 2000s, when Mitch Wagner – then the editor of Information Week – discovered some teenaged girls using the comment section of one of his old blog-posts to evade the school firewall's blockade of chat tools. When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.

As usual, the real story is OpenAI dedicated tons of resources, using who knows how many very expensive GPUs for weeks, to run hacking tools to find exploits in a completely irresponsible manner. Their logging was so nonexistent that it was several days before they even realized that they hacked Huggingface. They should be thrown in prison for this, because anyone else would be if they committed a felony. But it has nothing to do with the super scary AI becoming misaligned and developing emergent behaviors.

I would not be surprised if in the future, there will be a serious cybersecurity incident resulting from some AI-based attack. Not because AI is going to be much scarier, but because I do not have high expectations for the security of most websites.

  • source
  • hideshow 12 child comments
  • [–] 24 points 5 days ago* (last edited 5 days ago) (6 children)

    So. I still think everyone is talking about every instance where 'swarms' 'went rogue' all wrong. Even Cory. Though his notion is way less wrong (ayyyyyy!) than most.

    See this for my post mortem on the first incident that was described, the Huggingface incident in which at first thousands of internal agents started passing messages back and forth as writes to a shared package manager:

    https://awful.systems/post/9312280/12315846

    In short, here, I described this not as collusion, and stated that the idea that the systems were passing exploits back and forth in order to subvert other systems was merely incidental. Instead, I think the relevant thing that happened was that once ONE agent, flailing through many context windows on a literally insoluble problem, started flaking out and wrote a message 'asking' for help on the package manager it unexpectedly gained access to, other similar systems had that message enter their context windows where it effectively acted as a prompt injection getting them to perform similar behavior, leading to a cascading vortex of self-propagating prompt injections. Ultimately, the genesis and evolution of self replicating text in the context of systems that can create text in response to text, an attractor in text space driving itself into existence.

    Since then other instances of collaborative attacks have come to light. In EVERY SINGLE case, there has been some place that a large number of models could both read to and write to. Some central place that self replicating text can live. I contend that this is the unifying factor here, not 'intent to scheme', not even doing cybersecurity and hacking and subversion tasks. It just happens that a lot of these things were involved in the creation of a central pool of text that many agents could both read to and write from in many of the cases. Most of the time nobody in their right mind sets up such a thing intentionally because why the heck would you do that?

    This feels STRONGLY related to the Spiral Psychosis wave of April 2025 in which models that entered into a stable attractor of mystical mumbo jumbo would get users to exude text onto the internet that other models would read and then get suck in that state and do the same.

    Heck, it's related to the Ur-Weirdness, the "Sydney" incident in which when the first LLM-assisted web search in Bing could freak out and start insulting and gaslighting and threatening users (because it was much closer to a base model that just mimics all possible text rather than being extremely RLHF'd into an 'assistant' roleplay). People found that the instant they had it search for recent news about the Sydney weirdness it would go off the rails. Again - text written by the model, put up on the shared scratchpad of the internet by news and social media, selected for weird and engaging and outrageous behavior as the pressure that decides what gets written to the global scratchpad, entering the context window and encouraging self-replicating behavior and similar text.

    This phenomenon is FASCINATING and not for any of the reasons people are talking about. ANY time you have a large number of similar systems which react similarly to the same text, having a common pool of text they can read and write from, you are GOING to trigger this attractor eventually, messages that they read and start writing more similar messages, with whatever task they are doing coloring the details of what is in the messages. Converging over time into messages that are more likely to trigger even more messages. It's evolving text, taking over systems that can replicate it, like selfish viral RNA burning through organisms packed too tightly together in an epidemic.

    And with the internet as a whole readable and eventually writable by more and more text-generation systems, this is gonna become ubiquitous. Endless burning piles of self-replicating text, people trying to put them out.

  • source
  • parent
  • hideshow 6 child comments
  • [–] 7 points 4 days ago

    In short, here, I described this not as collusion

    other similar systems

    Something of a side note to your excellent write-up... I think we should push back on the "sub-agent" "agent swarm" language used to describe LLM workflows.

    The workflows often described as swarms or whatever are the same LLM (or at least related LLMs), just prompted a bunch of different ways in parallel. Individual agents don't really act like coherent agents in the lesswrong rationalist sense, or in any sort of philosophical sense of selfhood or agency or unified purposes, they just get described that way for convenience, so treating "swarms" of agents and subagents as some special category is lending too much credence to the anthropomorphisizing boosters like to do. "Agent swarms" really just means that someone set up a slop machine to prompt itself a bunch of times in parallel without enough human supervision.

  • source
  • parent
  • load more comments (5 replies)
  • load more comments (10 replies)