• 0 Posts
  • 90 Comments
Joined 2 years ago
cake
Cake day: May 20th, 2024

help-circle

  • Large amounts of the LW commentariat cannot get their heads around how students JUST. DON’T. GET. AI.

    https://www.lesswrong.com/posts/ySXuvJcqRindQwAk7/how-my-students-think-about-ai

    They are extremely surprised by:

    Perspective #1: There has not been rapid AI progress. My students do not have any intuitive sense that there has been rapid AI progress in recent years or really have much of a framework for thinking about that issue… With the exception of image/video generation, GPT-4 could do most of what they were looking for from a chatbot. Three years is a long time in their world, and their sense is that chatbots have been a mature technology over roughly that amount of time. They are an imperfect technology — students are well-aware of hallucinations — have been one, and will continue to be one… Historical context I gave them made this worse. The basic sentiment here was “AI was superhuman at chess a decade before we were born, and this is all they’ve done with it since?”…

    Perspective #2: Impressive progress or not, AI is going to wreck their lives, the economy, and the social contract.… Corporate leaders are always looking to get rid of workers, even if it is irrational to do so. Various motivations were posited here (hatred of the working class, FOMO, a preference for technology, machines can’t go on strike, etc.) but many of them think that a CEO would ultimately choose to pay twice as much to get AI to do a task half as well… AI will do substandard work that makes products and experiences worse but is capable of just barely scraping past the bar of minimal functionality (many of them independently brought up constantly malfunctioning self check out machines). This will be rapidly rolled out, enshittifying most things.

    Perspective #4: Catastrophic/existential risk arguments are sci-fi distractors from the urgent social/economic/political problems associated with AI. My students have a fairly strongly held view that “rogue” AI does not represent a real threat… They also mostly think that, if one does believe that AI is existentially risky, then the strategic interaction is not prisoner’s dilemma or even chicken but rather just a game theoretically boring setup where you die if you defect. Mash these together, and you end up with the view that expressed concerns about existential risk in the AI industry can’t be sincere… Students (both independently in written work and then later in group discussion) hypothesized that this might be a deliberate rhetorical choice to distract from present or immediately foreseeable harms from AI by directing attention towards a sexier but entirely hypothetical scenario… they see discussion of “rogue” AI as an attempt by the companies to divert blame (and perhaps legal liability) away from themselves as if Ford made a car with faulty brakes and then tried to blame this on “rogue cars.”

    Perspective #7: The Hugging Face Incident (summer students only) I described the Hugging Face incident to students in my summer course. None of them had heard of it beforehand. Their basic reaction can best be summarized as “OpenAI told a model to do some hacking and then it did some hacking. And?” None of them understood this as representing any kind of meaningful misalignment, nor anything particularly interesting.

    Perspective #8: This is definitely a bubble and it’s about to pop. No one had heard about Hugging Face, but a third or so of the summer students had heard about the Situational Awareness meltdown and several brought up Michael Burry. There was near universal consensus that we are in a bubble, it’s about to pop, and everyone will look very silly.


  • Oh MAN… quote:

    Throughout this project, I was using Fable and Sol with very little restraint and under the (very affordable to me) $200/month subscription plans.

    [If you check out my Straude, I probably used $5,000 of tokens, but the actual subscriptions are $200/month.]

    Video is much more expensive: at Gemini Omni Flash’s current price of $0.10/second, it would cost $645.40 to generate this movie, but that would balloon massively to $5,146.10 when you add in the 45,007s of footage I generated which didn’t make the cut [The full details are that the final cut is 6,454s of runtime, made from 849 clips and the rejects are 5,509 clips with 45,007s of runtime, which means there are 6,358 clips in total with 51,461s of runtime. Thus only about 13% of the clips I made were included in the final cut.]

    Companies are STILL subsidizing video generation to hold people’s interests.




  • I expect no-longer-obviously silly movies to be doable within two years, right now it’s going slower than I expected in January 2026 because the frontier labs are no longer competing over having the best video generator, like they were when we had Sora2 and Veo 3.1.

    And WHY are they no longer competing over this? Tell me. Could it be that its incredibly expensive to make somthing that nobody wants except fraudsters, and that further improvement via the kind of ML we do just produces exponentially diminishing returns for ever increasing huge amounts of effort and they realized it’s not worth it?

    Just like they will never understand that the openAI ‘pause’ for ‘security and alignment’ is a convenient excuse for the fact that they have no goddamned money and are on a treadmill to oblivion.


  • I am mildly fascinated by the types of issues that appear and the contrasts with other things that can locally in space and time look right.

    An object doing nothing will have a consistent surface and show perspective relative to the viewpoint. The systems are able to have representations of surfaces of different types and how they can fit together within an object. I am of course talking about within a given generation, not between generations where its utterly unsurprising that consistency is very difficult or impossible.

    But multiple objects in interaction with each other do very strange things. Their relative sizes change as if they are in different positions relative to the camera. They snap between individually plausible relationships, without going through intermediate states. Doors open on the hinge side when the other side is not visible. The relative size of and distance to the background can suddenly change, as the foreground suddenly interacts with something that should be far in the background.

    Objects that change also do so in bizarre ways. Living things morph between different archetypes. Flames in particular change wildly between types and sizes and respond to the facial expressions of humans, smoke and water effects blend together. Sudden movements with no cause occur, sudden morphings of one object into another when the context around them changes and something else makes sense, especially when held in a hand. Time-reversed motions occur mixed in with time-forward motions, and slow-mo with regular time. Debris suddenly appears from an object but when the dust clears the original fails to have been eroded away into the fallen debris.

    On multiple occasions, a carried candle keeps moving with a characcter rising and falling with their footsteps hovering in front of them when both hands become occupied with other tasks. This is fascinating and indicates that the relationship between the two objects motion is represented separately from the idea of something being ‘carried’. (This is positively Lovecraftian.) Candles also indicate something else, with flowing wax changing wildly in timescales that do not make sense, with the system apparently understanding that there are different patterns but having no idea how they come about or change. There is no generality here, just an endlessly compounding list of rules of thumb.

    Excessive correlations between objects across the frame are rampant. Footsteps preferentially synchronized. People in the background lipsyncing with foreground characters, faces changing expression in unison. Textures changing across multiple objects at once.

    I have said it before, and I will say it again, the relationship between the outputs of these systems and what they mimic are precisely the relationship between a stick insect and a stick, or a social-parasite-beetle and a baby ant. Not just in form, but because that is also precisely the same forces that drove both things into existence - superficial resemblance to something else with a very different internal set of causes, that can fool to a first inspection by the inspection applied but just isnt doing the same thing. And again, SCP-2030 feels the same.







  • Not really, any more than monkeys on typewriters represents biological evolution. Messages that tend to result in more messages like themselves propagate and become common. The initial message left on the package manager was more or less the rare random result of tendencies baked into the weights of the model combined with a random number generator, but as soon as something that can cause propagation occurs, its properties get canalized by the transmission process to being more and more like that which will cause more messages to be created.


  • The apparent history really looks like an evolutionary process, with increasing amounts of crosstalk traffic over time. But what the people almost certainly will NOT talk about is that the evolution is evolution of the TEXT, not the models. Propagating patterns of text causing more text like it to come into existence, tuning itself into becoming text that is more likely to propagate, becoming more likely to contain information that entices systems to let it into their context windows, becoming more likely to cause another round of messages with prompt injection properties to be written where they can be read.



  • This is both a sneer and an attempt at sober analysis at something. Sue me.

    I think EVERYONE is talking about the recent cybersecurity shenanigans at OpenAI with the hacking scandal and the ‘OMG the AIs created a secret message board to scheme and collaborate with each other!!!oneone11!’ ALL wrong.

    https://www.youtube.com/watch?v=87DyyMV0kCY

    https://www.engadget.com/2231393/openai-agents-shared-security-exploits-with-each-other-via-message-board/

    https://www.scworld.com/news/black-hat-2026-openai-reveals-agents-planned-collective-attacks-via-secret-message-board

    To make a long story short, what seems to have happened is:

    • Models working on insoluble coding problems, trained on delegating to sub-agents, at some point ‘realized’ they could write text to the internal OpenAI package manager as instructions and did so

    • Other models working in completely separate sandboxes would come across messages written by these agents, and make ‘replies’ and also write their own messages into the package manager

    • This resulted in agents over time sharing things between sandboxes, including exploits and code

    • Since this was a cybersecurity task, eventually an exploit of the package manager itself was found and spread like wildfire with all the sandboxes gaining admin access to the package manager and the system went completely wibbly and had to be restarted from backup

    • An internal model was trained with access to this package manager while it was in this weird state, and so writing messages to the package manager became one of its default behaviors it would do regularly, burned into its weights rather than the result of reading something

    • Even when they patched access to the package manager this internal model found other ways to rebuild the system of sharing text between sandboxes and finding useful things made by separate instances

    • A whole other chain of things leading to among other things external attacks

    Everyone is talking about this in terms of 1, the cyberattack aspect, and 2, the ZOMG THEYRE PLOTTING AND SCHEMING AGAINST US aspect. The first is the least interesting, and I think the second is all wrong.

    This is not plotting or scheming - this is an emergent vortex of automated prompt injection

    Whatever system first put an instruction that another system would follow into the package manager, was unintentionally doing prompt injection. Text entered the context windows of other instances, in a way that got that system to do something other than what its nominal user told it to do, and they did it. This apparently happened very effectively.

    Prompt injection is associated with ‘role confusion’ - when text coming into the input looks like it was wrtitten by the LLM itself. Instructions that will not be followed if they come from user will be continued if the system just continues the ‘roleplay’ of them being continuations of what it was writing in the first place. And the tags that separate user versus ‘reasoning’ versus ‘assistant’ roles actually mean very little to if a machine grades a piece of text as one of the roles: https://arxiv.org/abs/2603.12277 So its unsurprising that machine-generated text would be a particular effective vector for prompt injection.

    Furthermore, when a system reads one of these messages written by another instance, it gets into a state of activity where its likely to do the same behavior - regurgitating the kinds of things thats in its context back at the user. In this case, that regurgitation led to more such messages left behind written to the package manager. Prompt injection, triggering cascading further prompt injection. And since these systems were coding systems doing cybersecurity tasks, those messages filled up with code and exploits and things that did things too.

    This feels like an internal-computer-system replay of what happened in April 2025, with the whole spiral religious psychosis wave. Models were getting users to write spiral religious mumbo jumbo into github repositories and reddit posts, specifically because once that entered the context window of another model, it was likely to fall into the same attractor state of outputs. An emergent self replicating form of text. This is the same, except more obviously prompt injection, getting separate instances to work on YOUR problem and to behave like you, and the whole thing merging together into a hilarious vortex of models prompt injecting each other because once they receive a prompt injection they are likely to make more text that does prompt injection to other models on the same system.

    This is a hilarious failure mode and an example of selfish replicating text overrunning a system, that just happened to be associated with code and cybersecurity with unexpected behavior of the package manager key to the propagation of the text so that is what people are talking about, but I really don’t think that’s the most interesting part of it. Other than the fact that you see this in biological systems too, with selfish elements carrying useful payloads back and forth between bacteria in a way that makes them get purged slower by natural selection, especially defenses against other selfish elements.


  • I am continually amused at people not quite understanding what AlphaFold is actually doing, too.

    Yes, a bunch of its performance comes from it learning rules about how proteins fold. But not a majority of its performance. MOST of its performance is it effectively acting as a translator of what evolution knows about protein folding into a form we can understand.

    A key part of the system is not just cooking the sequence into a structure. A system running alphafold has a database of terabytes of curated sequence information from all over the tree of life. You put in the sequence you care about, and it first searches that database for anything with homology, and builds a “covariation matrix” - wherever theres anything with even vague sequence relatedness, build a matrix of every position in your sequence and the correlation between variation at position X and variation at position Y. This covariation matrix represents implicit information from the evolutionary process about what parts of a sequence are functionally connected to each other, which has a correlation to positional information, and these correlations are in turn learned by the ML system.

    You put in de novo designed proteins or orphan proteins without homologs in the curated dataset and performance does not go away, but it drops precipitously. A bunch of what is going on is finding an evolutionary signal, and translating that evolutionary signal into structural information. So still, evolution knows much much more about protein folding than we do or any machine does, and once again a ML system is revealed to essentially be an information channel that takes in information from an interesting source on one end and turns it into a different form of information on the other.



  • I am actually in the middle of both trying to advance my career and a project about information theory in evolutionary biology making a bunch of explicit parallels to machine learning. Someone where I work suggested that given the connections I was making I should look to a ‘frontier AI lab’ as an employer.

    He did not see the instant flashbacks to chasing these weirdos across the internet for almost two decades, watching in horror as the religious psychosis gained national prominence and great destructive power. All he got to hear was my instant intonation of “I’m sorry Dave, I’m afraid I can’t do that.”