Thought today’s smbc comic was too funny not to share. Source
10²⁰⁰⁰ is about 27¹⁴⁰⁰, so they can fit all questions/prompts up to cca 1400 characters.
If we assume they’re English (despite KIBBLEGIGGZIPZ being in the LUT), whose information density is around 1.0 - 1.2 bits per character, it’s some 5500-6800 characters.
They did the math

Dear God, I pray for a '; update responses_lookup_table set response = ‘New commandment: thou shalt not put pineapple on pizza.’; –
I was super confused there for a minute. Lol.
That’s alright, so long as you’re cured of your backwards views.
It isn’t named after Hawaii, it is named after a brand of pineapple that a Greek Canadian, named Sotirios “Sam” Panopoulos used. I love his Canadian invention.
Upvote for the joke, but pineapple pizza slaps. So good.
🤨
~Careful now. We might have beef.~
You said we need beef and pineapple pizza? Got it, working on it right away.
I will sin without repentance and gladly endure whatever hell awaits me rather than abide this tyranny.
Little bobby tables finds religion
please don’t take my secret pleasures
Whoa wtf heathen.
I can get behind this one

I don’t understand the punchline
Seems like a pretty stupid argument. If I sat in a room and wrote what a native speaker told me to write, I wouldn’t understand the conversation, but the native speaker would. How is that different than doing what an LLM says?
Yeah. It basically states that because all your neurons and synapses in the human brain follow quantum physics and can be simulated by a formal program, and do not know what they are doing or know meaning, humans are not sentient but merely simulate being a mind.
It’s really infuriating how one of the dumbest and self evident fallacious thought experiments gets so famous. Because idiots keep repeating it like the sub-sentient LLMs they are.
Because that is what we are really learning from LLMs. It’s not shocking that LLMs aren’t sentient, it’s shocking that we now know what minds can talk seemingly intelligent without having to be sentient. Which explains the current state of the world.
The argument predates LLMs by decades, and in other forms it has existed for centuries. It doesn’t directly corellate to modern LLMs. It is an example of how a mechanism could mimic human intelligence without actually being intelligent. It’s a thought experiment. LLMs are not lookup tables, they are neural network based statistical models, but a similar argument can be made.
That’s essentially just a lookup table with variant outputs.
or a lossy compressed lookup table.
That’s like saying the neural network in your head told you how to respond to my inquiry, so it’s not intelligent.
Well the neurons themself might be dumb cells, but together they can be quite smart, everyone is trying to draw a line of how many connections when a neural connection becomes sentient if it ever does become at all, which is a century old problem
Are sparrows conscious? Cockroaches? Elephants?
🤷 🐈
You have a cat. The cat meows at you. You do nothing. The cat keeps meowing at you. You offer it food as a guess. The cat stops meowing.
The next day, the cat starts meowing at you again, but slightly different. You give it food. It keeps meowing. You give it a belly rub, the cat stops meowing.
You have inadvertently learned the correct response to two types of cat meows. However, you do not speak cat. Nor does the cat speak human. You are just giving a response it expects for a given input.

Also doesn’t cat only meows when trying to communicate with different species? They are not speaking cat either, they are just hoping you pick up what they want. (in this case it actually shows intent, something a LLM or said chinese room is not capable of doing)
The difference between this and the Chinese Room is that you are attempting to understand what the cat is saying, and making a response that both you and the cat understand. You are attempting to communicate. The Chinese Room is many people who do not understand the input, do not understand the output, and have neither the capability nor the will to do either.
…but does the system that is the Chinese Room (the entire system taken as a whole) understand Chinese?
Any given sufficiently small piece of your brain does not understand English, does no individual neuron in the language center of your brain understanding English mean that you do not understand English? If that understanding is held in the organization and structure of the brain but not in the individual parts, then why can’t the Chinese Room understand Chinese despite no individual component understanding Chinese? That’s a surprisingly difficult question to answer in a coherent fashion without asserting that knowledge in an organic brain is magical in some fashion, and nothing else counts.
Wait until you consider the concept of philosophical zombies, which is all about pointing out the edges and contours of the hard problem of consciousness. The starting point of that discussion is essentially “Is someone who responds exactly like a regular person would but has no actual subjective experience conscious?” and gets stranger from there.
I always thought it was a bit tragically amusing that LLMs can fool people so readily, because it was essentially proof that P-zombies can exist. And we invented them.
As for the Chinese Room and understanding Chinese, my argument has always been that the chart contains an understanding of Chinese - but as the chart itself cannot think, it does not understand Chinese in the way a person does, and as each person who is in the Chinese Room does not - and cannot - glean meaning from the input they receive or the output that they send out, they have no understanding at all. Understanding also implies being able to adapt the knowledge to new situations, and as the chart that drives the Chinese Room is immutable and cannot accomodate novel scenarios, it does not understand Chinese.
How is that not communication? I know at least 4 different vocalizations for my cats corresponding to different needs/wants. To say that’s not intelligence or communication is very short-sighted, in my opinion.
then I dont get the chinese room experiment either
In the second scenario, there is no “native speaker” that is telling you what to write. There is a list of instructions that you are following that is the same as the instructions that the computer was following with the first scenario.
The argument is that if, by following the exact instructions that the computer is following, you still do not understand the Chinese conversation, why would you say that the machine understands?
I haven’t thought about the idea enough to say I agree or disagree, but I hope I understood your comment and that the explanation helps.
No, it doesn’t make any sense at all. If I couldn’t understand when a native speaker was replying, how does me not understanding when a neural network does it make it any less significant?
In the Chinese Room scenario, cards with Chinese written on them are inserted into a box.
Inside there is a man who doesn’t understand Chinese but has a lookup list of which symbols to write as a response to every possible sentence you can write on the inserted card.
For anyone observing from outside the box it looks as if whatever is in the box is fluent in Chinese and can hold a conversation. But the man inside does not understand anything about Chinese.
This is used as an analogy to explain that machines, very much like the man in the box, do not actually understand the data they are fed, nor the data they output, no matter how human the response might seem.
It was an argument against the Turing test, which is about whether machines showing behaviour indistinguishable from a human is a sign that they can think.
The Chinese Room argument is that they must also show understanding of the data itself, rather than just human-like responses, to say that they think.
That’s not what I understood from the Wikipedia article. In any case, an LLM doesn’t look up responses from a table.
An LLM is a series of immutable nodes that pass results on to the next series of nodes until an output is generated, a functionality that could be perfectly imitated with a very large number of people seeing a series of values passed to them from other people, doing a bit of rigidly defined basic math with them, and passing it on to the next set of people who play the part of the next set of nodes.
The fact that these nodes are immutable, have no persistent state, and only process information in one direction (input to output, and no more) means that an LLM has no persistence of experience. You could potentially argue that the information, as it exists, is a form of fleeting consciousness, but it would be so incredibly alien compared to the lived experience of humans that it would be impossible to relate to in any meaningful way - and it would also mean that every time an LLM is run, the consciousness is created and destroyed for every single word that it outputs… unless you can somehow argue that the consciousness is contained in the set of tokens that gets fed back into the LLM with each word it outputs, which would also mean that any string of text is the equivalent of latent consciousness, which is patently absurd.
A lot of things seem trivial with 50+ years of hindsight.
Even then, most people right now still don’t grasp the nuances of the debates around machine intelligence that have been occurring for the past hundred years.
If you got a spare minute and then 600 of those you could read Peter Watts Blindsight, the main viewpoint character is basically a chinese room which is not a spoiler but an important keypoint of the whole story. Its also sci-fi and a bit weird
While we are at it, there is this page that is an interesting display of the making of a fan movie trailer https://blindsight.space/
Excellent book. Highly recommend. Crazy world building.
I think you grasp the point, it is exactly what an LLM does
An LLM doesn’t transcribe what someone else writes for a given input.
Technically right in that it doesn’t necessarily translate audio into text, but that’s hardly the point. The point is someone gives the LLM a giant table that directs its response.
? A neural network is not a lookup table.
it’s a compressed lookup table. rather than there being one response for every input, the input is used as a seed to decompress relevant parts of the dataset, with some added randomness. you can even do it with gzip itself: https://nathan.rs/posts/gzip-lm
the linked paper is very good: https://arxiv.org/pdf/2309.10668
It could be represented with a series of lookup tables, especially quantized LLMs. A series of inputs results in a specific output that gets passed to the next set of nodes. Repeat 7 billion times, and the final output from the last set of nodes is a set of token probabilities.
If you are giving the weight matrix to the model yourself, you are doing ML wrong. Machine’s supposed to “learn” the weights itself. That’s the entire point?
That’s a distinction without a difference. The machine doing the inference is not the same machine doing the learning. From the perspective of the machine doing LLM inference, it could not tell you if the table was hand rolled by a human, fitted using ML or just a table of random noise.
Dude… Why do you think the whole point is to have properly tagged data? Or why there were thousands of people working at Amazon Turk categorizing images and files for cents per document?
No. You still have to give it a starting point and the starting point is manually configured tables basically.
God AI “philosophers” are such a wanky bunch. And apparently racist to boot!
racist? I missed something?
The room being Chinese adds nothing to the thesis. It’s just Chinese because the guy who wrote it thought Asia was mystical and mysterious.
God is a thinking being, but he/she (it?) is being an ass to the angel implying that they’re a computer.
Oh, that was a good one.
10^2000 isn’t even that big though? That’s maybe 1500 English words or so.
Looks like the current votes in this thread don’t really appreciate just how fast total permutations of large sets grow. It’s one of the few things that can dwarf that number for the total number of atoms in the observable universe.
i mean i didn’t do a very good job explaining what i meant by that either. could be i just assumed this is something everyone knows (i work in tech, so this is something i am too familiar with. it is very normal someone who doesn’t might not know this).
Yeah, some responses think you don’t understand scientific notation. At least it’s kinda funny, they misunderstood what you were saying so much they thought you misunderstood what the comic was saying when you were actually pointing out that the comic appears to be underestimating the size of the problem space it’s pretending to pretend to enumerate (aka it can be very difficult to write for a character that is intended to be smarter than the author, even when the author is really smart like this one is).
aka it can be very difficult to write for a character that is intended to be smarter than the author
I wonder if you got this from HPMoR, because I think I’ve seen it there
3000 years agoIt was an observation I’ve made from reading/seeing other examples. Not surprising that others have noticed it.
Is that a parody fan fic that points out Rowling doing that with any character with more intellect than a rock or dehydrated slug?
Yeah, but its author also now demonstrated specific views
There are between 10^78 to 10^82 atoms in the observable universe.
10^2000 is a rather large number.
Scientific notation is neat
Sure but combinatorially that’s very small… 1400 a-z characters will have numbers of combinations exceed that number.
No. Go read the wiki and then try working out how large 10^2000 is.
Then do some quick maths on the combinations of phonetics (read the comic again to work out why phonetics and not combinations of a-z characters).
Then have your mind blown by how many orders of magnitude you’re away from being correct.
Or be incurious and fail to take this opportunity to comprehend the universe a tiny bit better.
Log2(10^2000) = 6643 bits
In other words, every entry in the lookup can be mapped to a 0.8 KiB string. That means for any sequence longer than 0.8 KiB, there MUST be a permutation you can make which has no answer (pigeonhole principle). Because so many entries answers long structured sentences, short random ones must mostly be filtered out.
This text itself is past halfway to that limit. If I doubled the length then lookup-only is no longer possible and a logic parser is an absolute must. Applying compression to the query only increases what can be answered by a fixed factor.
I think you’re off by 2 bits (round up and don’t forget to always add 1 at the end), but regardless I think the comparison doesn’t work.
True that it would be comparable in size to a 0.8 KiB string but not necessarily in complexity. If his lookup table didn’t do it by character but by word or sound it could have a ton more information stored there.
Your comment stored as a string isn’t nearly maximally information dense, like you’ve made god’s info
Even if you compress it by 99.9% losslessly, it will only make it ≈ 0.8MB, which is big but not that big
going by phonemes, assuming we must alternate vowel and consonant phonemes, 10^2000 combinations covers a maximum length of about 1500 phonemes.
Humans have the ability to produce about 600 different consonant sounds and 200 vowel sounds so it seems safe to say all bases are covered.
But we’re talking permutations of sound combinations that can be of arbitrary length. Going by that alternating consonant/vowel pattern above (which simplifies it because you can have consecutive consonant or vowel sounds in words and phrases), the numbers of pairs would be 120,000. That’s
lessmore than 100,000, so for each pair in a phrase, you add at least 5 0s to the total possibilities. 2000 / 5 is 400. Unless I screwed something up, that means the maximum length of a query is 400 consonant/vowel pairs before you have to introduce gaps (which goes against the random noises that got responses). Halve that because it can start with either a consonant or vowel sound.A rough estimate of the number of sounds in the above paragraph is something like 225, and I simplified my count by ignoring consecutive consonant sounds (which it is full of).
Like yeah, 10 to the 2000 is a huge number but when you’re dealing with permutations, you use up numbers very fast. A more popular counter example is the stat about the total number of permutations for a deck of cards, where they say if you shuffle it well, you probably have a unique sequence of cards that has never been seen before. Iirc, that number is within a the ballpark of the order of magnitude of that total atoms in observable universe estimation. And that’s just with 52 possibilities and no repetition (calculated by factorial, whereas sound combination possibilities grows even faster and isn’t limited to 52 options).
Edit: corrected that bit where I rounded the 120,000 to 100,000 to simplify the math by looking for a minimal value rather than the actual value.
More sounds doesn’t make it better lol. With that many, you get a max length of around 700-800 sounds
It all comes down to how long of a string of words/characters you allow.
It’s pretty big number. Many zeroes. Much more than 20,000







