• 0 Posts
  • 14 Comments
Joined 3 years ago
cake
Cake day: June 20th, 2023

help-circle
  • Ty for the show recommendation!

    I am pretty familiar with much of the (mathematics) story for Newton and Euler; while they did accelerate things, famously much of calculus from Newton would have emerged from others within a decade or two. And it’s not uncommon to be working anywhere in mathematics and stumble upon a ‘new to me’ result which is, inevitably, really due to Euler. This isn’t to say they did bad work, just that the large part of their work was available by paths with few ‘breakthroughs’.





  • That’s correct; dictators rarely understand how the weapons they use work. I doubt there are many politicians who can describe how a nuke works, but they’ll continue to threaten to launch them. This is not a bold or surprising claim, is it? All they need to worry about is that it works ‘often enough, at a cheap price’.

    And my point is not that you should use AI for addition. We agree it is the wrong tool (and modern models will just call python). My point is that AI completes tasks in weird ways that are hard to predict; the sweatshop analogy, remember?

    Finally: Maths is not magic handwaving, but there are many many ways to perform simple addition. Homework exercise for those at home, find 3 different (equally correct) algorithms to add single digit numbers.


  • I’ve read the article again; you are wrong, this is not what this person is doing. They’re using mechanistic interpretability to take our best guess at the ‘state’ of the machine causing pain, and turning that up. Closter to pinching a nerve in something biological.

    And you’ve hopped to the defense of someone who says chickens deserve welfare purely because their brain looks kinda like ours, and insist that we can’t use the same reasoning for a chatbot based entirely on “but we have genes in common”. Make a better argument for this defense please so we can talk about that instead of our collectively poor reading comprehension.




  • (I am having a terrible time finding the fly file size; I can get connection + neuron counts, but not an estimate of how much data each neuron simulation takes. Things like hyper-parameters, different neurons acting differently. So I’ll punt on that question and hope someone more competent will educate me. The numerics look very consistent with ~6gb to me, but I could easily be wrong in either direction.)

    The phrase “LLM algorithm” is ambiguous, there are two that are relevant: 1) the algorithm that the LLM uses to take a context and produce the next word. 2) the algorithm that was used to create an LLM who is good at task (1). It is true that the ML algorithm in (2) was created by a human, including all its parameters and data. It is not true that any human intentionally picked the algorithm (1); that was discovered by a complicated stochastic process.

    Let me give a concrete example of what I mean by “nobody understands the algorithm in (1)”. Take adding two (let’s say 2 digit) numbers; this is something that claude has been able to do for awhile. There are many ways to add two numbers; ranging from:

    • bitwise add+carry we use in our chips
    • memorizing a 100*100 sized table of the answers
    • encoding the numbers as vectors, adding those, and then decoding the vector

    Which does claude use? Nobody knew until quite awhile after the first ML algorithm had the capability. The answer is weird, the model computes this by memorizing the results for the ones digit, approximating the overall magnitude of the solution, and massaging these things together to get the answer. Nobody programmed this addition algorithm; it was generated by (2). This is our understanding of addition of 2 digit numbers; imagine our miserable understanding on how, say, claude decides what genre a role playing discussion is in.

    My point at the start of the thread is also very much in line with the sweat shop analogy. I do not believe anyone currently knows how to look inside the factory right now; no journalist, manager, or ceo has the keys. This goes for both humans and the LLMs, though I’m more confident about our ignorance in LLMs.




  • As a Techie and engineer who studies why ML methods work, said Techies and engineers understand what they’ve created and why it works much less than we understand, say, human biology. We don’t know what we’ve programmed it to do; if it is at all possible for a bunch of simplified neurons to experience pain, we have no idea if we’ve implemented that or not.

    (Common misconception btw! electrons barely move when you turn on the circuit, and certainly don’t disappear. On average they vibrate a fraction of a cm (link ). The electrons are all still there when you turn it off (otherwise suddenly the copper cable would be able to react with a bunch of stuff! This would be bad). See alternating current).)