As we continue developing our software, we accumulate a growing amount of technical debt just to keep the system running. But I believe we are on the brink of an even larger issue. Cognitive debt.

Hope you enjoy this reading, all feedback is welcome.

  • Shin@piefed.socialOP
    link
    fedilink
    English
    arrow-up
    1
    ·
    2 days ago

    No, the random part happens in the query.
    There is also random in the training, but the query also generate more random numbers.
    Otherwise this would be a deterministic procedure, and it’s not.

    • fruitycoder@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      1
      ·
      2 days ago

      We are saying the samethings but you are adding no to it.

      Right, during inference random numbers are generated, that plus the numbers from input are added to the weight values and the matrix multiplication happens. If you used the same random numbers and inputs it is deterministic. For regular use you don’t do that because you want a stochastic output, if you wanting to do forensics and trace what led to an output you would benifit from that determinism.

      • Shin@piefed.socialOP
        link
        fedilink
        English
        arrow-up
        2
        ·
        2 days ago

        Gotcha,

        But this means we need to provide not only the same query (and be sure that the tokenizer is the same) but also provide the seed for the number generation. In this case we will have a deterministic outcome. (Unless we provide the list of numbers used, which for me feels wasteful)

        But at this point none of the providers have this feature.

        And I don’t think the open source have this also. But open source can be updated/changed.

        • fruitycoder@sh.itjust.works
          link
          fedilink
          English
          arrow-up
          1
          ·
          2 days ago

          Not that I see either looking at opensource inference engines. vLLM supports setting the seed so if you have that you do, and maybe it’s just as simple as harness setting and recording that. At least for replay this request type of operation but that does not show where in model each random numbers would have been applied.