• 0 Posts
  • 11 Comments
Joined 3 years ago
cake
Cake day: September 2nd, 2023

help-circle
  • At some point you have to assume some kind of knowledge though. If you have to explain everything from the basics, the moment a concept is slightly complicated, the blog gets thousands of pages long.

    You don’t need to explain how an integer is represented in binary in memory in a blog post about atomic integers. Most people that read about atomic integers already know how integers are represented in memory.






  • Good practice != Something that ensures security. As I said, it’s an even better practice to just host your own fork if what you want is security. The reason to use git hashes is mostly so you don’t have to trust the author following semver in there version number. It’s a matter of ensuring that your code will always compile, not a security feature.

    Have you read the article? Your arguments derive from an assumption of “sha1 is mutable, sha256 is immutable”. First of all, every hashing algorithm is going to have collisions, that’s an unavoidable “feature” of hashes. And sha1 is in no way “mutable”, it takes a great amount of effort (and money) to generate a collision. Furthermore, those collisions are not arbitrary. You have to calculate them beforehand. As the article says: what is more likely? Paying 40k€ in compute time to generate a single collision? Or just paying an open source maintainer 40k€ to let you do 1 new commit that most people are going to download anyway?



  • If someone has access to write to a repo you are downloading code from, and you don’t trust that someone, you shouldn’t be running the code in that repo.

    The only “security” vulnerability is: I have audited and verified that the code in commit 05682ab36ca. And I will use only that version.

    What you do then is: download that commit and store it in a local repo, and package it so you can distribute it to your clients.

    Auditing external software is tremendous effort, even if it is open source. Using your local repo instead of downloading that commit from GitHub every time is very little extra effort in comparison.

    And if you really need it. You can always hash with sha256 yourself and verify that the commit with that sha1 still produces the same sh256. So you can keep downloading it each time, just need to store the sha256.



  • I tried to implement my own git from scratch. And when I got to choose a hashing algorithm, I reached the exact same conclusion as the author pretty fast.

    The hashing algorithm doesn’t really matter as long as it’s good enough to prevent accidental collisions. And if there is a collision, it is extremely easy to check.

    If you are about to create a new object and an object with that hash already exists, you compare the 2 objects. If they are the same, nothing happened. If they are different, you just found a collision. Throw some obscure error that will only be witnessed once in the lifetime of the universe and be done with it. Tell the user to change a single bit of the input and be done with it.

    The purpose of object hashes in git was never to provide security. Security is achieved through other means.


  • Can confirm. I have to do this all the time. Some coworkers just delete it, and I have to explain the situation every time when they see in the git blame that I put that semicolon there.

    As someone pointed out in another comment, in switch statements you should just start a new scope with {}. But in actual labels it’s just better to put a load-bearing semicolon. You don’t need to put it in another line though, I just put them like this:

    int main() {
        int *a = malloc(sizeof(*a));
        if (init_a(a) <= 0) {
            goto err;
        }
        do_something(a);
    err:;
        int ret = 0;
        if (a == NULL) {
            ret = -1;
        }
        free(a);
        return ret;
    }